Skip to content

For GA4 usersFrustrated with GA4 attribution? Upload your GA4 export, see causal insights in 5–10 minutes for €99 pay-per-use.

ROAS & Incrementality

5 min read

Incrementality Testing for Ecommerce: A GA4 Playbook

Most incrementality tests fail before they start: the effect is smaller than the noise, or the geography cannot clear significance. This playbook runs the arithmetic first, then the test, on GA4 and order data you already own.

Share
Quick Answer·5 min read

Incrementality Testing for Ecommerce: Most incrementality tests fail before they start: the effect is smaller than the noise, or the geography cannot clear significance. This playbook runs the arithmetic first, then the test, on GA4 and order data you already own.

Read the full article below for detailed insights and actionable strategies.

The attribution problem

One sale. Four channels. 400% credit claimed.

100
1 sale
Meta
100%
claimed
Google
100%
claimed
TikTok
100%
claimed
Klaviyo
100%
claimed

Reported revenue: 400 · Actual revenue: 100 · Gap: €300

An incrementality test for an ecommerce brand has five steps, and the first three happen in a spreadsheet before any spend is withheld: qualify the channel against the smallest lift your revenue noise lets you detect, count the regions you can test on, fix the pre-period, then run at least eight weeks and register the design before you look at anything. Skip the first three and the test produces a confident report that was decided before launch.

This playbook uses two inputs you already have: order data from the store, and the GA4 export. It follows The Price of Being Found, Chapters 15, 16 and 20, and it is written for the person who has to defend the result afterwards.

Step 1: qualify the channel

Two numbers. A is the channel's spend divided by total revenue. iROAS is the incremental return you would honestly defend to a CFO, not what the platform reports. Multiply them: that is roughly how much total revenue would move if the channel were switched off.

Compare the product to the smallest lift your design can see. The book's geo table, for the standard design of 26 weeks of history and 8 weeks live, gives 8.3% for a typical DTC brand whose weekly regional revenue wobbles by about 8%, 6.2% for a steady one, 12.5% for a lumpy one. A channel at 5% of revenue with an honest return of 2 moves revenue 10% and clears the floor. At 2% of revenue it moves 4% and does not. The book's verdict on the second case: the test cannot answer the question, and a null result will be read as an answer anyway.

Your wobble is one spreadsheet formula: for each region, the standard deviation of weekly revenue divided by its mean, averaged across regions. Can a brand your size measure Black Friday lift at all? reproduces both of the book's tables.

Step 2: count your regions

With N regions and one treated, placebo inference compares the real result against N minus 1 pretend ones, so the smallest p-value the design can return is 1 divided by N. Twelve Dutch provinces give 0.083, which is above 0.05 before a euro is spent. Forty COROP regions give 0.025. Fifty US states give 0.02. Treating two units instead of one raises the number of placebo assignments and lowers the floor at the cost of a smaller control pool. A Dutch province geo test cannot reach p below 0.05 has the table by geography.

Step 3: fix the pre-period and write the design down

The pre-period is what the model uses to work out what would have happened. Twenty-six weeks is the standard; fifty-two is better. Fix it before you look at anything. Then write one page: hypothesis, treated and control units, start and end dates, the minimum detectable effect from step 1, the analysis method, and the decision rule for each outcome. Put it somewhere you cannot edit. The book calls this the cheapest defence against the thing that destroys most marketing experiments, which is not bad statistics but reinterpreting an inconvenient result.

Step 4: run it, eight weeks minimum

Withhold the channel from the control units and change nothing else. Do not add a promotion in the treated regions, do not shift budget from the held-out regions into the others, do not end early because the first fortnight looks good. Twelve weeks if you can bear it. If a peak sits inside the window, the baseline stops meaning anything; the last day to start a holdout before Black Friday is 2 October works through that case.

Step 5: read it against the design, not against hope

Three outcomes. Lift above the floor with an interval clear of zero: scale with a number behind it. Interval includes zero and the effect you expected was below the floor: you learned nothing, say so, and manage the channel on stated judgement. Interval includes zero and the expected effect was above the floor: the channel is not doing what the report says, and the budget it held is the only money that was ever going to move on evidence.

Report all three the same way: point estimate, interval, design, and the floor. "12%, with an interval of 4 to 20, from a design whose minimum detectable effect was 8.3%" is a sentence a CFO can act on. "12% lift" is not.

Where the GA4 export fits

A holdout answers one channel per test. Between tests you are running on a model, and the book's advice is to say so and to re-anchor quarterly. A causal read on the GA4 export is that observational counterpart: per-channel incremental ROAS with confidence intervals from data you already hold, in 5 to 10 minutes, which is the fastest way to choose which channel earns the next holdout and to see the rest with their intervals stated. It is not an experiment and it does not claim to be one.

What to do this week

  • If you have to defend the number: compute the wobble and A times iROAS for every channel, and bring two lists to the owner: measurable at current scale, and not.
  • If you own the budget: pick the one channel with the largest margin over its floor and book its start date. One qualified holdout is worth more than three that could not have found anything.

As of 9 September 2026. The detectable-lift table, the A times iROAS rule and the placebo floor are from The Price of Being Found (Edition 2.10), Chapters 15 and 16, and carry the book's caveats.

Get attribution insights in your inbox

One email per week. No spam. Unsubscribe anytime.

Key Terms in This Article

Related Articles

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Ready to see your real numbers?

Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.

Full refund if you don't see value.

Stay ahead of the attribution curve

Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.

Which one are you? Optional.

No spam. Unsubscribe anytime. We respect your data.

Frequently Asked Questions

How long should an ecommerce incrementality test run?

Eight weeks live is the standard design in The Price of Being Found, on 26 weeks of pre-period history, and twelve weeks if the business can bear it. Shorter tests raise the smallest lift the design can detect: four weeks live moves a typical DTC brand from 8.3% to about 11%.

What data do I need for an incrementality test?

Weekly revenue by region for the pre-period from the store or finance system, the channel's spend as a share of revenue, and a count of the regions you can treat and hold out. The GA4 export is not required for the test itself; it feeds the observational read between tests.

Can I run an incrementality test without pausing a channel everywhere?

Yes. A geographic holdout withholds the channel from a subset of regions while spend continues elsewhere. What matters is intensity, not budget: the book notes the size of the holdout drops out of the detectability arithmetic to a first approximation.

Related reports

Real reports on this topic.

Anonymised reports from the Attribution Report Library tagged with roas & incrementality.

Browse all related reports

Find your wasted ad spend in 5–10 minutes.

Watch the model work on a sample store first, no signup. Then upload your last 40–90 days of GA4 sessions and get incremental ROAS with confidence intervals. No pixel, no SDK. €99 per read.

Prefer to talk it through? Book a 20-min call, or read how it works.

Last-click guesses.We run the math.

Causal attribution for ecommerce brands. Watch the model work on a sample store first, then upload your GA4 export and see which channels really drove revenue in 5–10 minutes. €99, pay-per-use. Pro at €299/mo when you want it continuous.

No signup for the demo. Book a 20-min call or compare plans.