Skip to content

For GA4 usersFrustrated with GA4 attribution? Upload your GA4 export, see causal insights in 5–10 minutes for €99 pay-per-use.

Causal Inference

4 min read

Incrementality Testing vs A/B Testing: What Each Proves

An A/B test tells you which variant wins. An incrementality test tells you whether running the thing at all beats not running it. Teams routinely use the first to answer the second.

Share
Quick Answer·4 min read

Incrementality Testing vs A/B Testing: An A/B test tells you which variant wins. An incrementality test tells you whether running the thing at all beats not running it. Teams routinely use the first to answer the second.

Read the full article below for detailed insights and actionable strategies.

Channel comparison

Reported vs. true incremental ROAS

Data relevant to: Incrementality Testing vs A/B Testing: What Each Proves

Platform reported
Causal (true)
Meta Ads+122% inflated
5.1x
2.3x
Email+167% inflated
12.0x
4.5x
Google Ads+62% inflated
6.8x
4.2x

An A/B test and an incrementality test both randomise, which is why they get confused. They compare different things. A/B compares two versions of an activity that is running for everybody in the test. Incrementality compares the activity against its absence, using a group that receives nothing at all. Teams routinely run the first and then answer a question that needed the second.

What each design contains

In an A/B test, group A sees version one and group B sees version two. Both groups are exposed. The result is the difference between variants: this subject line beat that one, this creative beat that one. Everything the two versions share is invisible to the test, because it is present on both sides.

In an incrementality test, the treated group gets the activity and the holdout group gets nothing. The result is the difference between doing it and not doing it. That is the only design that puts a number on whether the spend was worth making at all.

The shared baseline in an A/B test is exactly what an incrementality test is built to measure.

The substitution that costs money

A team A/B tests retargeting creative, finds that variant B lifts conversions 8% over variant A, and concludes retargeting works. The test established nothing of the kind. Both groups saw retargeting. If retargeting adds nothing over showing no ad at all, variant B is 8% better than variant A at adding nothing.

This matters most in the channels where the substitution is easiest to make: retargeting, brand search, abandoned-cart flows. All three are behaviourally triggered, all three sit close to purchase, and all three are routinely optimised through A/B tests without anyone ever testing them against absence.

What each one is good for

A/B testing is the right tool for creative, copy, landing pages, subject lines, offers and anything else where the decision is how to run something you have already decided to run. It is fast, cheap, repeatable and well understood.

Incrementality testing is the right tool for whether to run it. Cutting a channel, defending a budget line, answering a CFO who asks what happens if we stop, and any claim that spend caused revenue all require a holdout.

Neither is more rigorous than the other. They are answers to different questions.

What incrementality costs that A/B does not

Withholding has a real price. The holdout group is customers you deliberately did not market to, and if the channel does work you forgo that revenue for the duration. That makes these tests expensive and infrequent, and most brands can run one or two per quarter rather than one per campaign.

It also makes design discipline non-negotiable. Fix the window before you start. Choose the holdout at random rather than by any behavioural rule, since selecting on engagement rebuilds the bias. And compute the smallest effect the design could detect at your volume beforehand: if your spend share times an honest return sits below that floor, the test cannot answer the question, and a null result will be misread as proof the channel does nothing. Whether your channels are measurable at all has the arithmetic, and the playbook has the procedure.

What runs between tests

If you can afford one holdout per quarter, eleven weeks have no experiment in them. A causal read on a 40 to 90 day export estimates each channel's incremental contribution from the variation already in the data, reports an interval, states its coverage, and labels itself observational rather than experimental. It is the map; the holdout is the territory, and the read is what tells you which channel earns the next one. Incrementality versus attribution covers why observed credit cannot substitute.

What to do this week

  • If you have to defend the number: list every experiment run last quarter and mark which compared two versions and which compared against absence. Most lists come back entirely in the first column.
  • If you own the budget: pick the one channel where the answer would change a decision, and design a holdout for it. One is enough.

The interactive demo shows the read on a sample store, no signup.

Design and measurability material is from The Price of Being Found (Edition 2.10), Chapters 15 and 19, with the book's caveats.

Get attribution insights in your inbox

One email per week. No spam. Unsubscribe anytime.

Key Terms in This Article

Related Articles

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

See what you get

Confidence-scored results in minutes. Full refund if you don't see it.

Full refund if you don't see value.

Stay ahead of the attribution curve

Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.

Which one are you? Optional.

No spam. Unsubscribe anytime. We respect your data.

Frequently Asked Questions

What is the difference between incrementality testing and A/B testing?

An A/B test compares two versions of something that is running for both groups, so it identifies the better variant. An incrementality test compares running something against not running it at all, using a holdout group that receives nothing, so it identifies whether the activity adds anything.

Can an A/B test measure incrementality?

Only of the difference between the variants. If version B beats version A by 5%, that is the incremental effect of B over A, and it says nothing about whether either version beat spending nothing, because no group in the test was unexposed.

When should I run a holdout instead of an A/B test?

Whenever the decision is whether to keep funding an activity rather than how to run it. Cutting a channel, justifying a budget line or answering a CFO all need a holdout, because all three compare the activity against its absence.

Related reports

Real reports on this topic.

Anonymised reports from the Attribution Report Library tagged with causal inference.

Browse all related reports

Find your wasted ad spend in 5–10 minutes.

Watch the model work on a sample store first, no signup. Then upload your last 40–90 days of GA4 sessions and get incremental ROAS with confidence intervals. No pixel, no SDK. €99 per read.

Prefer to talk it through? Book a 20-min call, or read how it works.

Last-click guesses.We run the math.

Causal attribution for ecommerce brands. Watch the model work on a sample store first, then upload your GA4 export and see which channels really drove revenue in 5–10 minutes. €99, pay-per-use. Pro at €299/mo when you want it continuous.

No signup for the demo. Book a 20-min call or compare plans.