Skip to content

Holdout Test

A holdout test is an experiment that keeps a random part of an audience from seeing a campaign and compares its results with those of the part that could see it. The held-out part is the control group, so the gap in outcomes estimates what the campaign added.

By , Founder & CEOUpdated 7 min read

Run the numbers for your store: the free holdout test planner.

What is Holdout Test?

In a holdout test, a random share of the people a campaign would reach is deliberately left out of it. Meta's help center describes the holdout group as the part of the ad audience that is deliberately kept from seeing the ads, and notes that it is also called the control group. Because assignment is random, the two groups are alike on average, so the difference in their results estimates the causal effect of the campaign. Randomization does not make them identical: a single finite sample still differs by chance, which is why a result comes with an interval, a range of plausible values.

The unit that gets randomized can be people or places. Google Research describes geo experiments: regions that do not overlap are assigned at random to control or treatment, and ads are geo-targeted to match. For people, Meta's Conversion Lift uses an intent-to-treat design: the test group is everyone eligible to see the ads, including people who never actually see one, and the control group is not eligible. Meta also says its lift tests count all conversions in the test and holdout groups, and Google says its Conversion Lift does not apply the usual conversion tracking settings or attribution rules. So a holdout compares total outcomes by group, not only the orders a platform credits to the campaign.

Holdout size is a trade-off. Google says a larger holdout increases the sample but has a higher opportunity cost, because control users are not shown the ad, while a smaller holdout needs a longer study to detect a difference. A holdout test also answers a narrow question. A lift study measures what a campaign adds on top of any other active marketing, and the result describes the people and period tested: IAB notes that incrementality is not static, since shifts in competition, buyer behavior or tactics can change the outcome.

In machine learning the word has another meaning: a holdout is a random part of the data kept aside to test a model. That is a different use of the term and not an experiment on customers.

For illustration, suppose a store keeps a random 5,000 of 50,000 prospects out of a campaign, and 900 of the 45,000 who could see it order, against 90 of the 5,000 held out. The rates are 900 / 45,000 = 2.0 percent and 90 / 5,000 = 1.

8 percent, so the campaign adds 0.2 percentage points, which is 0.002 x 45,000 = 90 orders, or 90 / 900 = 10 percent of the orders in the test group.

A real test would also report how uncertain that estimate is.

Why Holdout Test matters for ecommerce

Platform dashboards and attribution reports show which ads and clicks were credited with an order. They cannot show how many of those orders would have happened without the campaign: the 2025 guidelines from IAB and IAB Europe on incremental measurement say attribution and ROAS describe what happened and cannot say whether marketing caused it. A holdout test measures that gap for one campaign.

The gap can be large. In eBay's field experiments, paid search returned only a fraction of what non-experimental estimates implied, because clicks on search ads and purchases are correlated, and brand-keyword ads showed no measurable short-term benefit. A 2019 Marketing Science paper on a set of Facebook experiments found that observational methods often did not reproduce the effects the randomized experiments showed.

Holdouts have costs: held-out people are not shown the ad, and a Quarterly Journal of Economics paper by Lewis and Rao reports that sales are volatile enough for informative advertising experiments to be costly or out of reach for many firms, so a small store may get an inconclusive result.

How to use Holdout Test

  1. Pick one campaign or channel and one outcome, such as orders or revenue per customer, and count the outcome the same way for both groups, for example from your store's order records.
  2. Choose what you can randomize: people on a customer or subscriber list you control, an audience inside an ad platform's lift tool, or regions. Use the unit the campaign can actually be switched off for.
  3. Assign by chance before launch. In a spreadsheet, give each unit a random number, hold out those below your cut-off, and save the list with the date. Do not assign by behavior, such as who opened or clicked.
  4. Withhold the campaign from the holdout completely and change nothing else for either group. Check whether other channels, devices or regions can still reach the holdout, and do not change creative or audiences mid-test.
  5. Fix the end date in advance and cover the lag between seeing the campaign and ordering. Google's guidance is to set the duration to capture the average conversion lag.
  6. Compare the groups: orders per unit in the test group minus orders per unit in the holdout, times the number of units in the test group. Report it with an interval, or with a test of whether the two groups differ. For a yes-or-no outcome such as ordered or did not order, the NIST/SEMATECH e-Handbook of Statistical Methods describes a z test of equal proportions for reasonably large samples and Fisher's exact test for small ones.
  7. Check the result. Pass: the groups had similar order history before launch, the holdout stayed clean, and the interval excludes zero by a margin that would change a budget decision. Fail: the groups differed before launch, the holdout was reached, or the interval includes zero. Then the answer is inconclusive, not zero.

Formula

Incremental orders = (orders per unit in test group - orders per unit in holdout group) x units in test group

Relative lift = (orders per unit in test group - orders per unit in holdout group) / orders per unit in holdout group

Where:
- orders per unit in test group: orders divided by the people or regions that could see the campaign
- orders per unit in holdout group: the same measure for the group the campaign was withheld from
- units in test group: how many people or regions could see the campaign

Under random assignment, the difference in group averages estimates the average effect of the campaign. Dividing by group size first puts groups of different sizes on the same footing. The estimate has sampling noise, so report an interval around it.

Common mistakes

  1. Choosing the holdout by behavior or convenience, such as people who did not click or the least engaged customers. They differ from the rest before the campaign starts, so the gap measures who they are, not what the campaign did. Cunningham's textbook, Causal Inference: The Remix, says that when people choose their own treatment, selection bias is the first thing to suspect.
  2. Letting the holdout be reached through another channel, device or region. Google says contamination of the control group lowers the estimated difference and reduces the reported incrementality, so a campaign that works can look as if it does not.
  3. Testing a subset of campaigns and reading the result as the value of the whole channel. Google says holdback users might still see the unmeasured campaigns, which reduces the chance of detecting lift.
  4. Reading a null result as zero effect. Google says a no-lift result does not necessarily mean the ads were ineffective, and random noise can produce a result with no lift even when the ads create lift.
  5. Treating one result as permanent. IAB says shifts in competition, buyer behavior and tactics can change outcomes, so a holdout describes the conditions of that test.

Frequently asked questions

  • What is a holdout test in marketing?
    A holdout test is a randomized experiment: a random group is withheld from a campaign and its results are compared with those of the group that could see it. Meta's help center uses holdout group and control group for the same thing.
  • How big should a holdout group be?
    It depends on how much data you want and how quickly, according to Google. Its lift tool accepts a holdback from 1 to 50 percent. A larger holdout gives more data but a higher opportunity cost, because control users are not shown the ad, while a smaller one needs a longer study.
  • How long should a holdout test run?
    Long enough to cover the delay between seeing the campaign and ordering, because orders placed after the window ends fall outside the comparison. Google says to set the duration to capture the average conversion lag. Its tool allows studies as short as 7 days but typically recommends more than 14 days.
  • Is a holdout test the same as an A/B test?
    Not according to Meta's help center, which says a lift test is not the same as an A/B test. In its description, A/B tests have no randomized holdout group, meaning no group of people deliberately kept from the chance to see the ads.

Further reading

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.