Skip to content

How long should an incrementality test run?

Usually at least two full weeks, and longer than your buyers take to decide. Cover your purchase lag in whole weeks and add a cooldown for late buyers. Then keep going until the platform's power estimate says the test can see the lift you need.

By , Founder & CEOUpdated 7 min read

Run the numbers for your store: the free holdout test planner.

Usually at least two full weeks, and longer if your buyers take longer to decide. Set the length to cover your purchase lag, run whole weeks, and keep counting late buyers for a while after the ads come back. If your store sells little, the test usually needs more weeks or a bigger holdout to show anything.

Ask the ad platforms and you get a floor, not a fit. Google lets a Conversion Lift study run for as little as 7 days, but usually wants more than 14. For campaign experiments, which pit one campaign setup against another, Google says at least 4 to 6 weeks, or longer with a long conversion delay. Meta's A/B tests run for 30 days at most, and Meta warns that tests under 7 days may produce inconclusive results.

Those numbers answer different questions. Four things set the length for your store: your buyers' lag, your weekly wobble, the lift you need to see and your calendar.

What one store's data shows

One store's anonymised GA4 export, 1 January 2024 to 21 August 2026. It holds shares of revenue only: no ad spend, no order counts.

What the export showsShare of revenueSource cell
Journeys with 1 touch (0.5 days to buy)79.5%Journeys sheet, 1 touch row
Journeys with 2 to 3 touches (12.5 days to buy)12.2%Journeys sheet, 2 to 3 touches row
Journeys with 4 to 9 touches (16.9 days to buy)5.4%Journeys sheet, 4 to 9 touches row
Journeys with 10 or more touches (16.0 days to buy)3.0%Journeys sheet, 10 or more touches row

That store runs on two clocks. Journeys with 1 touch carry 79.5% of revenue in the Journeys sheet, and they took 0.5 days to buy. A test of any length sees buyers that quick.

The second clock is slow. In the same Journeys sheet, journeys with 2 to 3 touches took 12.5 days and those with 4 to 9 touches took 16.9 days. Together, the multi-touch rows hold 20.6% of revenue in the export, about a fifth.

Now set those days against the platform floors. A 7-day test in that store would end before the average 2-to-3-touch journey had finished. A 14-day test would still stop short of the 16.9 days in that store's 4-to-9-touch row. A three-week test covers every row's average.

An average is a middle, not a maximum. A test sized to the average can still miss the slowest buyers, and a cooldown catches them.

The export cannot show how fast evidence piles up. It holds no order counts, so it can't say how quickly that store's test would fill. Your Shopify sales report can.

Why does the two-week rule mislead?

Two weeks is a fine floor and a poor fit.

It counts days, not evidence. A test is long enough when both groups hold enough sales to tell apart. Google lists study duration next to budget, holdback size and your choice of conversions as inputs to study power. Lewis and Rao studied 25 large field experiments with major U.S. retailers and brokerages. In the median one, the plausible range for return on investment was over 100 percentage points wide. Informative experiments, they wrote, can easily need more than 10 million person-weeks. Fewer people means more weeks.

It ignores your lag. Google's advice is to make a study long enough to capture your average conversion lag. That is the time between an impression and a conversion. A floor written for every advertiser can't know whether your buyers decide over lunch or over a month. Google also found up to a 17% drop in absolute lift when studies with a long conversion lag ran under 14 days.

It stops counting too early. Meta's Conversion Lift counts conversions relative to the test's start and end dates. A buyer who saw an ad on the last day and pays a week later falls outside. For geo tests, Google recommends a cooldown when your conversion cycle runs longer than a few weeks. The ads go back to normal, and the counting carries on.

It invites peeking. Results start to show while a study runs, and the early line wobbles. Google recommends waiting until the end of a study for the most accurate results. Meta suggests waiting until a test has finished. Stop the first time it looks good, and you are often stopping on noise.

What can a longer test not fix?

A test that is too small. Google says smaller holdouts need a longer study to collect enough data. It also says conversions that happen more often raise certainty. If even a long window leaves the power low, test something bigger: a whole channel, more regions or a larger holdout. Google allows holdouts of up to 50%, at the price of more people not seeing your ads.

A messy middle. A long test gives everyone more time to fiddle. During a study, Google advises sticking to business-as-usual changes such as bids or budgets. New creatives and audiences mid-study make the result hard to read. Lock the creative before you lock the dates.

The wrong season. A test that runs through a sale mostly measures the sale. If your next quiet stretch is shorter than your lag plus a cooldown, wait for a longer one. Or test a channel with quicker buyers first.

What happens months later. A test of a few weeks sees the sales during it and shortly after. It can't see whether those buyers come back next season. Read it as the short-run answer it is.

A longer test is a bigger net, not a better one. Size the net to the fish.

What to do this week

  1. Read the lag on the paths that pay. In GA4, click Advertising, then Key event attribution paths under the Key events dropdown. Sort the table by Purchase revenue and note Days to key event on the top multi-touch paths. Pass: those days fit inside the test you plan, plus a cooldown. Fail: they run past it, so lengthen the plan before anyone books dates.
  2. Measure how much your weeks wobble. In Shopify, go to Analytics > Reports, filter the Category by Sales and open Total sales over time. Group it by week across the last year. Pass: neighbouring weeks sit close together, so a few weeks of testing can show a modest lift. Fail: weeks swing hard, so plan a longer test or a bigger holdout.
  3. Ask Google for the power at your dates. In Google Ads, open the Lift studies tab under Campaigns > Experiments in the Campaigns menu. Select the plus button, choose Conversion Lift based on users, enter your dates and read the study power. Pass: it reaches 90%, Google's target for the most reliable results. Fail: it falls short, so extend the dates or raise the holdback first.

Check the homework. Your GA4 Attribution paths export already holds the evidence. Causality Engine reads that one file and shows what each channel caused next to what last-click gave it, in 1 to 2 minutes, for €99 once (excluding VAT), refundable within 30 days. Check the homework

Sources, 1 October 2026: Set up Conversion Lift based on users (Google Ads Help); Experiments FAQs (Google Ads Help); Best practices for A/B testing (Meta Business Help Center); Differences between Conversion Lift test results and other reporting tools (Meta Business Help Center); Understand your Conversion Lift based on geography measurement data (Google Ads Help); Similar performance between test and holdout groups in a test (Meta Business Help Center); The Unfavorable Economics of Measuring the Returns to Advertising (The Quarterly Journal of Economics, via RePEc); Key events attribution paths report (Google Analytics Help); Sales reports (Shopify Help Center).

Frequently asked questions

  • Can I stop an incrementality test early if the result looks clear?
    Usually not. Early results wobble while late buyers and noise arrive. Google recommends waiting until the end of a study for the most accurate results, and Meta suggests waiting until a test has finished. Set the end date before launch and read the result once.
  • Is a longer incrementality test always better?
    No. Google notes that every incrementality study has an opportunity cost, because the holdout misses your ads. Once the test covers your buying cycle and reaches the platform's power target, extra weeks mostly add exposure to sales, launches and seasons.
  • Does a smaller holdout group mean a longer test?
    Yes. Google says smaller holdouts need a longer study to collect enough data to tell the groups apart. Larger holdouts, up to 50%, fill up faster but cost more, because more people miss your ads while the test runs.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.