How Long Should an Incrementality Test Run? A Sample-Size Guide for Shopify Brands: Most incrementality tests fail before they start, because the brand running them never had enough orders to detect the lift they were looking for. Here is the arithmetic, a duration table by order volume, and what to do when the numbers say you cannot run the test.
Read the full article below for detailed insights and actionable strategies.
The attribution problem
One sale. Four channels. 400% credit claimed.
Reported revenue: €400 · Actual revenue: €100 · Gap: €300
How Long Should an Incrementality Test Run?
Run your test until each arm has accumulated roughly 16 ÷ (expected relative lift)² orders. To detect a 20% lift you need about 400 orders per arm; a 10% lift needs about 1,600; a 5% lift needs about 6,300. At a 50/50 split, divide by half your weekly order volume to get the number of weeks. Below that threshold the test cannot distinguish a real effect from noise.
That is the whole answer. The rest of this guide explains where the number comes from, gives you a duration table by order volume, and — more importantly — tells you what to do when the arithmetic says your brand is too small to run the test at all. That last part is where most Shopify brands actually live, and almost nobody writes about it honestly.
Why Most Incrementality Tests Are Dead on Arrival
Incrementality testing is the gold standard for measurement, and for good reason: it is the only method that directly observes what happens when you turn spend off. Every other approach infers it. So the advice "just run a holdout" has become the reflexive answer to any attribution question.
The problem is that a holdout test is a statistical instrument with a resolution limit, and nobody checks the limit before pointing it at their business. A brand doing 400 orders a month runs a two-week Meta holdout, sees the test group order 3% more than the control, and concludes Meta drove 3% incremental revenue. It did not conclude anything. With that sample size, a 3% observed difference is entirely consistent with Meta driving 0% — and equally consistent with it driving 25%. The test had no power to tell those stories apart.
This is not a subtle statistical footnote. It is the single most common way incrementality testing goes wrong, and it produces confidently wrong budget decisions in both directions: brands cutting channels that worked, and brands scaling channels that did not.
The three numbers that determine whether your test can work
Before you design anything, you need three inputs:
- Baseline conversion rate or order volume. How many orders happen in a given period without any change.
- Minimum detectable effect (MDE). The smallest lift you would actually change your budget over. Not the lift you hope for — the lift that would change a decision.
- Tolerance for error. Conventionally 95% confidence and 80% statistical power, meaning a 5% chance of seeing an effect that is not there and a 20% chance of missing one that is.
Almost everyone skips step 2. They design a test to "see if Meta works" rather than to "detect whether Meta drives at least 15% incremental revenue." Those are different tests with wildly different costs.
The 16-Over-Lift-Squared Rule
Here is the piece of arithmetic worth memorising. For a standard two-arm test at 95% confidence and 80% power, the required number of orders per arm is approximately:
Orders per arm ≈ 16 ÷ (relative lift)²
Where relative lift is expressed as a decimal — 0.10 for a 10% lift.
This is a simplification of the two-proportion sample size formula, and it is accurate to within about 3% for the conversion rates typical of ecommerce (1–5%). The elegance is that it does not depend on your conversion rate. Whether your store converts at 1.2% or 4%, the number of orders you need is the same; only the number of sessions required to generate them changes.
| Relative lift you want to detect | Orders needed per arm | Total orders needed |
|---|---|---|
| 5% | ~6,300 | ~12,600 |
| 10% | ~1,600 | ~3,200 |
| 15% | ~700 | ~1,400 |
| 20% | ~400 | ~800 |
| 30% | ~175 | ~350 |
| 50% | ~65 | ~130 |
Read that table carefully, because it contains the uncomfortable truth of this entire article: detecting small effects is quadratically expensive. Halving the effect you want to detect quadruples the data you need. A 5% lift costs sixteen times more evidence than a 20% lift.
Duration Table: Weeks to Significance by Order Volume
Convert orders into calendar time with a 50/50 split:
Weeks ≈ 32 ÷ (relative lift² × weekly orders)
| Weekly orders | Detect 5% | Detect 10% | Detect 20% | Detect 30% |
|---|---|---|---|---|
| 100 (~430/mo) | 128 weeks | 32 weeks | 8 weeks | 3.6 weeks |
| 250 (~1,080/mo) | 51 weeks | 13 weeks | 3.2 weeks | 2 weeks* |
| 500 (~2,150/mo) | 26 weeks | 6.4 weeks | 2 weeks* | 2 weeks* |
| 1,000 (~4,300/mo) | 13 weeks | 3.2 weeks | 2 weeks* | 2 weeks* |
| 2,500 (~10,750/mo) | 5.1 weeks | 2 weeks* | 2 weeks* | 2 weeks* |
* Never run shorter than two full weeks regardless of what the arithmetic says. You need to cover complete weekly cycles, or day-of-week composition differences between arms will contaminate the result.
Anything above roughly eight weeks should be treated as impractical, not merely slow. Over a long window your creative changes, competitors move, seasonality shifts, and the assumption that the only difference between arms is the treatment quietly stops being true.
The Test Feasibility Ladder
Rather than asking "can I run an incrementality test," ask which rung of this ladder your brand is standing on. Each rung has a different correct answer.
| Rung | Weekly orders | What you can detect in ≤8 weeks | Correct measurement strategy |
|---|---|---|---|
| 1. Below resolution | Under 100 | ~30%+ only | Do not run channel holdouts. Use structural reasoning and causal modelling on historical data. |
| 2. Coarse | 100–400 | ~15–20% | Test only whole channels you suspect are near-zero. Never test creative or audience variants. |
| 3. Practical | 400–1,500 | ~10% | Channel-level holdouts work. Geo tests become viable. |
| 4. Fine-grained | 1,500+ | ~5% | Campaign-level and always-on holdouts are realistic. |
The ladder matters because the advice ecosystem is written by and for Rung 4 brands, then consumed by Rung 1 and 2 brands who cannot act on it. If you are doing 300 orders a week and an agency proposes a two-week creative-level incrementality test, the proposal is not aggressive — it is arithmetically void.
A Worked Example: The €40,000 Question
Illustrative example. Figures are constructed to demonstrate the method, not drawn from a specific customer.
A Shopify skincare brand spends €40,000 a month on Meta and does 900 orders a month (about 210 a week) at €95 average order value. The founder wants to know whether Meta is worth it.
The correlational read. Meta's dashboard reports 640 conversions and 4.2 ROAS. Last-click in GA4 reports 310. The gap is the usual platform-versus-reality discrepancy, and neither number answers the question, because both count conversions that would have happened anyway.
The causal read they attempt. They run a two-week 50/50 geo holdout. Result: treatment cities ordered 6% more than control. They conclude Meta drives 6% incremental revenue — roughly €5,100 a month against €40,000 of spend — and prepare to cut Meta entirely.
Why that conclusion is unsupported. Two weeks at 210 orders/week is 420 total orders, 210 per arm. Consult the rule: 210 orders per arm supports an MDE of √(16 ÷ 210) ≈ 27.6%. The test could only reliably detect a lift of about 28% or larger. The observed 6% falls deep inside the noise band; the true incremental lift could plausibly be anywhere from roughly −15% to +28%. The test did not show Meta is weak. It showed nothing at all.
What the arithmetic actually recommends. To detect a 15% lift — the threshold at which this brand would genuinely change budget — they need ~700 orders per arm, or 1,400 total: about 7 weeks, not two. That is a real cost: seven weeks of suppressed spend in half their geography. Whether that cost is worth paying is a business decision, but at least it is now a decision made with the price tag visible.
The alternative. Seven weeks of deliberately withheld advertising is a genuine expense. The same brand can instead apply causal inference to the GA4 data it already has — using the natural variation in spend, timing, and geography already present in eighteen months of history — and get a channel-level incremental estimate without pausing anything. It is a modelled estimate rather than an experimental one, and it deserves slightly wider error bars, but it costs €0 in withheld revenue and arrives in minutes rather than quarters.
Common Mistakes
- Designing without an MDE. "Let's see what we find" is not a test design. Decide the lift threshold that would change your budget before you start.
- Peeking and stopping early. Checking results daily and stopping when p < 0.05 inflates your false positive rate dramatically. Fix the duration in advance and honour it. Log the pre-registration in a marketing experiment tracker.
- Unbalanced splits for "safety." A 90/10 holdout feels safer but is far weaker than 50/50. Effective sample size is governed by the smaller arm; a 10% holdout needs roughly 2.8× the total volume of a balanced test.
- Testing during anomalous periods. A holdout spanning Black Friday measures Black Friday, not your channel.
- Ignoring cross-arm contamination. If your control geography sees your ads through spillover, or your holdout audience still receives branded search and retargeting, your control is not clean and your measured lift is biased toward zero.
- Confusing a null result with a zero. "Not statistically significant" means "this test could not tell." It is not evidence the channel does nothing.
- Forgetting regression to the mean. Testing a channel right after an unusually bad month will flatter the result. Regression to the mean is a persistent contaminant in short tests.
Pre-Flight Checklist
Before launching any incrementality test:
- I have written down the MDE — the lift that would change my decision.
- I have computed required orders per arm: 16 ÷ (MDE)².
- I have divided by half my weekly orders and the answer is 8 weeks or fewer.
- The window contains no promotions, launches, or seasonal peaks.
- Split is as close to 50/50 as the business will tolerate.
- Arms cannot contaminate each other via audience overlap or geographic spillover.
- I have pre-registered the duration and stop rule, and I will not peek.
- I have calculated the revenue cost of the withheld arm and accepted it.
- I know what I will do if the result is null.
If any box is unchecked, you are not ready. If box three fails, you are on the wrong rung of the ladder and should be modelling rather than experimenting.
When You Cannot Run the Test
Most Shopify brands sit on Rungs 1 and 2, which means the honest answer to "should I run a holdout" is often no. That does not leave you with last-click. It leaves you with quasi-experimental and modelled approaches that extract causal signal from variation that already occurred:
- Difference-in-differences — exploit a change you already made, comparing before/after against an unaffected comparison group.
- Synthetic control — construct a weighted composite of untreated regions to serve as the counterfactual.
- Propensity score matching — compare exposed and unexposed customers with similar pre-exposure characteristics.
- Bayesian causal attribution — estimate incremental contribution across all channels simultaneously from historical spend and conversion data, with explicit uncertainty intervals rather than false precision.
These are weaker than a clean randomised experiment and stronger than any rules-based attribution model. For a brand that cannot afford to withhold spend for seven weeks, that trade is usually correct. The attribution maturity model is not a ladder you climb by buying a more expensive tool; it is one you climb by matching method to the evidence your order volume can actually support.
Key Takeaways
- Required orders per arm ≈ 16 ÷ (relative lift)². Memorise it.
- Detection cost scales quadratically: halving the MDE quadruples the data required.
- Most incrementality tests reported by small DTC brands are underpowered and their conclusions are unsupported in both directions.
- Never run shorter than two weeks or longer than eight.
- A null result means the test could not tell. It is not evidence of zero effect.
- Below ~400 weekly orders, channel holdouts mostly cannot resolve decision-relevant effects. Model instead of experimenting — that is a legitimate strategy, not a consolation prize.
Incrementality testing earned its reputation honestly. But a method that requires more evidence than your business generates is not rigorous when you use it anyway — it is just expensive guessing with a scientific vocabulary. The discipline is knowing which instrument your data can support, and reaching for causal inference on historical data when the experimental one is out of range. Brands running attribution without a data team and those under €50K monthly spend face this constraint most acutely, and it is worth understanding before you commit a quarter to a test that was never going to answer your question.
For €99, upload any historical GA4 period and get causal attribution for every channel in 5–10 minutes — no pixel, no migration. Go Pro at €299/mo for continuous attribution, an AI chatbot for your data, and a developer API.
Get attribution insights in your inbox
One email per week. No spam. Unsubscribe anytime.
Key Terms in This Article
Attribution Model
An Attribution Model defines how credit for conversions is assigned to marketing touchpoints. It dictates how marketing channels receive credit for sales.
Causal Attribution
Causal Attribution uses causal inference to determine which marketing touchpoints genuinely cause conversions, not just correlate with them.
Causal Inference
Causal Inference determines the independent, actual effect of a phenomenon within a system, identifying true cause-and-effect relationships.
Incrementality Testing
Incrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
Propensity Score
A propensity score is the probability a unit receives a specific treatment given observed characteristics. It reduces selection bias in observational studies, enabling causal inference when randomized experiments are not possible.
Propensity Score Matching
Propensity Score Matching is a statistical method that estimates the causal effect of a treatment from observational data. It matches individuals with similar likelihoods of receiving treatment to isolate its impact.
Quasi-Experiment
A quasi-experiment estimates the causal impact of an intervention without random assignment. It applies when random assignment is not feasible or ethical.
Regression to the Mean
Regression to the Mean describes the phenomenon where an extreme variable measurement tends to be closer to the average on subsequent measurements. This can bias before-and-after studies, falsely attributing change to an intervention.
Related Articles
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Ready to see your real numbers?
Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.
Full refund if you don't see value.
Stay ahead of the attribution curve
Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.
No spam. Unsubscribe anytime. We respect your data.
Frequently Asked Questions
How long should an incrementality test run?
Until each arm accumulates roughly 16 ÷ (expected relative lift)² orders. Detecting a 20% lift needs about 400 orders per arm, a 10% lift about 1,600, and a 5% lift about 6,300. Divide by half your weekly order volume to get weeks. Never run shorter than two weeks or longer than eight.
What is a minimum detectable effect (MDE) in marketing?
The MDE is the smallest lift your test can reliably distinguish from zero. It is set before the test by your sample size, confidence level, and power. Critically, it should be chosen as the lift that would actually change your budget decision — not the lift you hope to see.
Why do small Shopify brands struggle to run incrementality tests?
Because detection cost scales quadratically with precision. A brand doing 100 orders a week would need 128 weeks to detect a 5% lift and 32 weeks to detect a 10% lift. Only effects of roughly 30% or larger are visible within a practical window, so most channel-level tests at that volume cannot resolve decision-relevant differences.
Is a 50/50 split better than a 90/10 holdout?
Yes, statistically. Effective sample size is governed by the smaller arm, so a 90/10 holdout requires roughly 2.8 times the total order volume of a balanced test to reach the same power. A 90/10 split feels safer commercially but is substantially weaker as a measurement instrument.
What does it mean if my incrementality test result is not statistically significant?
It means the test could not tell the difference between the effect you observed and zero. It is not evidence that the channel has no effect. Underpowered tests routinely produce null results for channels that genuinely drive incremental revenue, and treating a null as a zero is one of the most expensive misreadings in marketing measurement.
Can I measure incrementality without running a holdout test?
Yes. Quasi-experimental and modelled methods — difference-in-differences, synthetic control, propensity score matching, and Bayesian causal attribution — extract causal estimates from variation that already exists in your historical data. They carry wider uncertainty than a clean randomised experiment but cost nothing in withheld revenue, which makes them the right choice for most brands below roughly 400 weekly orders.
Why should I not stop a test early when it reaches significance?
Repeatedly checking results and stopping at the first significant reading inflates the false positive rate well beyond the nominal 5%. Random fluctuation will cross the threshold at some point in almost any test. Fix the duration in advance, record it, and evaluate only at the end.
Does my store conversion rate change how many orders I need?
Barely. The 16-over-lift-squared rule depends on the relative lift you want to detect, not your conversion rate. A store converting at 1.2% and one converting at 4% need roughly the same number of orders per arm — the low-converting store simply needs far more sessions to produce them.