Skip to content

Can a small store run incrementality tests?

Usually yes, if it asks a big question. Switch one whole channel off in half your regions, compare total orders, and size the test in orders, not days. A small store can't see small effects, and the ad platforms' own lift studies come with minimums.

By , Founder & CEOUpdated 7 min read

Run the numbers for your store: the free holdout test planner.

Usually yes, if you test something big. A small store can switch one whole channel off in half its regions and compare total orders with the other half. What it usually can't do is see small effects, or qualify for Meta's and Google's own lift studies. Test the channel you doubt most, and size the test in orders.

The usual answer is a polite no: come back when you're bigger. It mixes up two problems. The platforms' own lift studies have an entry ticket, and a test you run yourself has none. Size mostly buys the power to see small effects.

What one store's data shows

One store's anonymised GA4 export, 1 January 2024 to 21 August 2026. It holds shares of revenue only: no ad spend, no order counts.

What the export showsShare of revenueSource cell
Journeys with 1 touch: 14 distinct paths, 0.5 days to buy79.5%Journeys sheet, 1 touch row
Journeys with 10 or more touches: 1,792 distinct paths, 16.0 days to buy3.0%Journeys sheet, 10 or more touches row
All channels in the touched view, added up110.4%Channels sheet, Touched column total

Start with the trap. The Journeys sheet lists 3,670 distinct path sequences, which looks like a busy store. It isn't a count you can test with: one path can stand for one purchase or hundreds. The export holds no order counts, so it can't say whether this store is small, or size a test.

What the rows can do is aim a test. In that store's Journeys sheet, 14 one-touch paths carry 79.5% of revenue, and those buyers took 0.5 days to buy. That store's 1,792 paths with 10 or more touches carry 3.0%. The money sits in a few short journeys, not the long tail.

If your export looks like that, a channel on quick, one-visit journeys shows its effect within days of going dark. A channel that mostly turns up in long journeys plays for a thin slice of revenue. Even a real effect there drowns in weekly wobble.

The touched view in that store's Channels sheet adds up to 110.4%, because a journey that met two channels counts in both. Ad platforms count the same way. So treat a platform's claimed share of your orders as a ceiling for the effect, and size the test for less.

The export can't show that store's order count, its weekly wobble or what any channel caused. Those take Shopify and a test.

Why does the usual answer mislead?

It treats the platform's door as the whole room. The doors are real:

Those are each tool's rules, not the rules of statistics. Nobody needs a ticket to switch ads off in half their regions: Google's own geo guide calls that a go-dark test.

What size buys is resolution. Economists Lewis and Rao examined 25 large ad experiments run with big American retailers and brokerages, most reaching millions of customers. They concluded that informative ad experiments can be so costly that they may be out of reach for many firms. Their trouble was seeing small effects against very noisy sales. A small store's question is usually blunter: does this channel pay at all?

Here is how resolution scales, with an even split and only chance in the way. Order counts wobble by about their square root. For illustration, a half that should sell 100 orders will often sell 90 or 110 by luck alone. The gap between two halves wobbles by the square root of their orders added together. A gap you can trust is about three times that.

For illustration: the effect you want to seeOrders the half with ads needs, at least
40% of its salesabout 100
20% of its salesabout 400
10% of its salesabout 1,700
5% of its salesabout 7,000

Halve the effect and you need about four times the orders. Real weeks wobble more than pure chance, so treat these as floors.

For illustration, a store selling 50 orders a week can split its regions into halves of about 25 a week each. Say the test runs four weeks: the half with ads collects about 100 orders. If the channel causes 40% of sales, that is enough to see it; if it causes 10%, it is nowhere near. A small store can run a test. It just has to ask a big question.

What can a small test not tell you?

  • That a channel does nothing. A flat result means the gap was smaller than the test could see. Meta reads its own lift tests the same way: similar groups mean the difference was not conclusive or significant. If your test could only see 40%, read a flat result as "less than 40%" and decide on that.
  • Whether add-to-carts are sales. Google suggests upper-funnel actions, such as page views, as secondary results when purchases don't have enough data. A lift in add-to-carts says people responded. It doesn't say they paid.
  • A clean answer, if you flick switches. Meta advises against turning ad sets on and off by hand as a test. Split by region, so both halves share the same weeks.

What to do this week

  1. Count your test budget. In Shopify, open Analytics > Reports, filter the Category by Sales, then open Total sales over time by week. Halve your usual weekly orders and multiply by the weeks you could test. Pass: if you get 400 or more for the half with ads, effects near 20% are within reach. Fail: if you get under 100, only a channel causing 40% or more of sales will show.
  2. Read the claim as a ceiling. In Meta Ads Manager's main table, compare last month's Results for purchase campaigns with Shopify's orders. Pass: the claim is a big share of your orders, so even half of it could show. Fail: it is a small share, so a small test won't see it; judge that channel on margin.
  3. Check whether a split would starve your ads. In Ads Manager, read the Delivery column for the ad sets you would test. Pass: they have left learning, which Meta says usually happens after about 50 results in the week after the ad set's last significant edit. Fail: they read Learning limited, so cutting their regions in half can starve them further. Combine similar ad sets first, as Meta suggests, and allow more weeks.

Check the homework. Your GA4 Attribution paths export already holds the evidence. Causality Engine reads that one file and shows what each channel caused next to what last-click gave it, in 1 to 2 minutes, for €99 once (excluding VAT), refundable within 30 days. Check the homework

Sources, 1 October 2026: Set up Conversion Lift based on users (Google Ads Help); About Conversion Lift (Meta Business Help Center); View and understand holdout test results across Meta technologies (Meta Business Help Center); Implement campaigns for geo experiments (Google Ads Help); The Unfavorable Economics of Measuring the Returns to Advertising (The Quarterly Journal of Economics, via RePEc); Similar performance between test and holdout groups in a test (Meta Business Help Center); About A/B testing (Meta Business Help Center); Sales reports (Shopify Help Center); Tools to create A/B tests on Meta technologies (Meta Business Help Center); About the learning phase (Meta Business Help Center)

Frequently asked questions

  • How many orders do I need for an incrementality test?
    It depends on the effect you want to see. For illustration, with an even split, seeing a 40% effect takes about 100 orders in the half with ads. If you want to see 20%, it takes about 400, and 10% takes about 1,700. Real stores wobble more than chance, so plan for extra.
  • Does a 50/50 holdout cost me half my sales?
    No. It costs whatever the channel was adding in the dark half, for the weeks it stays dark. If the channel was adding nothing, the test costs nothing. That makes the channel you doubt most the cheapest one to test first.
  • Can I measure lift on add-to-carts instead of purchases?
    As a second signal, yes. Google suggests tracking upper-funnel actions as secondary results when purchases don't have enough data, since they show higher lift. An add-to-cart lift tells you people responded. It doesn't tell you they paid, so set budgets on orders.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.