How to run a small-store incrementality test, step by step
Count your orders in Shopify, size the test to a question they can answer, and see whether Google or Meta will run it for you. If not, switch one whole channel off in half your regions, let the ads settle, and compare total orders by half.
By Joris van Huët, Founder & CEOUpdated 9 min read
Run the numbers for your store: the free holdout test planner.
Usually in eight moves. Count your orders and pick a question big enough for them. Check whether Google or Meta will run the test for you. If not, switch one whole channel off in half your regions, let the ads settle and run whole weeks. Then compare total orders by half, and read a flat result as a ceiling.
This is the version for a store with dozens of orders a week, not thousands. You need Shopify's reports, the location settings in Meta or Google Ads, and a calculator. Every menu path below comes from the platform's own help pages.
Step by step
- Count the orders you can test with. Halve your usual weekly orders from the last three months. That is roughly what each half of an even split collects. Multiply by the weeks you could run. Path: Shopify admin > Analytics > Reports > Category filter > Sales > Total sales over time > Group by > Week.
- Pick a question your orders can answer. Halve the effect you want to see and you need about four times the orders. If the half with ads collects only about a hundred orders, only an effect near 40% of its sales will stand out. So ask one blunt question about a whole channel: does it cause anything like the share it claims? The holdout test calculator turns your daily orders into test days. Path: Meta Ads Manager > main table > Results column, set against Shopify's orders for the same dates.
- Ask the platforms first. A platform study splits people at random and does the sums for you, so use one if it will have you. Google shows a study power estimate before launch and says to aim for 90% certainty. Meta says some lift tests show spend minimums or recommended spend when you start creating them. Path: Google Ads > Campaigns menu > Campaigns > Experiments > Lift studies tab > plus button; in Meta, Experiments > Create test.
- Split by regions, not by on and off weeks. Meta does not recommend testing by turning ad sets or campaigns on and off by hand. Regions let both halves share the same weeks, paydays and weather. List your cities by past orders, pair cities that sell alike, and flip a coin for each pair. Path: Shopify admin > Analytics > Reports > Total sales over time > Filters > ⊕ > Billing city > is one of.
- Switch the channel off in the dark half. On Meta, edit each ad set and go to Audience, then Locations. Hover over your country and click... to exclude cities. If Reach more people likely to respond to your ads appears, uncheck it so Meta stays inside your locations. On Google Ads, target the half with ads only and set Location options to Presence. Path: Meta Ads Manager > ad set > Audience > Locations; Google Ads > Campaigns menu > Campaigns > Settings icon > Locations.
- Let the ads settle first. Meta treats any change to targeting as a significant edit, which sends the ad set back into learning. Make the switch, then start counting once delivery is stable again. Meta says ad sets usually leave learning after about 50 results in the week after the last significant edit. If yours read Learning limited and never leave, wait one full week and start anyway. Path: Ads Manager > Columns > Last significant edit, then the Delivery column.
- Run whole weeks and touch nothing. Stop when the half with ads has the orders you sized for. Then keep counting through a cooldown for slow buyers. Leave bids, budgets, creative and offers alone, and write down any change with its date. Path: Ads Manager > Columns > Last significant edit, checked each week.
- Read total orders by half and do the sums. Pull each half's orders for the weeks before the test and for the test itself. Use the before-test ratio to predict what the dark half would have sold with ads; the shortfall is what the channel added. Then add both halves' orders and take the square root: a gap under about three times that is too close to luck to act on. Path: two saved Shopify reports, one per half, each with its own Date range filter.
A worked example
For illustration, with round numbers made up for the example. Say a store sells about 60 orders a week, almost all in one country, with Meta as its only paid channel. Suppose Meta claims about 24 purchases a week, or 40% of orders, and the store spends €800 a week there.
Say the store's cities pair into two halves that each sold about 30 orders a week before the test. A coin sends one half dark for four weeks. Both halves then run two cooldown weeks with Meta back on.
| For illustration | Half with Meta on | Dark half |
|---|---|---|
| Orders a week before the test | 30 | 30 |
| Orders over 4 dark weeks and 2 cooldown weeks | 176 | 158 |
| Gap between the halves | 18 | |
| Chance wobble: square root of 176 + 158 | about 18 | |
| Generous ceiling: gap plus two wobbles | about 54 orders |
For illustration, a gap of 18 orders is about one wobble, so this test can't tell the ads from luck. A small test often can't show that ads work. It can still put a ceiling on them.
Say the store's orders average €50. For illustration, the ceiling of about 54 orders is then about €2,700 of sales. If the dark half normally gets half of Meta's €800 a week, four dark weeks held back about €1,600 of spend. For illustration, that is a return of about 1.7x, even at the ceiling. For illustration, Meta claimed about 12 purchases a week in that half before the test, so about 48 over the four dark weeks.
Now the yardstick, from one store's export. Its Break-even sheet shows that at a 40% margin, break-even ROAS is 2.5x, which is 1 divided by 0.40. If your margin is also 40%, even the generous ceiling of this test loses money. Real weeks wobble more than chance, so the true ceiling sits a little higher. For illustration, it would have to rise by about half again to reach 2.5x.
So the call is to cut Meta back hard or rework it, then test again, rather than trust its claim. Where did two cooldown weeks come from? One store's Journeys sheet shows journeys with 2 to 3 touches taking 12.5 days to buy. If your buyers look like that, two cooldown weeks cover that row's average.
What to check when the report looks wrong
- Orders still arrive from dark cities. Expect a little. Meta excludes people by current and home location, so someone who lives in a lit city can still see ads while visiting a dark one. Billing cities also include gifts and travellers. Judge the gap, not a perfect zero.
- The half with ads got worse in week one. The targeting change restarted learning. Start the count once delivery settles, as in step 6, and drop the unsettled days from both halves.
- Spend in the half with ads jumped. On a capped Google Ads budget, removing regions pushes unspent money into the rest. Google says this inflates the baseline. Cut the daily budget to that half's usual share before day one.
- One city went its own way. A local launch, event or outage moved it. Google's geo guide says to exclude regions with major launches during a test, so drop that pair from both halves.
- The result flips from week to week. Small counts swing. Meta recommends waiting until a test has finished before judging it, and the same goes for yours. Read once, at the end.
What to do this week
- See how your orders spread. In Shopify, open Total sales over time for the last three months. Add Billing city as a column from the Dimensions menu. Pass: no single city holds most of your orders, so two halves can match. Fail: one city dominates, so split by postal code instead, which Meta accepts in Locations.
- Dry-run the exclusion in Meta. Open one ad set, go to Audience, select Locations, hover over your country and click... to find Exclude cities. Pass: you can exclude a city, and Reach more people likely to respond to your ads is unchecked or absent. Fail: that box is checked, so plan to uncheck it on the day you switch.
- Check Location options in Google Ads. If you run Google Ads, open the campaign's Settings, expand Locations, then Location options. Pass: Presence is selected, the setting Google's geo guide uses to prevent location leakage. Fail: another option is selected, so change it a week before the test, not on day one.
Check the homework. Your GA4 Attribution paths export already holds the evidence. Causality Engine reads that one file and shows what each channel caused next to what last-click gave it, in 1 to 2 minutes, for €99 once (excluding VAT), refundable within 30 days. Check the homework
Sources, 1 October 2026: Sales reports (Shopify Help Center); Filtering and editing your reports (Shopify Help Center); Tools to create A/B tests on Meta technologies (Meta Business Help Center); About the learning phase (Meta Business Help Center); Set up Conversion Lift based on users (Google Ads Help); Create a brand survey test (Meta Business Help Center); Best practices to get started with Experiments (Meta Business Help Center); About A/B testing (Meta Business Help Center); Use location targeting (Meta Business Help Center); Implement campaigns for geo experiments (Google Ads Help); Significant edits and learning phase (Meta Business Help Center); View and understand holdout test results across Meta technologies (Meta Business Help Center)
Related answers
Frequently asked questions
Should a small store hold out 10% or half of its regions?
Usually half. For a given number of orders, an even split gives the clearest answer. Google allows holdouts of up to 50% in its own studies, at a higher opportunity cost. Meta's single-cell holdout controls are a tenth of the test group's size, which suits big accounts better.Can I test with on and off weeks instead of regions?
You can, but it is the weaker design. Meta advises against switching ad sets or campaigns on and off by hand. A pause of 7 days or longer also sends an ad set back into learning. If regions are impossible, use long blocks, let a coin pick their order and expect only big effects to show.What if one city brings in most of my orders?
Split inside it. Meta lets you target and exclude postal codes, and Google's geo guide works with city IDs or postal codes outside the US. Pair postal areas that sold alike, flip a coin per pair, and keep both halves inside the same big city.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- AttributionAttribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
- Control GroupControl Group is a segment of an audience intentionally not exposed to a marketing campaign, used to measure the campaign's true causal impact.
- ConversionConversion is a specific, desired action a user takes in response to a marketing message, such as a purchase or a sign-up.
- ExperimentsExperiments are scientific procedures that test hypotheses or demonstrate facts. In marketing, experiments like A/B tests determine the causal effect of campaign changes, enabling data-driven decisions.
- Google AdsGoogle Ads is an online advertising platform where advertisers bid to display ads, service offerings, and product listings.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.