How do I run a geo holdout test?
Pick comparable regions, switch the ads off in a random set and leave the rest running. Then compare both groups with how they moved before the test, run a full purchase cycle, and keep counting after the ads return.
By Joris van Huët, Founder & CEOUpdated 7 min read
Run the numbers for your store: the free holdout test planner.
Usually like this: split your market into comparable regions, switch the ads off in a random set of them and leave the rest running. Then compare sales in both groups with how they moved before the test. If a gap opens and holds after your slowest buyers have bought, that gap is what the ads added.
What one store's data shows
One store's anonymised GA4 export, 1 January 2024 to 21 August 2026. It holds shares of revenue only: no ad spend, no order counts.
| What the export shows | Share of revenue | Source cell |
|---|---|---|
| Direct, in last click, first click and touched views | 57.7% | Channels sheet, Direct row |
| Paid Social, in all three views | 0.0% | Channels sheet, Paid Social row |
| Journeys with 1 touch (0.5 days to buy) | 79.5% | Journeys sheet, 1 touch row |
| Journeys with two or three touches (12.5 days to buy) | 12.2% | Journeys sheet, two-to-three touches row |
Start with Direct. In that store's Channels sheet, Direct holds 57.7% of revenue in every view, so almost every journey that touches Direct also starts and ends there. Click-based attribution cannot say what sent those buyers: a podcast, a friend, an ad they saw and never clicked. A geo holdout does not need to know. It counts every sale in a region, Direct included, so the ads are judged on all revenue, not on the clicks they can claim.
Paid Social sits at 0.0% in all three views of that store's export. The file holds no spend, so it cannot say whether paid social ran at all. If you run it and your own GA4 shows the same zero, click credit has nothing more to tell you. A holdout is one of the few ways left to see what it does.
Then the timing. In the Journeys sheet, one-touch journeys carry 79.5% of revenue and take 0.5 days to buy. Journeys with two or three touches take 12.5 days in that store's export. Score a test on the day the ads come back and you miss part of that slower lane. Google's fix is a cooldown: campaigns return to business as usual while you keep counting sales.
What the export cannot show is what any channel caused. Shares of revenue are credit, not effect, and with no spend in the file there is no return to work out for that store. Settling cause is the job of a test.
Why does the usual answer mislead?
The usual answer is to turn the ads off in a few cities and watch sales. Right idea, four holes.
- Regions picked by hand. In Google's classic design, regions are assigned to test or control at random. Hand-picking the quiet cities bakes your own guess into the result. Google's researchers also name the traps: few regions, revenue that is lopsided across them, and sales that swing over time.
- A run that is too short. Meta's geo-testing guide says a test should cover at least one purchase cycle. It sets a minimum of 15 days with daily data, or 4 to 6 weeks with weekly data. Google's guide wants clean history of at least 3 times the test length before you start.
- The platform as referee. Scoring the test in the ad platform's own conversions lets the ads grade their own homework. Google's geo study counts unattributed conversions by region. Its cross-platform guidance asks for an independent first-party data source. For a store, that is sales by region in Shopify.
- Budget that leaks back. Remove regions from a capped campaign and the platform may spend the savings in your control regions. Google warns that this inflates the baseline, and says to cut the daily budget to the control regions' run rate.
If you sell in one small country, the classic design runs thin. Google's researchers built a time-based method for tests with few regions, such as smaller countries or a single pair of matched markets.
What can a geo test not tell you?
- Small effects. Geo tests are noisy. Google says its geo-based lift study typically needs a higher budget than the user-based one. Meta recommends people-based tests where possible, because they have more statistical power.
- Which ad did it. Google's own comparison says geo results won't be sliced by categories such as demographics. You learn what a channel or a group of campaigns added, not which creative earned it.
- What the next euro buys. A go-dark test values the spend you already run. Whether more budget pays is a separate test, which Google calls a heavy-up: test regions get extra budget.
- That a channel does nothing. No significant lift means the gap was too small to tell apart from noise. It is not a measured zero, so check the size of the test before you cut.
A geo test is a blunt, honest instrument. It will not flatter a channel, and it will not explain one either. Use it to settle the one question that moves the most money.
What to do this week
- Ask Google Ads whether it will run the test for you. Open the Goals menu, then Measurements > Lift measurement, and select the plus button. Pass: Conversion Lift appears under Based on Geo, and your campaigns target a single country. Fail: it is missing, which Google says happens on some accounts, so ask your Google representative or run a manual go-dark test.
- Count your usable regions in Shopify. Go to Analytics > Reports, open Total sales over time, and add Billing city from the Dimensions menu. Pass: 20 or more regions with sales in most weeks, the floor Meta's guide recommends. Fail: a few cities carry nearly everything, so merge regions or use a few-region design.
- Time your slow buyers in GA4. Click Advertising, open the Key events drop-down and choose Key event attribution paths. Read Days to key event. Pass: the multi-touch lag fits inside your test plus a cooldown. Fail: it runs longer, so plan a longer cooldown, as Google recommends when the conversion cycle is longer than a few weeks.
Check the homework. Your GA4 Attribution paths export already holds the evidence. Causality Engine reads that one file and shows what each channel caused next to what last-click gave it, in 1 to 2 minutes, for €99 once (excluding VAT), refundable within 30 days. Check the homework
Sources, 1 October 2026: Measuring Ad Effectiveness Using Geo Experiments (Google Research). Estimating Ad Effectiveness using Geo Experiments in a Time-Based Regression Framework (Google Research). Trimmed Match Design for Randomized Paired Geo Experiments (Google Research). Implement campaigns for geo experiments (Google Ads Help). Set up Conversion Lift based on geography (Google Ads Help). Understand your Conversion Lift based on geography measurement data (Google Ads Help). Comparing lift types (Google Ads Help). Welcome to GeoLift (Meta). GeoLift Best Practices (Meta). Key events attribution paths report (Google Analytics Help). Sales reports (Shopify Help Center). Filtering and editing your reports (Shopify Help Center).
Related answers
Frequently asked questions
How long should a geo holdout test run?
At least one full purchase cycle, and longer if your buyers are slow. Meta's GeoLift guide sets a minimum of 15 days with daily data, or 4 to 6 weeks with weekly data. Then add a cooldown after the ads return, so late buyers still count.Is a geo holdout better than a platform lift study?
Not better, different. Meta recommends people-based lift tests where possible, because they have more statistical power. A geo test needs no user-level tracking, can cover several channels in one study and counts sales the platforms never see. You pay for that with more noise and a bigger budget.Can I run a geo holdout if I only sell in one small country?
Often, yes. Few regions weaken the classic randomised design, so Google's researchers built a time-based method for tests with few regions, including smaller countries. Pick regions whose sales move together, keep a long clean history, and expect to detect only a large effect.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- AttributionAttribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
- ConversionConversion is a specific, desired action a user takes in response to a marketing message, such as a purchase or a sign-up.
- ExperimentsExperiments are scientific procedures that test hypotheses or demonstrate facts. In marketing, experiments like A/B tests determine the causal effect of campaign changes, enabling data-driven decisions.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Lift MeasurementLift Measurement: A method to determine the incremental impact of a marketing campaign by comparing exposed and control groups.