How to run a matched market test, step by step
Pull daily sales by region, group look-alike markets and let chance pick the test side. Check the fit on weeks you held back, then add budget in the test markets only. Predict each test day from the controls and add up the gaps.
By Joris van Huët, Founder & CEOUpdated 8 min read
Usually in seven steps. Pull daily sales by region, group look-alike markets, and let chance pick which ones get the change. Check the match on weeks you did not use to build it. Then make the change, keep everything else still, and compare real sales with what the control markets predicted.
This walkthrough adds budget in the test markets, which Google calls a heavy-up. Nobody loses ads while it runs, and Google's GeoX docs suggest it when current spend is too low to power a go-dark test. The same steps work for switching a channel off. You need Shopify reports, a spreadsheet and your ad account's location settings.
Step by step
- Draw markets people do not commute across. Use regions or groups of cities that you can target in your ad account and find in Shopify. Google's GeoX glossary says boundaries should not cross well-known commuting patterns, so daily travel does not spill ads into the control side. Path: Shopify admin > Analytics > Reports > Total sales by billing location.
- Export daily gross sales. Take three times the test length at least, or a year if sales follow the seasons, as GeoX advises. Use gross sales, because Shopify books a return as a negative number and GeoX advises against metrics that can go negative. Path: Analytics > Reports > Total sales over time > Dimensions menu > day, then ⊕ Billing city > Export.
- Group look-alike markets, then let chance split each group. Sort the markets into a few groups with a similar sales level, trend and seasonality. Inside each group, a coin decides which markets get the extra budget, the way GeoX assigns markets at random within each cluster. Path: Google Sheets, one row per market, with its group and the coin's verdict.
- Fit on old weeks, score on recent ones. Average the test markets' daily sales and the control markets' daily sales. Fit test against control on all but the latest eight weeks, then check those eight with RSQ, a simple stand-in for GeoX's out-of-sample fit. GeoX recommends an R-squared of at least 0.8 and calls under 0.5 unreliable. Path: Google Sheets, =SLOPE(data_y, data_x), =INTERCEPT(data_y, data_x) and =RSQ(data_y, data_x).
- Run a fake test first. Pretend the test began four weeks ago and predict those weeks from a line fitted on the weeks before. Add up the daily gaps: with nothing changed, they should land near zero, the kind of A/A check GeoX runs on every design. Path: the same sheet, with the fit ending where the fake test begins.
- Add budget in test markets only. In Google Ads, duplicate the campaign: one copy for control markets, one for test markets, both on Presence. Keep the control copy at the control markets' usual spend and raise the test copy's budget. Keep the 4 to 5 day learning period inside the test, as Google advises. Path: Google Ads > Campaigns menu > Campaigns > Settings icon > Locations > Location options; on Meta, Ad set > Audience > Locations.
- Hold still, then add up the gap. Change nothing else in either group, and write down anything you must change. After the test, predict each day from the control markets, subtract, and add up the gaps, cooldown included. Path: Google Ads > Campaigns > Insights & reports > When and where your ads showed > Matched locations, to confirm where the budget went.
That sum is GeoX's time-based regression in miniature: a line fitted before the test predicts each test day, and the gaps add up. Divide the extra sales by the extra spend and you have incremental ROAS.
A worked example
For illustration, with round invented numbers. Say you sell in 20 regions and want to know whether more budget for your non-brand Search campaign would pay. Say you sort them into five groups of four by sales level and trend. If a coin sends two of each group to test, 10 regions get more budget and 10 stay as they are.
If the test runs 30 days, the three-times rule asks for at least 90 days of daily history, and a year is better. Say your fit on the older weeks puts the test regions €20 a day above the control regions, with a slope of 1.0. If RSQ on the eight held-back weeks comes out at 0.86, the match clears GeoX's 0.8 bar. Say a fake test over the last four weeks sums to €300, close to nothing next to a month of sales in 10 regions.
Say the test copy of the campaign then gets an extra €38 a day in each test region.
| For illustration | Per test region, per day | All 10 test regions |
|---|---|---|
| Control regions' average sales during the test | €480 | |
| Predicted test sales (€20 + 1.0 × €480) | €500 | |
| Actual test sales | €550 | |
| Extra sales over 30 test days (€50 × 10 × 30) | €50 | €15,000 |
| Extra sales over a 14-day cooldown (€15 × 10 × 14) | €15 | €2,100 |
| Extra spend over 30 days (€38 × 10 × 30) | €38 | €11,400 |
| Incremental ROAS (€17,100 / €11,400) | 1.5x |
Why two weeks of cooldown? In one store's Journeys sheet, journeys of 2 to 3 touches took 12.5 days to buy. If your slow buyers take about that long, a two-week cooldown gives them time to buy.
Say your gross margin is 50%: break-even ROAS is then 1 / 0.50 = 2.0x. If the numbers hold, each extra euro brings back €1.50 of sales and €0.75 of margin: a loss of 25 cents. Keep the budget where it was, or test a smaller step up.
What to check when the result looks wrong
- The fake test showed a lift. The match is loose, or something hit one group. Regroup the markets or drop the odd one out, because GeoX rejects designs that fail its A/A check.
- One control market jumped on its own. A regional sale, a store opening or local press can do that. Meta's GeoLift guide asks you to keep local marketing constant across test and control, so drop that market and refit.
- Budget showed up in control markets. Check Matched locations and the Presence setting first. On Meta, untick Reach more people likely to respond to your ads, which Meta may tick by default for a city or region.
- The gap is real but small. Compare it with the smallest lift your design could detect, which GeoX calls the minimum detectable effect. A gap below it cannot be told apart from noise.
- Refunds turned some days negative. Run the test on gross sales, then apply your usual net-to-gross ratio afterwards, as GeoX suggests.
What to do this week
- Export a year of daily sales by city. Open Total sales over time, set the time unit to day, add Billing city and export. Pass: few empty days per city. Fail: GeoX warns when more than 30% of days have no data, so merge thin cities into regions.
- Read what each market costs today. In Google Ads, open Insights & reports, then When and where your ads showed, then Matched locations. Set the date range to last quarter. Pass: you can read spend for each candidate market. Fail: much of it sits at country level only, so test with bigger markets.
- Score one candidate split. In Google Sheets, put the two groups' daily averages side by side and run RSQ on the latest eight weeks. Pass: 0.8 or more, the floor GeoX recommends. Fail: under 0.5, which GeoX calls unreliable, so regroup the markets.
Check the homework. Your GA4 Attribution paths export already holds the evidence. Causality Engine reads that one file and shows what each channel caused next to what last-click gave it, in 1 to 2 minutes, for €99 once (excluding VAT), refundable within 30 days. Check the homework
Sources, 1 October 2026: Implement campaigns for geo experiments (Google Ads Help); Meridian GeoX: Types of experiments (Google for Developers); Meridian GeoX: Glossary (Google for Developers); Meridian GeoX: Prepare your pretest data (Google for Developers); Meridian GeoX: Design methodology (Google for Developers); Meridian GeoX: FAQs (Google for Developers); Meridian GeoX: Time-based regression (Google for Developers); Meridian GeoX: Data validation and quality checks (Google for Developers); SLOPE (Google Docs Editors Help); INTERCEPT (Google Docs Editors Help); RSQ (Google Docs Editors Help); Sales reports (Shopify Help Center); Setting and comparing time ranges for your reports (Shopify Help Center); Exporting reports (Shopify Help Center); Filtering and editing your reports (Shopify Help Center); Use location targeting (Meta Business Help Center); View matched locations and distance reports (Google Ads Help); GeoLift Best Practices (Meta)
Related answers
Frequently asked questions
Can I run a matched market test in a spreadsheet?
Yes, for a simple design. Fit the test markets to the controls with SLOPE and INTERCEPT, check the fit with RSQ on held-back weeks, then add up the gaps. Google's GeoX and Meta's GeoLift are code libraries that do the same with more safety checks.Should I use daily or weekly sales for a matched market test?
Daily. Google's GeoX accepts only daily data, and Meta's GeoLift guide strongly recommends daily over weekly. Weekly totals leave too few points to judge a match in a short test. If daily sales are too spiky, merge small markets into bigger ones or run the test longer.Can I test two channels in one matched market test?
Yes, with a separate group of test markets for each change and one shared control. Google's GeoX calls this a multi-cell design and warns that it needs significantly more markets and budget. With few regions, test one change at a time.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- AnalyticsAnalytics is the systematic computational analysis of data. It reveals customer behavior and measures campaign performance.
- AttributionAttribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
- CausalityCausality is the relationship where one event directly causes another, essential for identifying specific actions that drive desired outcomes in marketing.
- ExperimentsExperiments are scientific procedures that test hypotheses or demonstrate facts. In marketing, experiments like A/B tests determine the causal effect of campaign changes, enabling data-driven decisions.
- Google AdsGoogle Ads is an online advertising platform where advertisers bid to display ads, service offerings, and product listings.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.