A Dutch Province Geo Test Cannot Reach p Below 0.05: One treated province against eleven placebos gives twelve possible orderings, so the best p-value the design can return is 1 in 12. That is above 0.05 before a euro is spent. The floor by country, and the three fixes.
Read the full article below for detailed insights and actionable strategies.
Customer journey
The customer journey last-click attribution misses
One conversion. Five touchpoints. Last-click credits the final touch with 100%.
Last-click attribution
Every other channel gets zero credit, even though they created the demand.
Causal inference
A geo test that treats one Dutch province and checks its result against the other eleven as placebos can tell twelve stories, the real one and eleven pretend ones. The most extreme the real result can be is best of twelve, and one in twelve is 0.083. That is above the 0.05 the world calls significance, and no budget, model or vendor changes it. If your Black Friday geo test is designed on provinces with one treated unit, its p-value was decided before launch.
This is the floor under the floor. The measurability arithmetic in Can a brand your size measure Black Friday lift at all? asks whether the test is powerful enough. This one has nothing to do with statistics and everything to do with counting, and it catches tests that cleared the first bar.
The floor by geography
With N units and one treated, permutation inference over placebo assignments yields N possible orderings, so the smallest attainable p-value is 1 divided by N. The Price of Being Found tabulates it:
| Geography | Regions | Best p-value physically possible |
|---|---|---|
| Netherlands, provinces | 12 | 0.0833 |
| Germany, Bundesländer | 16 | 0.0625 |
| Netherlands, COROP regions | 40 | 0.0250 |
| US, states | 50 | 0.0200 |
| France, départements | 96 | 0.0104 |
| US, DMAs | 210 | 0.0048 |
The book notes the statistical literature reached the same point independently: Lei and Sudijono, in a synthetic-control methods preprint (arXiv:2401.07152v3, revised 19 April 2025, not peer-reviewed), observe that the placebo test has low resolution because its null distribution is built from only N reference estimates, which creates a barrier at common levels like 0.05 when N is small. It is a known property of the design, not a quirk of Dutch provinces.
Twelve that behave like 7.8
Dutch provinces are wildly unequal. Zuid-Holland holds 21.4% of the population, Zeeland 2.2%, and the top three are more than half the country. Weight the units by population and the effective number of regions is about 7.8, not 12. You have three regions that matter, a few that sort of matter, and a handful that barely register, and the placebo distribution is built from all of them as if they were equals.
The three fixes
- Treat more than one unit. Two treated provinces raise the number of placebo assignments to 66 and drop the floor to 0.015, at the cost of a smaller control pool.
- Go finer. COROP regions, of which there are 40, give a floor of 0.025 and a far more even population distribution. The book's verdict: COROP is the correct unit for Dutch geo testing and provinces are not, as a fact about arithmetic rather than about any vendor's software. Elsewhere, DMA rather than state, département rather than région.
- Use inference that is not permutation-based, and say so plainly. Model-based intervals are not bound by 1 over N. They carry their own assumptions, and the design statement on the slide has to name them.
What you must not do is run the province design anyway and report the number that comes out, because the number was fixed before the campaign started.
Why this matters more before Black Friday than at any other time
The peak is when a test's result is read most eagerly and questioned least. A December report that says "p = 0.083, not significant" on a province design will be read as "the channel does nothing", when the design could not have said anything else. The book's sixth question for any measurement vendor is the defence: what geographic unit will you test on, and what is the minimum attainable p-value at that granularity? Ask it in September. Ask your attribution vendor for a placebo test has the companion questions.
What to do this week
- If you have to defend the number: count your units. Write 1 over N on the design document next to the MDE, and if it sits above the level you intend to report against, change the design before 2 October, not the interpretation in December.
- If you own the budget: ask whoever is designing the test which unit of geography it uses. If the answer is provinces and the treated count is one, the test cannot clear 0.05 and you have just saved eight weeks.
The calendar has the dates. A causal read on the GA4 export is not bound by the placebo floor because it is not a permutation design; it states its interval and its assumptions instead, which is the third fix above.
As of 9 September 2026. The floor table, the population weighting and the three fixes are from The Price of Being Found (Edition 2.10), Chapter 16 and Appendix D.7; the province floor of 1 in 12 is arithmetic and the book rates it Established.
Get attribution insights in your inbox
One email per week. No spam. Unsubscribe anytime.
Key Terms in This Article
Attribution
Attribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
Black Friday
Black Friday is the day after Thanksgiving in the United States. It marks the start of the Christmas shopping season and is a major sales event for retailers.
Incrementality
Incrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
Related Articles
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Ready to see your real numbers?
Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.
Full refund if you don't see value.
Stay ahead of the attribution curve
Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.
No spam. Unsubscribe anytime. We respect your data.
Frequently Asked Questions
Why can't a one-province geo test in the Netherlands be statistically significant?
Placebo inference compares the treated province against each of the other eleven as if it had been treated, giving twelve possible orderings. The smallest p-value that design can produce is 1 in 12, which is 0.083, above the 0.05 threshold. More budget or a better model does not change the count.
What geographic unit should I use for a geo test in the Netherlands?
COROP regions. There are 40 of them, which gives a minimum attainable p-value of 0.025 and a much more even population distribution than the 12 provinces, where the top three hold more than half the country and the effective number of regions is about 7.8.
What is the placebo floor in a geo lift test?
With N regions and one treated, permutation-based inference can return at best a p-value of 1 divided by N. The fixes are to treat more than one unit, use a finer geography, or use a non-permutation inference method and state its assumptions on the design document.