Skip to content

For GA4 usersFrustrated with GA4 attribution? Upload your GA4 export, see causal insights in 5–10 minutes for €99 pay-per-use.

Causal Inference

5 min read

A Dutch Province Geo Test Cannot Reach p Below 0.05

One treated province against eleven placebos gives twelve possible orderings, so the best p-value the design can return is 1 in 12. That is above 0.05 before a euro is spent. The floor by country, and the three fixes.

Share
Quick Answer·5 min read

A Dutch Province Geo Test Cannot Reach p Below 0.05: One treated province against eleven placebos gives twelve possible orderings, so the best p-value the design can return is 1 in 12. That is above 0.05 before a euro is spent. The floor by country, and the three fixes.

Read the full article below for detailed insights and actionable strategies.

Customer journey

The customer journey last-click attribution misses

One conversion. Five touchpoints. Last-click credits the final touch with 100%.

Instagram
Day 1
Pinterest
Day 4
Google Shopping
Day 7
Purchase
Day 10

Last-click attribution

Google Shopping100%

Every other channel gets zero credit, even though they created the demand.

Causal inference

Instagram48%
Pinterest27%
Google25%

A geo test that treats one Dutch province and checks its result against the other eleven as placebos can tell twelve stories, the real one and eleven pretend ones. The most extreme the real result can be is best of twelve, and one in twelve is 0.083. That is above the 0.05 the world calls significance, and no budget, model or vendor changes it. If your Black Friday geo test is designed on provinces with one treated unit, its p-value was decided before launch.

This is the floor under the floor. The measurability arithmetic in Can a brand your size measure Black Friday lift at all? asks whether the test is powerful enough. This one has nothing to do with statistics and everything to do with counting, and it catches tests that cleared the first bar.

The floor by geography

With N units and one treated, permutation inference over placebo assignments yields N possible orderings, so the smallest attainable p-value is 1 divided by N. The Price of Being Found tabulates it:

GeographyRegionsBest p-value physically possible
Netherlands, provinces120.0833
Germany, Bundesländer160.0625
Netherlands, COROP regions400.0250
US, states500.0200
France, départements960.0104
US, DMAs2100.0048

The book notes the statistical literature reached the same point independently: Lei and Sudijono, in a synthetic-control methods preprint (arXiv:2401.07152v3, revised 19 April 2025, not peer-reviewed), observe that the placebo test has low resolution because its null distribution is built from only N reference estimates, which creates a barrier at common levels like 0.05 when N is small. It is a known property of the design, not a quirk of Dutch provinces.

Twelve that behave like 7.8

Dutch provinces are wildly unequal. Zuid-Holland holds 21.4% of the population, Zeeland 2.2%, and the top three are more than half the country. Weight the units by population and the effective number of regions is about 7.8, not 12. You have three regions that matter, a few that sort of matter, and a handful that barely register, and the placebo distribution is built from all of them as if they were equals.

The three fixes

  1. Treat more than one unit. Two treated provinces raise the number of placebo assignments to 66 and drop the floor to 0.015, at the cost of a smaller control pool.
  2. Go finer. COROP regions, of which there are 40, give a floor of 0.025 and a far more even population distribution. The book's verdict: COROP is the correct unit for Dutch geo testing and provinces are not, as a fact about arithmetic rather than about any vendor's software. Elsewhere, DMA rather than state, département rather than région.
  3. Use inference that is not permutation-based, and say so plainly. Model-based intervals are not bound by 1 over N. They carry their own assumptions, and the design statement on the slide has to name them.

What you must not do is run the province design anyway and report the number that comes out, because the number was fixed before the campaign started.

Why this matters more before Black Friday than at any other time

The peak is when a test's result is read most eagerly and questioned least. A December report that says "p = 0.083, not significant" on a province design will be read as "the channel does nothing", when the design could not have said anything else. The book's sixth question for any measurement vendor is the defence: what geographic unit will you test on, and what is the minimum attainable p-value at that granularity? Ask it in September. Ask your attribution vendor for a placebo test has the companion questions.

What to do this week

  • If you have to defend the number: count your units. Write 1 over N on the design document next to the MDE, and if it sits above the level you intend to report against, change the design before 2 October, not the interpretation in December.
  • If you own the budget: ask whoever is designing the test which unit of geography it uses. If the answer is provinces and the treated count is one, the test cannot clear 0.05 and you have just saved eight weeks.

The calendar has the dates. A causal read on the GA4 export is not bound by the placebo floor because it is not a permutation design; it states its interval and its assumptions instead, which is the third fix above.

As of 9 September 2026. The floor table, the population weighting and the three fixes are from The Price of Being Found (Edition 2.10), Chapter 16 and Appendix D.7; the province floor of 1 in 12 is arithmetic and the book rates it Established.

Get attribution insights in your inbox

One email per week. No spam. Unsubscribe anytime.

Key Terms in This Article

Related Articles

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Ready to see your real numbers?

Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.

Full refund if you don't see value.

Stay ahead of the attribution curve

Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.

Which one are you? Optional.

No spam. Unsubscribe anytime. We respect your data.

Frequently Asked Questions

Why can't a one-province geo test in the Netherlands be statistically significant?

Placebo inference compares the treated province against each of the other eleven as if it had been treated, giving twelve possible orderings. The smallest p-value that design can produce is 1 in 12, which is 0.083, above the 0.05 threshold. More budget or a better model does not change the count.

What geographic unit should I use for a geo test in the Netherlands?

COROP regions. There are 40 of them, which gives a minimum attainable p-value of 0.025 and a much more even population distribution than the 12 provinces, where the top three hold more than half the country and the effective number of regions is about 7.8.

What is the placebo floor in a geo lift test?

With N regions and one treated, permutation-based inference can return at best a p-value of 1 divided by N. The fixes are to treat more than one unit, use a finer geography, or use a non-permutation inference method and state its assumptions on the design document.

Related reports

Real reports on this topic.

Anonymised reports from the Attribution Report Library tagged with causal inference.

Browse all related reports

Find your wasted ad spend in 5–10 minutes.

Watch the model work on a sample store first, no signup. Then upload your last 40–90 days of GA4 sessions and get incremental ROAS with confidence intervals. No pixel, no SDK. €99 per read.

Prefer to talk it through? Book a 20-min call, or read how it works.

Last-click guesses.We run the math.

Causal attribution for ecommerce brands. Watch the model work on a sample store first, then upload your GA4 export and see which channels really drove revenue in 5–10 minutes. €99, pay-per-use. Pro at €299/mo when you want it continuous.

No signup for the demo. Book a 20-min call or compare plans.