How accurate is campaign attribution? What experiments show
Less accurate than a dashboard implies: in published Facebook experiments, observational methods often missed the randomized result. A paired-region holdout on one campaign shows how far off your own number is.
By Joris van Huët, Founder & CEOPublished 6 min read
Campaign attribution is an estimate, and the experiments that checked estimates like it found large gaps. In 15 advertising experiments at Facebook, observational methods "often fail to produce the same effects as the randomized experiments" (Gordon and colleagues, Marketing Science, 2019). In 663 later experiments, one statistical method's median lift for lower-funnel outcomes such as purchases was 24%, where the experiments measured 5% (Gordon, Moakler and Zettelmeyer, arXiv preprint, 2022). Neither paper tested a named attribution product, so the check that settles your numbers is a holdout on your own campaign.
What did the experiments compare?
Both papers use the same design. A randomized experiment holds some users back from a campaign, which gives the true lift. The authors then set the holdout aside, estimate the lift from observational data alone, and compare the two. Rules such as last click are accounting, not effect estimates, so the comparison matters when someone reads a credit as what the ad caused.
The 2019 paper did this with "multiple observational models" and "extensive demographic and behavioral variables". The 2022 preprint used double/debiased machine learning (DML) and stratified propensity score matching (SPSM) on "over 5,000 user-level features", data it calls "richer than what most advertisers or their measurement partners can access". It concluded that "neither method performs well".
How big were the errors, and which way did they run?
In the 2022 preprint, the experiments' median lifts were 29%, 18% and 5% for upper, middle and lower funnel outcomes: page views, add-to-cart and purchase, in the paper's terms. Using DML, the median lifts were 83%, 58% and 24%. With SPSM they were 173%, 176% and 64%. The authors call that "significant relative measurement errors", say DML is "less upwardly biased than SPSM", and describe "the remaining bias" as "substantial". The methods did "comparatively better for prospecting campaigns rather than for remarketing campaigns".
In field experiments at eBay (Blake, Nosko and Tadelis, Econometrica, 2015), "returns from paid search are a fraction of" the non-experimental estimates, and brand keyword ads had "no measurable short-term benefits". In both the preprint and the eBay paper, the observational number was the larger one. The 2019 abstract does not say which way its misses went, so read this as a pattern in two sources, not a law.
Can an experiment settle it for a small advertiser?
Not always. In 25 field experiments with major U.S. retailers and brokerages, "the median confidence interval on return on investment is over 100 percentage points wide" (Lewis and Rao, Quarterly Journal of Economics, 2015). Individual sales are so volatile that "informative advertising experiments can easily require more than 10 million person-weeks", which the authors say makes experiments costly and possibly out of reach for many firms. The same abstract calls selection bias "a crippling concern for widely employed observational methods".
So both sides are hard: observational estimates can be biased, and experiments can be too noisy to read. Fewer orders make a test noisier, not quieter. A small test that shows no gap does not show the campaign does nothing. It shows the test could not see one.
Google's data-driven model is trained and reported by Google. Its help page says the model compares people who saw an ad with "similar users in a holdback group", and it prints no accuracy measurement. Treat that as a claim to test.
How do you check your own campaign?
Run a paired-region holdout on one campaign:
- Decide what each answer changes. Write down what you would do if the campaign's true effect were half of what your tool credits, and what you would do if it matched.
- Pair your regions. Export Shopify orders by shipping region. Rank regions by orders, pair neighbours in the ranking, and flip a coin inside each pair to pick the region that pauses the campaign. Keep the running regions' budget where it was.
- Run it long enough. Google's lift-study guidance sets the duration to "a length that captures your average conversion lag". GA4's attribution paths report shows days to key event. Do not stop early because the gap looks good.
- Measure the gap. For each pair, take orders in the paused region minus orders in its running partner during the test, then subtract the same difference from the weeks before. Add up the pairs.
- Put it next to the noise. Work out the same pair difference for each earlier week. If the test's gap is no bigger than the usual weekly swing, the test could not see the effect.
Pass: the drop in orders is about what the tool credited to the campaign in those regions, and it is larger than the usual swing. Fail: the drop is far smaller than the credited orders, so the tool claims orders the campaign did not cause; or there is no gap beyond the swing, so the test needs a longer run or a bigger change.
For illustration: if a tool credits a campaign with 400 orders over the test weeks and half your orders sit in the paused regions, a correct attribution predicts a loss of about 200 orders there, and a measured loss of 60 would mean it credited about 3.3 times what the campaign caused (200 divided by 60).
Later, once one holdout has shown how far a dashboard can drift, a causal attribution read like Causality Engine's can give you a next step per channel, and how to test it, from a GA4 export.
Sources, 30 September 2026: A Comparison of Approaches to Advertising Measurement (Gordon and colleagues, Marketing Science, 2019; abstract via IDEAS/RePEc); Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement (Gordon, Moakler and Zettelmeyer, arXiv preprint, 2022; full text); The Unfavorable Economics of Measuring the Returns to Advertising (Lewis and Rao, Quarterly Journal of Economics, 2015; abstract via IDEAS/RePEc); Consumer Heterogeneity and Paid Search Effectiveness (Blake, Nosko and Tadelis, Econometrica, 2015; abstract via IDEAS/RePEc); Get started with attribution (Google Analytics Help, 2026); Set up Conversion Lift based on users (Google Ads Help, 2026).
Related answers
Frequently asked questions
How accurate is campaign attribution?
Less accurate than dashboards imply, and the miss can be measured. In a 2022 preprint of 663 Facebook experiments, one method's median lift for lower-funnel outcomes such as purchases was 24% where the experiments measured 5%. Those were statistical estimators, not attribution products, so check your own campaign with a holdout.Can a small brand test whether its attribution is right?
Yes, with a paired-region holdout on one campaign, but expect noise. Across 25 published retail and brokerage experiments, the median confidence interval on return on investment was over 100 percentage points wide. Compare the gap with your normal weekly swing, and read a small gap as unclear, not as zero.Does data-driven attribution fix the accuracy problem?
Not by itself. Google's help page says its model compares people who saw an ad with similar users in a holdback group, and it prints no accuracy measurement. Google trains and reports the model, so test the channel ranking it gives you against a holdout before moving budget.
Go deeper: Incrementality testing, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Causal AttributionCausal Attribution uses causal inference to determine which marketing touchpoints genuinely cause conversions, not just correlate with them.
- Confidence IntervalConfidence Interval is a statistical range of values that likely contains the true value of a metric. In marketing analytics, it quantifies uncertainty around estimates, indicating the precision of an outcome or causal effect.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Machine LearningMachine Learning involves computer algorithms that improve automatically through experience and data. It applies to tasks like customer segmentation and churn prediction.
- Propensity ScoreA propensity score is the probability a unit receives a specific treatment given observed characteristics. It reduces selection bias in observational studies, enabling causal inference when randomized experiments are not possible.
- Propensity Score MatchingPropensity Score Matching is a statistical method that estimates the causal effect of a treatment from observational data. It matches individuals with similar likelihoods of receiving treatment to isolate its impact.