Post-purchase surveys vs modelled attribution
A survey records what buyers remember and pick from your list, and the list changes the answer. Models fitted without a holdout miss experimental results too. Use each to check the other, then let a holdout decide budget.
By Joris van Huët, Founder & CEOPublished 5 min read
A post-purchase survey records what buyers remember and choose from your list, not what caused the order, and the list changes the answer. When Pew Research Center asked the same question two ways after the 2008 US presidential election, 58% chose the economy as the issue that mattered most when it was on the list, against 35% who volunteered it when the question was open. Modelled attribution has its own error, so use each to check the other and let a holdout settle budget.
What does a survey answer actually measure?
Memory first. Survey research has long found that questions about events in the past strain recall, and that respondents then fill the gaps with inference from partial memory (Bradburn, Rips and Shevell, Science, 1987). Then your answer list shapes what comes out. In the same Pew experiment, fewer than one-in-ten people (8%) gave an answer outside the five options they were read, against 43% when the question was open. That poll was about voting, not shopping, but the mechanism carries: the list you show decides which answers exist. A buyer who found you through a newsletter or a podcast you didn't list picks the nearest option, and you don't see the gap.
Order matters too. Pew says research suggests that in self-administered surveys, which is what a checkout survey is, people tend to choose items at the top of the list, a primacy effect. Randomising the order spreads that bias without removing it.
What a survey is good for is channels that leave no click of their own: a podcast, a friend, a creator's video. Click tracking has nothing to count there.
How wrong can modelled attribution be?
One large published test of modelling without a holdout used Facebook's own logged data. In 663 large-scale experiments, Gordon, Moakler and Zettelmeyer compared randomised lifts with two non-experimental methods given access to over 5,000 user-level features. For upper-funnel outcomes the median randomised lift was 29%, against 83% from double/debiased machine learning and 173% from stratified propensity score matching. For lower-funnel outcomes it was 5%, against 24% and 64%. The authors conclude that, even with those data, they could not reliably estimate an ad campaign's causal effect (arXiv, 2022).
Earlier work on 15 US advertising experiments at Facebook, with 500 million user-experiment observations, found that observational methods often failed to produce the same effects as the randomised experiments, even after conditioning on extensive demographic and behavioural variables (Gordon and colleagues, Marketing Science, 2019).
Read those as a warning about models with no holdout, not a score for any tool. The newer paper tested two statistical methods and the older one several observational models, both on Facebook data. Neither tested GA4's attribution models, a survey or any vendor's product.
How do you use both?
Run it on your own orders:
- Fix the question. Ask once, after purchase. Keep a text box for Other, randomise the order of the options, and read the Other answers until you stop finding channels you forgot to list.
- Check who answers. Compare the buyers who answered with the full set of buyers over the same weeks on order value, new versus returning, and country. If they differ, say so whenever you quote a share.
- Match each answer to its order, then to that order's source in GA4 using the transaction ID you already send with the purchase event.
- Tabulate survey channel against GA4 channel. Where they agree, carry on. Where they disagree, you have found your test list: a survey that says podcast while GA4 says Direct points at something GA4 doesn't see.
- Test the biggest disagreement with a holdout, then read the survey again.
What the result means:
- Pass: orders and the channel's survey share both fall when the channel is off. The two signals agree with the experiment.
- Fail, credit without cause: the survey credited the channel, but orders didn't move. The survey was measuring memory.
- Fail, blind spot: orders fell, but the survey didn't name the channel. The survey was measuring your answer list.
Where does a causal read fit?
Later, once the survey and your dashboards disagree about a channel, a causal attribution read like Causality Engine's can give a third view from your GA4 export, at the level of GA4 channel groups. The holdout still decides.
Sources, 30 September 2026: Writing Survey Questions (Pew Research Center); Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement (Gordon, Moakler and Zettelmeyer, arXiv, 2022); A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook (Marketing Science, 2019); Answering autobiographical questions: the impact of memory and inference on surveys (Science, 1987); Dimensions and metrics (Google Analytics Help, 2026). Pew's figures come from one question asked in one poll, and the Facebook studies are peer-reviewed or preprint research on one platform; no source read here compares a post-purchase survey with a holdout.
Related answers
Frequently asked questions
Are post-purchase surveys accurate for attribution?
They record what buyers remember, which helps for channels that leave no click. The answer list shapes the result: in one Pew poll, 58% chose an option when it was offered and 35% volunteered it unprompted. Treat survey shares as leads to test, not as credit.Should my how-did-you-hear question have an Other option?
Yes, with a text box, and read what people write. In Pew's test, 43% of people asked openly gave an answer outside the five listed options, against 8% when a list was read. The Other answers show channels you forgot to list.Can modelled attribution replace a survey, or the reverse?
No. They fail differently: surveys reflect memory and the options offered, and models fitted without a holdout can sit far from experimental results. Use each to check the other, then test the biggest disagreement with a holdout.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Attribution ModelAn Attribution Model defines how credit for conversions is assigned to marketing touchpoints. It dictates how marketing channels receive credit for sales.
- Attribution ReportAttribution Report shows which touchpoints or channels receive credit for a conversion. It identifies which campaigns drive desired actions.
- Causal AttributionCausal Attribution uses causal inference to determine which marketing touchpoints genuinely cause conversions, not just correlate with them.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- Machine LearningMachine Learning involves computer algorithms that improve automatically through experience and data. It applies to tasks like customer segmentation and churn prediction.
- Propensity ScoreA propensity score is the probability a unit receives a specific treatment given observed characteristics. It reduces selection bias in observational studies, enabling causal inference when randomized experiments are not possible.
- Propensity Score MatchingPropensity Score Matching is a statistical method that estimates the causal effect of a treatment from observational data. It matches individuals with similar likelihoods of receiving treatment to isolate its impact.