Designing a holdout an agent will respect: How to choose a holdout set for agent-driven catalogue changes so that it stays comparable, stays excluded, and survives contact with an automated process.
Read the full article below for detailed insights and actionable strategies.
The attribution problem
One sale. Four channels. 400% credit claimed.
Reported revenue: €400 · Actual revenue: €100 · Gap: €300
A holdout only works if it is chosen before results exist and excluded by something the agent cannot quietly override. Both halves fail regularly.
Choosing the set
Comparable, not random-looking. The held-back products should resemble the changed ones on the things that drive sales: price band, category, age in the catalogue, recent velocity.
A holdout made of your slowest products is not a control. It is a different population, and the comparison will show a difference that was there before you started.
Sizing it
Big enough that its own noise does not swamp the effect you are looking for. Small enough that the forgone upside is affordable.
There is no universal fraction. The honest version is that a holdout too small to resolve anything is a cost with no information, and a read that reports a confidence interval spanning break-even is usually telling you the design was underpowered rather than that the change did nothing.
Making the exclusion stick
| Weak exclusion | What goes wrong |
|---|---|
| A note in a doc | The agent never reads it |
| A tag the agent can edit | It gets rewritten in a later batch |
| Manual re-checking | Works twice, then nobody does it |
| A hard filter in the agent's instructions plus a post-batch diff | Holds |
The last row is the one that survives. Instruct the exclusion, then verify it after each batch by diffing the held-back set against its pre-batch state. An agent that was told to skip 50 products and skipped 47 has quietly ended your experiment, and only the diff will tell you.
Ending it properly
Decide the end date before you start, and read at the end date rather than when the numbers look interesting. Stopping a test early because it currently looks good is how a random fluctuation becomes a company policy.
The read
With a clean holdout and a logged boundary, a causal read on a Google Analytics export returns an estimate against the control rather than a before-and-after difference. It is 99 euro once, refunded if it does not move a budget decision. The interactive demo shows the output with no signup.
The stronger design, when the decision is expensive enough, is a geo holdout.
Related answers
Key Terms in This Article
Analytics
Analytics is the systematic computational analysis of data. It reveals customer behavior and measures campaign performance.
Attribution
Attribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
Causality
Causality is the relationship where one event directly causes another, essential for identifying specific actions that drive desired outcomes in marketing.
Confidence Interval
Confidence Interval is a statistical range of values that likely contains the true value of a metric. In marketing analytics, it quantifies uncertainty around estimates, indicating the precision of an outcome or causal effect.
Google Analytics
Google Analytics is a web analytics service that tracks and reports website traffic.
Incrementality
Incrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
Related Articles
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Ready to see your real numbers?
Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.
Full refund if you don't see value.
Stay ahead of the attribution curve
Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.
No spam. Unsubscribe anytime. We respect your data.
Frequently Asked Questions
How do I choose a holdout set for catalogue changes?
Pick products comparable to the changed ones on price band, category, catalogue age and recent velocity. A holdout made of your slowest products is a different population, not a control.
How do I stop an agent from changing the holdout?
Instruct the exclusion and then verify it with a post-batch diff against the pre-batch state. An agent told to skip 50 products that skipped 47 has ended the experiment, and only the diff reveals it.
When should I read the result?
At the end date you set before starting. Stopping early because the numbers currently look good is how a random fluctuation turns into company policy.