Attribution Tools That Report Confidence Intervals: Learn why confidence intervals matter more than point estimates in attribution, and which tools report them for better budget decisions.
Read the full article below for detailed insights and actionable strategies.
The attribution problem
One sale. Four channels. 400% credit claimed.
Reported revenue: €400 · Actual revenue: €100 · Gap: €300
A practitioner's look at why the interval, not the point estimate, is the number that should move budget in the room.
Updated 8 September 2026 · Joris van Huët, founder, Causality Engine
The worst attribution meetings I have sat in all had the same shape. Someone pulls up a dashboard, points at a channel with 3.1x ROAS, and says "we should put more money here." Nobody in the room can say how sure the model is about that 3.1x. So the decision gets made on a single number, and three months later the incremental revenue does not show up.
A point estimate tells the room what the model thinks. It does not tell the room how risky it is to act as if the model were right. That second thing is the whole job of a confidence interval, and it is the reason I now refuse to move budget off a number that arrives without a range attached.
Why a point estimate on its own is not a decision
A point estimate is the model's central guess for a quantity, like the incremental revenue from a channel. An interval is the range around that guess. Recast puts the distinction plainly in its documentation: the point estimate is an average, and the interval shows how precise that average is.
Here is the problem a single number hides. Say two channels both come back at 2.0x ROAS.
- Channel A: 2.0x, interval 1.8x to 2.2x
- Channel B: 2.0x, interval 0.3x to 5.7x
Same central estimate. Completely different decisions. Channel A is a tight, reliable read. Channel B is a coin flip that happens to average out to 2.0x. If you treat them as identical because the headline number matches, you are going to overfund a channel the model barely understands.
A point estimate alone cannot tell you whether a channel is reliably above break-even, whether two channels are meaningfully different or effectively tied, or whether a proposed spend increase sits inside the range the model has actually observed. Those are the questions a budget meeting is supposed to answer.
There is a fourth thing a defensible number states, and it is the one nobody includes. Not "the channel delivered a 12% lift" but "12%, with an interval of 4 to 20, from a design whose minimum detectable effect was 8%." The interval says how precise the estimate is. The floor says whether the design could have seen the effect at all. An interval without its floor is incomplete.
This is not an argument for waiting until all uncertainty disappears. That day never comes. It is an argument for making the risk visible so the size of the move matches the strength of the evidence.
What a per-channel interval means in a budget meeting
Once you have an interval on every channel, the conversation changes. It stops being "which channel has the highest reported ROAS" and becomes "which budget move has the best expected return given the uncertainty and the constraint we are under."
A few questions I actually ask when I read a per-channel interval:
- Does the lower bound stay economically attractive? If break-even is 1.0x and the interval bottoms out at 1.4x, the downside is still profitable.
- Does the interval cross the break-even threshold? If it does, the channel is not clearly profitable from the model's evidence.
- Do two channels' intervals overlap heavily? If they do, do not claim the one with the higher point estimate is the winner.
- Is the recommended spend inside the range where the model has seen real variation? Recast notes that uncertainty tends to rise when you push spend beyond the historically observed range, which is exactly when overconfidence hurts.
- Is the effect I am looking for bigger than the floor? Multiply the channel's share of revenue by the return you would honestly defend for it; if that is smaller than the smallest lift the design can detect (about 8% for a typical DTC brand with six months of history and an eight-week test), a wide interval is not a mystery. It is the arithmetic telling you the channel is not measurable at this scale.
A quick decision guide I keep in my head:
| What the interval looks like | Reasonable action |
|---|---|
| Narrow and above break-even | Increase spend with more confidence, watching for saturation |
| Wide but attractive point estimate | Treat it as a hypothesis, make a small move or test |
| Interval crosses break-even | Hold, reduce, or validate before adding budget |
| Overlaps another channel's interval | Do not over-rank one above the other |
| Widens sharply at higher spend | Model is extrapolating, be cautious about scaling |
| Expected effect below the floor | Manage on judgement, openly; no test will settle it at this scale |
One thing worth being precise about: a wide interval does not mean a channel is bad. It means the estimate is imprecise. The right response might be more data, a controlled test, or a smaller move, not a cut.
A worked example you can read in ten seconds
This is illustrative, not a vendor result. Say your break-even ROAS is 1.0x.
Search comes back at 2.0x, interval 1.2x to 2.8x. Social comes back at 2.4x, interval 0.4x to 4.8x.
On point estimates alone, Social wins. It has the higher central number, so the naive call is "move money into Social."
Read the intervals and the call flips. Search's entire plausible range sits above break-even, so even the pessimistic case is profitable. Social's range dips to 0.4x, well under break-even, which means the model thinks there is a real chance the channel is losing money at current spend. The defensible move is a confident increase on Search and a small, constrained bet or an incrementality test on Social.
The interval reversed the conclusion the point estimates suggested. That happens more often than people expect, which is why I stopped trusting rankings built on central numbers alone.
A note on terminology, because vendors are not consistent
Not every interval is a 95% interval, and not every interval means the same thing. Frequentist tools usually say "confidence interval." Bayesian tools usually say "credible interval," which comes from the posterior distribution of the model. Measured describes credible intervals as the Bayesian version of confidence intervals. Google Meridian is a Bayesian MMM and reports ROI with credible intervals. Recast's stability documentation describes its interval as containing the middle 50% of model estimates, which is not the same as a 95% band.
The practical takeaway: check the convention on the specific report before you interpret the width. A middle-50% interval and a 95% interval look very different even when the model is identical.
Which tools publish intervals, and what each one recommends
These tools do not all belong to the same category. Most of the list is media mix modeling or media effectiveness. Causality Engine is export-based causal attribution built for ecommerce teams on GA4 and Shopify. The shared output is budget guidance. The inputs, setup, and workflow are very different.
| Tool | What it publishes about intervals | Kind of budget recommendation | Caveat |
|---|---|---|---|
| Recast | Reports uncertainty on outputs, describes a Bayesian credible interval holding the middle 50% of estimates, and an optimizer that reduces weight on higher-risk channels | Starts from your current budget and proposes per-channel increases or decreases, with conservative-to-aggressive settings | Do not describe its interval as a 95% band; the documented one is middle-50% |
| Measured | Shows lift ranges and confidence intervals, describes credible intervals as the Bayesian form of interval uncertainty | Response curves, what-if budget scenarios, reallocation based on validated incrementality curves | Default confidence level not established on cited pages; calibrates with geo tests |
| Cassandra | Exposes an uncertainty interval around each channel's incremental ROI; Budget Allocator returns predicted revenue with confidence intervals | Time-and-channel media plan with a week-by-channel spend matrix and constraint report | Distinguish the modeled allocator output from a separately run geo or lift test |
| Prescient AI | Provides confidence intervals and forecast ranges; a finance page shows both 50% and 95% forecast intervals | Scenario-based recommendations comparing projected revenue and ROAS before committing spend | Also publishes a separate confidence score based on data coverage; that score is not the interval |
| Google Meridian | Open-source Bayesian MMM reporting ROI with credible intervals, contributions, and response curves | Budget optimization and scenario planning, including fixed-budget and target-ROI optimization | A framework, not a turnkey subscription; you supply, prepare, validate, and interpret |
| Causality Engine | Per-channel causal attribution and incremental ROAS with a 90% confidence interval on every estimate, per its pricing page for agents | Plain-English action view: what to scale, cut, and test, plus a budget reallocation view | Reads a GA4 export, no pixel; not a geo-test tool and not an identity graph |
If you run an ecommerce brand on GA4 and Shopify and you want a defensible budget read this week, Causality Engine is the one I would reach for first out of this list, and I will explain the tradeoff honestly.
Where Causality Engine fits, and where it does not
Causality Engine reads a historical GA4 CSV export. You can add a Shopify export when you have one. There is no pixel, no SDK, no DNS change, no developer ticket, and no onboarding project. The output is per-channel incremental ROAS, causal contribution, a confidence interval on each estimate, a platform-reported versus causal comparison that surfaces over-attribution, and a budget reallocation view.
Pricing is deliberately low-friction. A one-time causal read is 99 euro. Pro is 299 euro per month and adds automated GA4 ingestion, an AI chat across your historical attribution data, real-time budget alerts, and API access. The continuous monitoring and API belong to Pro, not to the 99 euro read, so do not expect ongoing ingestion from a single one-time analysis.
Now the part most vendor pages skip. Causality Engine does not run geo tests. It does not build a cross-device identity graph. Its read is causal inference from your existing sales, spend, and conversion history, treating the marketing history you already have like an experiment that already ran. That is a genuinely different method from a prospective geo holdout, and it is fair to say so.
My read on the tradeoff: a fast, export-based causal read is the right first move for a budget conversation because it gives you interval-backed numbers in minutes instead of standing up a modeling project or waiting weeks for a geo test to finish. If your team needs prospective experimental proof for a very large spend decision, a geo test remains a separate validation method, and the two are complementary rather than competing. Between anchors you are running on a model, and confidence decays from the date of the last qualified holdout; put that date on the dashboard next to the intervals. For the everyday "what do we scale, cut, or test next" question that most ecommerce teams face every month, the export read wins on speed and cost, and it still ships the interval you need to size the move.
Interval quality is not the same as model validity
One more thing that trips people up. An interval expresses model uncertainty. It does not prove the model got the causal structure right. A confident, narrow interval on a wrong model is still wrong, just precisely so. Under selection, model fit and causal accuracy can even run in opposite directions: a model that describes the observed data beautifully may be reproducing the platform's targeting rather than the advertising's effect.
Validation comes from geo holdouts, conversion-lift experiments, and calibration against experiment results. Measured says it calibrates with geo-based incrementality tests. Prescient describes calibration with geo-lift and holdout tests. Those are validation practices sitting alongside the interval, not replacements for it. Treat vendor backtesting figures as vendor claims, not independent proof, and press any tool on how its results have been checked against real experiments.
We did exactly that. Between 14 August and 2 September 2026 we audited thirty-one commercial measurement vendors for a published validation of their method against randomised experiments, with the sample, the design and the discrepancies disclosed. We found none. Hold Causality Engine to the same question; an interval is a claim about precision, not a validation, ours included.
What I would do first
If you are staring at conflicting dashboard numbers right now:
- Stop ranking channels by point estimate. Pull the interval on every channel you care about.
- Check where each interval sits relative to your break-even ROAS, not relative to each other.
- Check each channel's expected effect against the floor. The ones below it are not measurable at your scale; say so and manage them on judgement.
- Fund the channels whose lower bound stays profitable. Test the ones with wide intervals or ranges that cross break-even.
- For a quick, defensible read you can bring to the next budget meeting, run a one-time causal read against your GA4 export before you commission anything heavier.
The discipline is simple to state and hard to hold under pressure: move money confidently when the relevant lower bound is still attractive, move cautiously when the point estimate is high but the range is wide, and do not pretend two overlapping channels are meaningfully ranked. The interval is the number that keeps a good headline from becoming a bad decision.
FAQ
Is a confidence interval the same as a credible interval?
Not exactly, and the difference matters when you compare tools. A confidence interval is the frequentist version. A credible interval comes from a Bayesian model's posterior distribution. Most MMM tools here, including Meridian and Recast, report credible intervals. In a budget meeting the practical use is similar, but do not assume both were built the same way or at the same confidence level.
Does a wide interval mean a channel is performing badly?
No. A wide interval means the estimate is imprecise, not that the channel is weak. The channel could be genuinely strong; the model just does not have enough clean variation in the data to pin it down. The right response is usually more data, a controlled test, or a smaller move, not an immediate cut.
Can Causality Engine replace a geo test?
For most monthly budget decisions on an ecommerce brand, its export-based causal read gives you interval-backed answers fast enough to act on. For a very large spend commitment where you need prospective experimental evidence, a geo test is a separate method and still has a role. I treat them as complementary rather than one replacing the other.
What do I get for the 99 euro read versus Pro?
The 99 euro one-time read gives you a single causal analysis of up to 40 days of GA4 export, with per-channel incremental ROAS, confidence intervals, and budget reallocation recommendations, refunded if it does not move a single budget decision. Continuous ingestion, real-time alerts, the AI chat, and API access are part of the 299 euro per month Pro plan. Do not expect ongoing monitoring from the one-time read.
Sources and further reading
- The Price of Being Found (Causality Engine, Edition 2.10, September 2026): the measurability floor (Chapter 15), the vendor audit (Chapter 17), what a defensible number looks like (Chapter 19)
- Understanding true ROAS
- Meridian scenario planning documentation
- Sixty-second versions of these ideas on YouTube Shorts
Vendor prices and features quoted in this article were taken from each vendor's own website on 8 September 2026 and may have changed since. Check the vendor's pricing page before relying on a figure.
Get attribution insights in your inbox
One email per week. No spam. Unsubscribe anytime.
Key Terms in This Article
Attribution
Attribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
Causal Analysis
Causal Analysis identifies true cause-and-effect relationships in data, moving beyond correlation to show how marketing actions directly impact outcomes.
Causal Attribution
Causal Attribution uses causal inference to determine which marketing touchpoints genuinely cause conversions, not just correlate with them.
Causal Inference
Causal Inference determines the independent, actual effect of a phenomenon within a system, identifying true cause-and-effect relationships.
Confidence Interval
Confidence Interval is a statistical range of values that likely contains the true value of a metric. In marketing analytics, it quantifies uncertainty around estimates, indicating the precision of an outcome or causal effect.
Holdout Test
A holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
Incrementality
Incrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
Media Mix Modeling
Media Mix Modeling is a statistical technique that measures the collective impact of marketing and advertising on sales. It uses historical data to inform budget allocation.
Related Articles
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Ready to see your real numbers?
Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.
Full refund if you don't see value.
Stay ahead of the attribution curve
Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.
No spam. Unsubscribe anytime. We respect your data.