Skip to content

For GA4 usersFrustrated with GA4 attribution? Upload your GA4 export, see causal insights in 5–10 minutes for €99 pay-per-use.

Insights

17 min read

How to Vet an Attribution Vendor Before You Sign

Learn how to vet an attribution vendor with a 7-question checklist to ensure defensible, reproducible, and finance-grade marketing measurement.

Share
Quick Answer·17 min read

How to Vet an Attribution Vendor Before You Sign: Learn how to vet an attribution vendor with a 7-question checklist to ensure defensible, reproducible, and finance-grade marketing measurement.

Read the full article below for detailed insights and actionable strategies.

The attribution problem

One sale. Four channels. 400% credit claimed.

100
1 sale
Meta
100%
claimed
Google
100%
claimed
TikTok
100%
claimed
Klaviyo
100%
claimed

Reported revenue: 400 · Actual revenue: 100 · Gap: €300

A vendor-call checklist a marketing lead or CFO can run to test whether an attribution number is defensible, reproducible, and safe to spend budget on.

Updated 8 September 2026 · Joris van Huët, founder, Causality Engine

The fastest way to end an attribution vendor call is to ask one question: "Show me a report where a channel got no number because the data was too thin." Watch what happens. The confident ones will pull up an example or describe the exact rule. The rest will talk about their model's power, or say every channel always gets a score. That second answer tells you almost everything you need before the demo even loads.

Most attribution buying goes wrong because the evaluation focuses on the wrong thing. Buyers compare dashboards, integrations, and logo walls. What actually gets defended in a budget meeting is not a dashboard. It is a claim: "This channel drove incremental revenue, and here is why we believe it." A vendor either helps you make that claim stand up under a CFO's challenge, or it hands you a black box with a nice UI.

Below is a checklist you can run live on a call. Nine questions, each scored 0 to 2, for a maximum of 18. I use the IAB/MRC Retail Media Measurement Guidelines as the reference point, because they set a clean bar: methodology, definitions, assumptions, limitations, uncertainty, and data handling should all be accessible to the people using the measurement. That is the standard. Score the evidence a vendor will actually give you, not the branding words on the site. The Price of Being Found carries a longer list of twelve; the nine below are the ones that decide a call.

How to score the call

Use the same rubric for every question so the comparison stays honest:

  • 2 points: Specific, documented, testable answer. They can show the report, the methodology doc, the export, or the contract clause on the call.
  • 1 point: Plausible but incomplete. Verbal, proprietary, "we can share after implementation," or a high-level white paper only.
  • 0 points: Vague, evasive, unsupported certainty, or no defined process.

Totals map roughly like this:

  • 15 to 18: Strong transparency posture. Move to technical and legal review.
  • 10 to 14: Usable, but only with written conditions, a sample export, and a validation clause.
  • 0 to 9: Treat the output as a directional aid, not finance-grade measurement.

Do not award points for saying "AI," "incrementality," "full funnel," or "enterprise-grade." Those are not answers. They are decoration.

The nine questions

Question What a defensible answer includes What a bad answer sounds like
1. Do you publish your methodology? Data inputs, outcome and exposure definitions, model family, causal assumptions, covariates, priors, validation, robustness checks, exclusions, attribution window, and stated limitations. "The algorithm is proprietary." "It's AI-driven, the details won't help you."
2. What exactly is the counterfactual? A plain statement of the compared-against world, plus the estimand, treatment, outcome, unit, and time period. "The model finds the true contribution." "Incrementality is built in."
3. Have you validated this against randomised experiments on real data? The study, the sample, the design, and the discrepancies. Not backtests, not simulated data, not client testimony. "Our accuracy is industry-leading." "Clients see great results."
4. Do you report uncertainty? Per-channel intervals, stated level (90%, 95%), how they're computed, and how to read an interval that crosses zero or break-even. "It's statistically significant." "The point estimate is all you need."
5. What is my minimum detectable effect, and which of my channels are not measurable at my spend? A number per channel before you start, and a list. "All of them are measurable."
6. Can we reproduce it from an export we control? Exact input schema, channel mapping, exclusions, versioning, rerun procedure, machine-readable output. "The dashboard is the source of truth." "You can't run it outside our platform."
7. How do you explain gaps vs platform numbers? A reconciliation table: windows, dedup, view-through, overlap, reported vs incremental, with residual differences explained. "The platforms are wrong." "Our number is lower because we're more advanced."
8. What happens when data is too thin, and what would the system return if my advertising had no effect at all? Written thresholds, sparse-channel rules, missingness handling, a visible insufficient-data outcome, and a placebo test. "It works with any amount of data." "Every channel gets a number."
9. What does the contract lock us into, and where is data processed? Cancellation, renewal, refunds, retention, deletion, ownership, subprocessors, AI processing, residency, governing law. "It's month to month, nothing else to review." "EU hosting means all processing is in the EU."

1. Do you publish your methodology?

Ask for the methodology document before you sign, not after. It should let a competent analyst understand what was estimated and how. The IAB/MRC framing is that methodology, data collection, processing, calculations, reliability, and limitations should be accessible to measurement users.

Here is the distinction that trips people up. A published methodology does not make a result correct. It makes the result inspectable. That is the whole game. A vendor that refuses to describe its model is asking you to trust a number you cannot examine, and that number will not survive a serious budget challenge.

A vendor can legitimately describe its method without publishing every prior distribution or exact robustness test on a public page. That is fine. What is not fine is "the algorithm is too complex to explain." Complexity is not a reason to hide the shape of the model, the assumptions, and the limitations.

2. What exactly is the counterfactual?

This is the question that separates causal measurement from credit allocation dressed up in causal language. Incrementality means the effect of marketing above a baseline. So the vendor has to tell you what the baseline is.

Make them say it in plain English. "Sales if this channel's spend had been lower." "Conversions without exposure." "Outcomes in a comparable untreated market." Then ask for the estimand: incremental revenue, incremental conversions, average treatment effect, and over what window. If the answer stays at "the model assigns credit," you are looking at a multi-touch weighting scheme, not a causal estimate. Score it a 0 and move on.

3. Have you validated this against randomised experiments on real data?

Show me the study, the sample, and the discrepancies. Not a backtest, which measures how well a model fits data it has already seen. Not simulated data. Not a client quote.

Here is why this question matters more than the demo. Under selection, model fit and causal accuracy can run in opposite directions. The naive comparison of exposed to unexposed users in the best-known experimental study of this kind fit the observed data beautifully and was still wrong by more than a factor of four, because the people the system chose to expose were the people most likely to convert anyway. The model did not fail to fit. It fit the selection. There is experimental variation, or there is nothing.

Between 14 August and 2 September 2026 we put this question to the public materials of thirty-one commercial measurement vendors: websites, documentation, research pages, case studies. We found no published comparison of a vendor's estimates against randomised experimental results with the sample, the design and the discrepancies disclosed. Some publish accuracy claims, some publish case studies, some publish backtests. That is the state of the field, and it is why this question earns a row of its own.

4. Do you report uncertainty?

A single ROAS figure with no range around it is a marketing claim, not a measurement. The IAB/MRC guidelines are direct here: when sampling, modeling, or estimation introduces uncertainty, error rates and confidence intervals should be reported.

Ask for per-channel intervals, the confidence level, how they are calculated, and how to interpret specific cases. What does an interval that crosses zero mean for that channel? One that crosses 1.0x ROAS? A very wide interval? Those are the situations where budget decisions get made, and "the score is highly accurate" is the answer of a vendor that does not want you looking too closely. "We don't show intervals because clients find them confusing" is worse. It means they made the number less honest to make it more sellable.

5. What is my minimum detectable effect, and which of my channels are not measurable?

Every design has a smallest lift it can reliably distinguish from nothing, and it can be computed before anyone spends anything. Multiply a channel's share of revenue (spend divided by revenue) by the return you would honestly defend for it; that is roughly how much total revenue would move if the channel stopped. If that is smaller than the design's minimum detectable effect (about 8% for a typical DTC brand with six months of history and an eight-week geo test), the channel is not measurable at your scale by that method, and a null result will be read as an answer when it is not one.

A vendor who says "all of them are measurable" has just failed the arithmetic. The honest answer for most advertisers is a list, and it is not short. This is the question that separates a measurement partner from a measurement salesperson.

6. Can you reproduce it from an export you control?

If the only place the result exists is the vendor's dashboard, you cannot audit it, and neither can anyone who challenges it later. Ask what you export, the exact schema, the date range, the channel mapping, how missing or unattributed traffic is handled, timezone treatment, and whether the same files rerun to the same result. Ask for the channel-level output in CSV, not just a slide.

The strongest position is: you control the input files, you receive machine-readable output, and the methodology version travels with the report. "You can't reproduce it outside our platform" is a 0. So is any workflow that requires vendor-controlled tracking just to see how the number was built.

7. How do you explain disagreements with platform numbers?

Every attribution vendor's number will differ from what Meta, Google, and TikTok claim. That is expected. Platforms report attributed conversions inside their own windows, often counting the same conversion across multiple channels. A causal estimate reports incremental outcomes. Different question, different answer.

Do not score a vendor down because its number is lower than the platform's. Score it down if it cannot explain why. Ask for a reconciliation: platform-reported by channel, the platform windows, the vendor's estimate and window, deduplication treatment, and the gap between reported and incremental, with the uncertainty interval attached. The IAB/MRC guidance leans on consistent, disclosed windows and day-level granularity precisely so these figures can be reconciled. "The platforms are wrong" is not a reconciliation. It is a shrug.

Bring your own number to that conversation: every platform's claimed conversions added up, divided by the orders you actually shipped. That claim ratio costs nothing to compute, comes from a source that does not sell you media, and is the single most effective defence against a scoreboard you do not own.

8. What happens when the data is too thin, and what would you return if my advertising had no effect?

This is the tell. A vendor that guarantees a number for every channel regardless of data quality is not measuring, it is decorating. Ask for minimum data requirements, sparse-channel rules, how it handles a channel that never pauses or varies, and how it treats promotions, stockouts, launches, and tracking gaps.

A defensible vendor will pause, suppress the estimate, widen the interval, mark it directional, or say the effect is not identifiable. Ask to see a report where a channel is flagged or omitted for insufficient data. Then ask the placebo question: what would this system produce if my advertising had no effect at all, and have you tested that? If they cannot produce either, you are looking at a tool that will confidently rank a channel it knows nothing about, and you will find out the hard way when you move budget on it.

9. What does the contract lock you into, and where is the data processed?

A low monthly price is not the same as no commitment. Read cancellation timing, renewal, refunds, price-change notice, data retention, deletion, ownership, and export rights. Then read the data-processing terms separately, because "we host in the EU" and "all processing happens in the EU" are different statements. Subprocessors, AI inference providers, and cross-region transfers all live in the gap between those two sentences.

The accurate question is: which data is processed by which provider, in which country, for which purpose, and for how long? Get the subprocessor list and the DPA. "EU hosting means everything is in the EU" is a 0, because it is almost never true once you look at the AI inference layer.

Running the checklist against Causality Engine

I will use Causality Engine as a worked example, because it maps cleanly onto the checklist and because it is upfront about its edges. It runs Bayesian causal inference on a GA4 export, with Shopify data supported, and requires no pixel, SDK, or engineering ticket. For an ecommerce brand on GA4 that wants a per-channel read without an onboarding project, it is the option I would run this checklist on first, and I would still score it question by question.

On the counterfactual, its public material describes the compared-against world as one where a given channel's spend was lower, and it frames output as incremental ROAS, meaning revenue caused rather than revenue claimed. That is a real answer to question 2. On uncertainty, its pricing page for agents states 90% confidence intervals on all estimates. One precise thing to raise on the call: it publicly uses the term "confidence intervals," while a Bayesian implementation is often discussed as credible or posterior intervals. Ask which construction applies to the delivered report. That is not a gotcha, it is the kind of clarity that matters when someone in finance asks how the range was built.

On question 3, hold us to the same standard as everyone else. The audit above found no published validation against randomised experiments among the thirty-one vendors it covered; ask Causality Engine for one, with the sample and the discrepancies, and if we cannot show it to you, weigh the read accordingly. The read is a model on observational data, not an experiment, and it should be scored as one.

On platform comparison, it promises a platform-reported versus causal view and explains that platform dashboards can count the same conversion across multiple channels. That covers question 7, provided you ask them to show the windows and dedup treatment live. The true ROAS guide on their site is a reasonable primer on why the numbers diverge.

Now the edges, stated plainly. Causality Engine does not offer geo tests or geo holdout experiments as part of the described product. It runs an analysis from existing data rather than a designed holdout. It also does not present an identity graph in its public product pages; the offer is built around Shopify and GA4 exports, not a proprietary cross-device identity layer. If your evaluation weights geo testing or an identity graph heavily, factor that in. Neither absence is hidden, and for most GA4-based brands neither is a dealbreaker, but you should decide that on purpose.

Three things are not publicly confirmed and belong in your call. First, the exact GA4 export schema and the rerun procedure for question 6. The report is exportable, but ask to see the precise fields and whether the same files regenerate the same result. Second, the thin-data behavior for question 8. The public material describes reads on 40 to 90 days of history, states a fit floor of roughly 5,000 euro in monthly paid spend, and says the methodology document (prior, functional form, covariates, robustness checks) goes to any customer who asks. I did not find published sparse-channel thresholds or a stated insufficient-data outcome. Ask them to show a flagged channel. Third, question 5: ask for the list of your channels that are not measurable at your spend. It is the question we most want to be asked.

On contract and data, pricing is transparent: €99 for a one-time read, €299 per month for Pro, with the company registered in Utrecht and primary infrastructure under EU data residency. Its privacy policy names the AI inference providers it uses (OpenAI, Anthropic, Google, and PublicAI, for real-time inference only) and says source data for a one-time read is retained for 40 days, while Pro data is retained for the duration of the active subscription. That disclosure is a point in its favor. Still run question 9 fully, because "EU primary infrastructure" and "every processing location" are the two different statements from earlier.

What I would do first

Before any call, export a real GA4 and Shopify dataset you already control and keep it. That single file is your leverage. It lets you test reproducibility, hand the same input to two vendors, and later re-run an audit without depending on anyone's dashboard.

Then run the nine questions in order, score as you go, and refuse to be talked out of the low scores by a good demo. A slick interface is not a methodology. The vendor that scores a 16 and admits two open items is safer than the one that scores an 18 by never saying "we don't know."

FAQ

Is a published methodology enough to trust the number?

No, and this is the most common misread. Publishing the method makes it inspectable. Whether the number deserves weight depends on validation, appropriate uncertainty, and whether the data supported identification in the first place. Transparency is the entry fee, not the verdict.

Should I reject a vendor whose ROAS is far below what Meta reports?

Not on that basis. Platforms report attributed conversions inside their own windows and routinely double-count across channels. A causal estimate answers a different question. Reject the vendor only if it cannot reconcile the gap with windows, deduplication, and the reported-versus-incremental distinction.

Do I need geo experiments for defensible attribution?

It depends on your stakes and spend. Geo holdouts are a strong causal design, but they require setup and control, and they have a floor of their own: one treated region against N regions means the smallest p-value a placebo-based design can return is 1 in N. A well-specified causal model on data you control, with stated assumptions and visible uncertainty, is defensible for many GA4-based ecommerce brands as the instrument between experiments. Just know which one you are buying. Causality Engine, for instance, runs the model-based read and does not offer geo holdouts, so if you need those you would pair it with a separate testing approach.

What is the single most revealing question on the call?

"Which of my channels are not measurable at my current spend?" A vendor that hands you a list has done the arithmetic. A vendor that says all of them is selling confidence, not measurement. The thin-data question from the opening is the same question wearing different clothes.

Sources and further reading

Vendor prices and features quoted in this article were taken from each vendor's own website on 8 September 2026 and may have changed since. Check the vendor's pricing page before relying on a figure.

Get attribution insights in your inbox

One email per week. No spam. Unsubscribe anytime.

Key Terms in This Article

Related Articles

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Ready to see your real numbers?

Own the budget? Upload your GA4 export and see which channels drive incremental sales, with confidence intervals, in minutes. Have to defend it? Start with the live demo and take the read to your CFO.

Full refund if you don't see value.

Stay ahead of the attribution curve

Weekly insights on marketing attribution, incrementality testing, and data-driven growth. Written for the person who owns the budget and the person who has to defend it.

Which one are you? Optional.

No spam. Unsubscribe anytime. We respect your data.

Related reports

Real reports on this topic.

Anonymised reports from the Attribution Report Library tagged with insights.

Browse all related reports

Find your wasted ad spend in 5–10 minutes.

Watch the model work on a sample store first, no signup. Then upload your last 40–90 days of GA4 sessions and get incremental ROAS with confidence intervals. No pixel, no SDK. €99 per read.

Prefer to talk it through? Book a 20-min call, or read how it works.

Last-click guesses.We run the math.

Causal attribution for ecommerce brands. Watch the model work on a sample store first, then upload your GA4 export and see which channels really drove revenue in 5–10 minutes. €99, pay-per-use. Pro at €299/mo when you want it continuous.

No signup for the demo. Book a 20-min call or compare plans.