Evaluate an attribution platform in a week: a scorecard
Test four things in a week: totals that reconcile to Shopify, one number rebuilt from order rows, a known channel change, and a comparison with an experiment. Score each pass, partial or fail before you commit.
€99 once, excl. VAT. Full refund within 30 days, no questions asked. You keep the read.
By Joris van Huët, Founder & CEOPublished 5 min read
In a week you can test four things about an attribution platform without trusting its demo: whether its totals reconcile to Shopify, whether one number can be rebuilt from order rows, whether a channel change you already know about shows up the right way, and whether it can show a comparison with an experiment. The last matters most. In 15 Facebook experiments, observational methods "often fail to produce the same effects as the randomized experiments" (Gordon and colleagues, Marketing Science, 2019), so score each test pass, partial or fail before you commit.
A demo shows what the vendor chose to show. A trial should run on the week you pick.
What should reconcile in the first two days?
Day one: totals. Pick one closed week that ended at least three days ago: Google says it updates daily export tables for "up to 2 calendar days, plus today" after the table date, so a fresher week is still moving. Compare the tool's orders and revenue with Shopify's for that week. Expect gaps and ask for a reason for each. Shopify's Admin API separates an order's total "at the time of order creation" from its total "after returns", so ask which one the tool reports. Check time zones too: Google says Analytics reports in the property's time zone and Google Ads in the account's, and that different settings "might" give discrepancies.
Day two: one number. Choose one channel on one day and ask for the order IDs behind its credit. Every ID should exist in Shopify with a matching total. IDs must be unique: Google Analytics deduplicates purchases that share a transaction ID, and warns not to send an empty string as the ID.
Pass: every gap has a reason you can see in your own data. Fail: an unexplained gap, or no list of the orders counted.
What does it do with Direct and with declined consent?
Day three. Ask for credit by channel with Direct and Unassigned shown. Google's own models give Direct no credit unless the whole path is direct, so a tool that treats it differently should say how and show the effect on each channel. Then subtract the tool's order count from Shopify's. After cancellations and test orders, what is left is the orders it cannot see, and it should be labelled as such. If the tool reads GA4's BigQuery export, expect it to differ from the GA4 interface too; Google says the export "might differ".
Does it move when you know something changed?
Days four and five. Take a change you can date: a channel you paused or launched, or a promotion week. Check whether the tool's credit for that channel moves the right way, at the right time. Pass: direction and timing match what you did. Fail: no movement, or movement before the change. A tool can get direction right and size wrong, which is why the next test exists.
What was it compared with, and can you leave with the rows?
Day six: the experiment. Ask for any comparison of the tool's output with a randomized test: who ran it, the sample, the dates, what was compared. A platform lift study is the yardstick where you can run one. Google describes Conversion Lift as comparing users who saw ads with users "held back" in a control group, and says it "intentionally ignores the standard conversion tracking settings and attribution rules". So compare its incremental conversions with the attributed conversions for the same campaigns and dates. A vendor's own comparison counts if it says so and shows its method.
Day seven: exit. Ask for an order-level export and recompute your weekly totals from it. Ask how much Shopify history the tool can read: Shopify says only "the last 60 days' worth of orders" are accessible to an app from the Order object by default, and older orders need a request for access to all orders. Ask what is deleted when you cancel.
How do you score it?
| Test | Evidence you collect | Pass | Fail |
|---|---|---|---|
| Totals | Orders and revenue for one closed week | Every gap has a visible reason | Unexplained gap |
| One number | Order IDs behind one channel and day | IDs exist, unique, totals match | Missing or duplicate IDs |
| Direct and consent | Direct credit, orders the tool cannot see | Both stated and labelled | Silence |
| Known change | A dated pause, launch or promotion | Credit moves the right way on time | No movement, or early movement |
| Experiment | A named comparison with a randomized test | Method, sample, dates, results | An accuracy claim with no test |
| Export | Order-level rows, history, deletion terms | Totals recompute from the rows | Dashboards only |
A platform that fails the experiment row can still be useful for reconciliation and path reports. The score tells you how much weight its channel credit deserves.
Sources, 30 September 2026: A Comparison of Approaches to Advertising Measurement (Gordon and colleagues, Marketing Science, 2019; abstract via IDEAS/RePEc); Set up BigQuery Export (Google Analytics Help, 2026); Select attribution settings (Google Analytics Help, 2026); Minimize duplicate key events with transaction IDs (Google Analytics Help, 2026); Get started with attribution (Google Analytics Help, 2026); Understand your Conversion Lift based on users measurement data (Google Ads Help, 2026); Order object (Shopify GraphQL Admin API documentation, version 2026-07).
Related answers
Frequently asked questions
How do I evaluate an attribution tool during a trial?
Use a closed week of your own data. Reconcile its orders and revenue to Shopify, rebuild one channel-day from order IDs, check a change you can date, ask what it was compared with, and export the rows. Score each test pass, partial or fail.What should an attribution tool's numbers match?
Shopify's orders for the same week, with every gap explained: returns, cancellations, time zones and visitors who declined consent. Shopify's Admin API separates an order's total at creation from its total after returns, so check which one the tool reports.How do I ask a vendor to show its attribution is accurate?
Ask for a comparison with a randomized test: who ran it, the sample, the dates and what was compared. In 15 Facebook experiments, observational methods often failed to match randomized results. A vendor's own comparison counts if it is labelled as such and shows its method.
Go deeper: Incrementality testing, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Attribution PlatformAttribution Platform is a software tool that connects marketing activities to customer actions. It tracks touchpoints across channels to measure campaign impact.
- Attribution SoftwareAttribution Software measures campaign impact by tracking customer interactions across touchpoints. It assigns value to each channel, showing what drives conversions.
- Control GroupControl Group is a segment of an audience intentionally not exposed to a marketing campaign, used to measure the campaign's true causal impact.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Multi-Touch AttributionMulti-Touch Attribution assigns credit to multiple marketing touchpoints across the customer journey. It provides a comprehensive view of channel impact on conversions.