Skip to content

A 30-day ecommerce attribution playbook

Reconcile platform claims to your Shopify orders first, fix tagging second, size one holdout test third and start it fourth. The plan ends at launch, because a holdout has to be big enough and long enough to be readable.

By , Founder & CEOPublished 6 min read

Run the numbers for your store: the free UTM coverage calculator.

In a month you can go from dashboards that disagree to one holdout test that is sized, written down and running: reconcile in the first week, fix tagging in the second, size the test in the third, start it in the fourth. The plan ends at launch, not at a verdict, because Google recommends a minimum of 14 days for its own lift studies (Google Ads Help, 30 September 2026). And a test too small to see the effect you need comes back inconclusive: across 25 large field experiments, the median confidence interval on return on investment was over 100 percentage points wide (Lewis and Rao, Quarterly Journal of Economics, 2015).

What do you do in the first week?

Reconcile. Use one complete week that ended at least five days ago, because Google says modelled conversions can take up to 5 days to stabilise in its reporting (Google Ads Help, 30 September 2026). For that week, in one currency and one time zone, write down:

  • For each ad platform: the purchases and revenue it credits, and the attribution window it used.
  • For Shopify: orders and revenue after refunds, all sources together.
  • For GA4: revenue by Session default channel group, and the share that sits in Unassigned.

Pass: no platform credits more purchases than the store took orders. Fail: one does, which points to duplicate events or a date mismatch; Google Ads vs Shopify revenue and Meta Ads vs Shopify revenue list the usual causes. Then add up the revenue the platforms credit and divide by Shopify revenue. That claim ratio is expected to be above one, because every platform grades its own homework and they overlap. The excess is the double counting, not a fault.

By the end of the week, decide which platform numbers are a ranking, not revenue.

What do you fix in the second week?

Tagging, then baselines. Put every paid link on one naming list (UTM naming rules has them), then read the Unassigned share again. Then write three baseline numbers: your break-even ROAS, which is one divided by your contribution margin; your MER, total revenue over total ad spend; and the claim ratio from the first week. Which metric to trust has the method.

Pass: the Unassigned share falls, or you can name the links that cause it. Fail: you can't say where it comes from. Don't size a test on numbers you can't explain.

How do you size the test in the third week?

Pick one channel: the one where a wrong call costs most and the evidence is thinnest. If your account qualifies, a platform lift study is randomised and run by the platform, a sound option. Qualifying is the catch. Google's users-based Conversion Lift needs at least 1,000 observed conversions and a minimum campaign budget of $5,000 USD, and "isn't available for all Google Ads accounts". TikTok's Conversion Lift Study is a managed service for eligible accounts that meet minimum ad spend requirements. Otherwise run a regional holdout, the design Google's geo-based lift page describes: comparable geographic groups, the channel switched off in some and kept on in the rest. Compare total Shopify revenue by region.

Then check that the test can see what you need it to see. Google's geo-based lift page defines minimum detectable iROAS as the effect size needed "to have a high chance to detect significant lift", and doesn't recommend proceeding on low feasibility. Do the sum by hand. For illustration: a channel whose spend is 5% of revenue and that returns 2 euros of revenue per euro would take about 10% of revenue with it in the regions where you switched it off (0.05 x 2 = 0.10). In this illustrative case, a design that can see only a 15% change can't answer, and one that can see 8% can. The holdout test planner gives the days a design needs for a smallest lift you choose. Pass: the change your decision needs is larger than the smallest change the design can see. Fail: it isn't, so extend the run, hold out more regions or pick a bigger channel.

Write the rule before you launch: the revenue gap between regions that means cut, hold or scale, and the date you will read it.

What do you start in the fourth week, and what do you decide at the end?

Start. Switch the channel off in the holdout regions, freeze price changes and promotions in both groups so they don't blur the comparison, and run it at least as long as your buyers take to convert. Google's users-based lift page says to set the study duration to a length that captures your average conversion lag, allows studies as short as 7 days and recommends a minimum of 14 (Google Ads Help, 30 September 2026). It also reports up to a 17% drop in Absolute Lift in studies with a long conversion lag that ran for less than 14 days (published by Google, not independently audited). Measuring conversion lag in GA4 shows yours.

On the last day of the plan, have four things on one page:

  1. The channel you are testing, and why it comes first.
  2. The holdout regions and the start date.
  3. The rule: what revenue gap means cut, hold or scale.
  4. The date you will read it, and what you will do if the answer is inconclusive.

Inconclusive is not "no effect". It means the test was too small or too short to see one, and the next step is a longer or larger test, not a cut.

What comes after the first holdout?

Later, once that first holdout has settled one channel, a causal attribution read like Causality Engine's can show what each of your other channels caused from a GA4 export, with a next step for each and a way to test it.

Sources, 30 September 2026: The Unfavorable Economics of Measuring the Returns to Advertising (Lewis and Rao, Quarterly Journal of Economics, 2015); Set up Conversion Lift based on users (Google Ads Help, 2026); Set up Conversion Lift based on geography (Google Ads Help, 2026); About modeled online conversions (Google Ads Help, 2026); About Conversion Lift Study (TikTok Business Help Center, 2025); Default channel group (Google Analytics Help, 2026). The 5% and 2 euro sums are arithmetic on assumed numbers, not benchmarks.

Frequently asked questions

  • How long should an ecommerce holdout test run?
    At least as long as your buyers take to convert. Google allows its own lift studies as short as 7 days and recommends a minimum of 14 (Google Ads Help, 30 September 2026). A shorter test with a long conversion lag can understate the effect.
  • Can I run a lift test with a small ad budget?
    Not always. Google's users-based Conversion Lift needs at least 1,000 observed conversions and a minimum campaign budget of $5,000 USD (Google Ads Help, 30 September 2026). Below that, a regional holdout using your own Shopify orders is the free route, and the holdout test planner shows how many days it needs.
  • What should I decide before the test starts?
    The channel, the regions that go dark, the revenue gap that means cut, hold or scale, and the date you will read it. A rule written after the result lets the result choose the rule.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.