Reporting vs analysis: what a dashboard cannot decide
On its own, a dashboard cannot decide whether spend caused sales: it reports credit for orders that happened, and a decision needs the comparison it lacks. A holdout supplies that comparison, and your Shopify orders settle it.
By Joris van Huët, Founder & CEOPublished 5 min read
A dashboard reports what happened; a budget decision needs what would have happened without the spend, which no report observes. A peer-reviewed paper by Gordon, Zettelmeyer, Bhargava and Chapsky (Marketing Science, 2019) used data from 15 U.S. advertising experiments at Facebook, comprising 500 million user-experiment observations and 1.6 billion ad impressions, to contrast experimental results with multiple observational models. The observational methods often failed to produce the same effects as the randomized experiments, even after conditioning on extensive demographic and behavioral variables.
What does a dashboard actually report?
Credit. Google Analytics Help defines attribution as "the act of assigning credit for important user actions to different ads, clicks, and factors along the user's path to completing the action", and an attribution model as a rule, a set of rules or a data-driven algorithm that assigns that credit. That is bookkeeping on sales that happened. Google's documentation also says its data-driven model uses a counterfactual approach that contrasts what happened with what could have occurred, so the question behind a budget decision is the one the model is reaching for. Nobody observes the sale that did not happen, so the answer has to come from a comparison.
What do field experiments say about reading effects off the data?
They say it is hard either way. Lewis and Rao (Quarterly Journal of Economics, 2015), another peer-reviewed paper, report that twenty-five large field experiments with major U.S. retailers and brokerages, collectively representing $2.8 million in digital advertising expenditure, showed that measuring the returns to advertising is difficult. The median confidence interval on return on investment was over 100 percentage points wide, and informative advertising experiments can easily require more than 10 million person-weeks. They call selection bias, due to the targeted nature of advertising, a crippling concern for widely employed observational methods.
Both studies looked at advertising in the settings they describe, and neither looked at your store. No public measurement tells you how far your own dashboard sits from the truth, which is why the check below uses your orders.
How does a holdout supply the missing comparison?
Google Ads Help describes the design: Conversion Lift is a controlled experiment that splits an audience into a treatment group that sees the ads and a control group that does not, and the difference in conversions is the lift. Google describes incrementality experiments as "the way to measure the causal impact of ads". Google notes that Conversion Lift is not available for all Google Ads accounts, so the version below, with regions and your own Shopify orders, is the one anyone can run. For illustration, with invented numbers:
Dashboard credit for Paid Social, region A: 600 orders
Region A, ads on: 5,000 orders
Region B (matched before the test), ads off: 4,600 orders
Caused by the ads: 5,000 - 4,600 = 400 orders (two thirds of 600)
Would have arrived anyway: 200 orders
To run it on your own orders:
- Pick two regions that tracked each other. Chart the ratio of their weekly Shopify orders for the weeks before the test. Its normal swing is your noise.
- Change one channel in one region and nothing else. No promotion, launch or price change that touches only one side.
- Compare Shopify orders by shipping region, not platform conversions. Shopify's reports offer shipping region as a dimension.
- Pass: the gap moves by more than its normal swing and stays moved. Fail: it stays inside the swing. The test did not have the volume to tell, so say that instead of rounding it up to a result.
- Write the dates and the change next to the number.
The holdout test calculator counts the weeks your order volume needs, and cut a channel the right way covers the decision that follows.
What is a dashboard still good for?
Reporting, and it is quick at it: whether revenue fell, whether tracking broke, whether GA4 and Shopify agree. Use it for those, and ask the counterfactual question only for decisions that move budget.
Later, once a holdout has shown you what one channel did, a causal attribution read like Causality Engine's can show what each channel caused from a GA4 export, next to what last-click gave it.
Sources, 30 September 2026: A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook (Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, 2019); The Unfavorable Economics of Measuring the Returns to Advertising (Lewis and Rao, Quarterly Journal of Economics, 2015); About Conversion Lift (Google Ads Help, 2026); Get started with attribution (Google Analytics Help, 2026); Analytics data points (fields) reference (Shopify Help Center, 2026).
Related answers
Frequently asked questions
What is the difference between reporting and analysis in ecommerce?
Reporting states what happened: orders, revenue and the credit a model assigns. Analysis asks what would have happened without the spend, which needs a comparison group such as a holdout. A dashboard does the first; a designed comparison does the second.Can observational data measure ad effectiveness?
Not reliably in the published tests. In Facebook experiments reported in Marketing Science, observational methods often failed to match the randomized results, and Lewis and Rao call selection bias a crippling concern for them. Your own data may differ, so check it with a holdout.How big does a holdout test need to be?
Often bigger than expected. Lewis and Rao report that informative advertising experiments can easily require more than 10 million person-weeks. Count the days your order volume needs with a holdout calculator, and treat a gap inside the normal week-to-week swing as no answer.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Attribution ModelAn Attribution Model defines how credit for conversions is assigned to marketing touchpoints. It dictates how marketing channels receive credit for sales.
- Causal AttributionCausal Attribution uses causal inference to determine which marketing touchpoints genuinely cause conversions, not just correlate with them.
- Confidence IntervalConfidence Interval is a statistical range of values that likely contains the true value of a metric. In marketing analytics, it quantifies uncertainty around estimates, indicating the precision of an outcome or causal effect.
- CounterfactualCounterfactual is a hypothetical outcome that would have occurred if a subject had received a different treatment.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Treatment GroupTreatment Group is the set of users exposed to a specific marketing intervention. Comparing this group to a control group shows the intervention's causal impact.