Merchandising changes: measure the page, not the ad
Split visitors at random and judge the page's own numbers, because a before/after comparison mixes the change with ad and promotion changes. Watch conversion or revenue per session, add to cart rate and order value, not ROAS.
By Joris van Huët, Founder & CEOPublished 5 min read
Run the numbers for your store: the free marketing ROI calculator.
Measure a merchandising change on the page it changed, with visitors split at random, because a before/after comparison mixes the change with everything else that moved that week, your ad mix included. Shopify's documentation says a new experiment defaults to a 50% control and 50% treatment split, and its metrics for a theme experiment are conversion rate, bounce rate, reached checkout rate and add to cart rate. None of those metrics is ROAS, and that is the point: the ad platform credits the click that brought the visitor, not the page that converted them.
Is the page worth testing?
Baymard's research says the product list decides whether people find a product, but its evidence is from usability tests, not from stores. Baymard tested 19 leading sites in its product list research and reports abandonment rates of 67-90% on sites with mediocre product list usability and 17-33% on sites with a slightly optimized toolset, as published by Baymard and not independently audited. Test-lab numbers for people told to find a product are not your conversion rate. Baymard also reports that across a benchmark of 343 top-grossing US and European sites, 36% had product list flaws severe enough to harm users' ability to find and select products. Baymard's guidance on badges is narrower: if most items carry a badge, it stops working as a highlight and can have the opposite effect. No public measurement gives a typical conversion effect for collection order, bundles or badges, so this post cites none.
Why does a before/after comparison mislead?
Because time changes everything at once. A sale, a new ad set, a stock-out or a weekday mix can each move conversion, and none of them shows up on the page. Even with randomised visitors, effects drift: Microsoft's Experimentation Platform reported that, across a month of its experiments with both 7-day and 14-day scorecards, adding a second week of data increased the error bars on treatment effects 17% of the time (20 November 2024). Microsoft's Experimentation Platform also warns customers that they should usually run experiments for at least a week, to avoid weekday and weekend effects. A before/after adds every other change on top of that drift.
Shopify's documentation adds a practical rule: editing the treatment or control while an experiment is active can affect its results.
How do you test a badge, a bundle or a collection order?
Match the method to what the change touches:
- A badge or layout change: a theme edit is a documented rollout change, and an experiment compares only the changes you add to the rollout. Your store needs the Grow plan or higher.
- A bundle or quantity offer: Shopify lets you add a discount change to an experiment. Its analytics do not report the revenue or orders attributed to a discount, and because visitors can share a discount code, a code in an experiment can reach people outside the group you are testing.
- A collection order: it is not on Shopify's list of rollout changes, so you need another way to split visitors. If you have none, fall back to a before/after comparison on the same weekdays with no ad or promotion changes, write down everything else that moved, and treat the result as a hypothesis to re-test.
Which numbers do you watch?
Write these down before you start, and read them once at the end date:
- The primary metric: conversion rate or revenue per session for the sessions that saw the tested page, not platform-reported conversions.
- A diagnostic: in GA4, the share of list views that become a selected item. Google's developer guide has a store send a
view_item_listevent when a visitor is shown a list and aselect_itemevent when they pick an item from it, so check that your setup sends both. - A guardrail: average order value, which Shopify lists for launches but not for theme experiments, so check it yourself before you call a winner on conversion.
Pass: the visitor split matches the ratio you set (how to check it), the primary metric moved by more than the noise you computed from the sample size, and the guardrail did not fall. Fail: the split is off, nothing else was held still, or the only movement is in a metric you did not name in advance.
Sources, 30 September 2026: Product Lists UX research and What is a UI Badge? (Baymard Institute, 2026); Types of rollouts and changes, Rollout analytics and Managing rollouts (Shopify Help Center, 2026); External Validity of Online Experiments (Microsoft Research, Experimentation Platform, 2024); Measure ecommerce (Google Analytics developer guide, 2026).
Related answers
Frequently asked questions
How do I test a change to my collection page on Shopify?
Use a Shopify experiment for a theme edit: Shopify says a new experiment defaults to a 50% control and 50% treatment split. A collection's sort order is not on Shopify's list of rollout changes, so split visitors another way or fall back to a before/after comparison and treat it as a hypothesis.Should I use ROAS to judge a merchandising change?
No. ROAS credits the ad click that brought the visitor, not the page that converted them. Judge the change on conversion rate or revenue per session for visitors who saw the page, with average order value as a guardrail.How long should a merchandising test run?
Until it reaches the sample size you calculated in advance, in whole weeks. Microsoft's Experimentation Platform tells customers to usually run experiments for at least a week, to avoid weekday and weekend effects.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Conversion rateConversion Rate is the percentage of website visitors who complete a desired action out of the total number of visitors.
- ExperimentationExperimentation in marketing conducts controlled tests to determine the causal impact of specific actions. This includes A/B testing and other controlled experiments to establish causality.
- ExperimentsExperiments are scientific procedures that test hypotheses or demonstrate facts. In marketing, experiments like A/B tests determine the causal effect of campaign changes, enabling data-driven decisions.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Split TestingSplit Testing is a method of running controlled, randomized experiments. It tests website changes to improve conversion rates.
- Treatment EffectTreatment Effect is the causal impact of an intervention on an outcome. In marketing, this means the change in a metric like conversion rate directly caused by a campaign or pricing adjustment.