On-site search and recommendations: how to measure the lift
Hold a search or recommendation block back from a random share of visitors and compare orders per visitor. A widget report counts orders that passed through the block, including ones that would have happened anyway.
By Joris van Huët, Founder & CEOPublished 5 min read
Run the numbers for your store: the free holdout test planner.
Hold a search or recommendation block back from a random share of visitors and compare orders per visitor, because clicks on the block overstate what it adds. In a randomised field experiment at an online book retailer, personalised recommendations raised shoppers' propensity to buy by 12.4% and basket value by 1.7% (Li, Grahl and Hinz, 2022). In one site's browsing logs for 2.1 million users, at least 75% of the activity from recommendation click-throughs would likely have happened without the recommendations (Sharma, Hofman and Watts, 2015). A widget report counts what passed through the block. Only a holdout shows what the block added.
What does a widget report count?
Shopify's Search & Discovery reports give recommendations a purchase rate: the number of customers who purchased a product that they discovered in recommendations (Shopify Help Center, official documentation). The search reports do the same for search results. That is attribution: the order passed through the block. It says nothing about whether the shopper would have bought anyway, and the observational estimate above puts most recommendation click-throughs in that group.
The same trap sits in any comparison of shoppers who used search with shoppers who didn't. In Baymard's usability testing, roughly half of participants turned to search as their preferred way to find products, and the rest used the main navigation (Baymard Institute, 2026). People choose to search, so that comparison sets two kinds of shopper against each other, not search against no search.
What do randomised tests show about recommendations and search?
Three results from peer-reviewed field experiments, each at one retailer or platform, so read the direction and not the size:
- They can add orders. In the book-retailer experiment above, random assignment to recommendations raised the chance of buying and basket value, so the effect showed up in orders and not only in clicks.
- They move sales between products. On a fashion retailer's site, conditional on a product's page view, its own sales decreased by 1.9% and the sales of its recommended substitutes increased by 9%. On average, a product's recommendation links produced an 11% gain in total sales of the product and its recommended substitutes (Kumar and Hosanagar, 2019). Count orders across the set, not one product's sales.
- They trade use with search. In a randomised experiment with 555,800 customers on a large ecommerce platform, less relevant home-page recommendations significantly increased use of the search channel, which the authors read as a partial substitution effect (Yuan and colleagues, 2025). Change one block and the other moves, so read total orders. For search on its own, no source read here reports a randomised test of a search change on orders, so treat search as the case you test, not one the literature settles.
How do you hold a block back on your own store?
Run it on your own traffic:
- Pick one block: the recommendation row on product pages, or the search suggestions. One change per test.
- Assign visitors, not page views. Give each visitor a group on the first visit and keep it, so nobody sees both versions.
- Label the group in GA4 before launch. Send it as an event parameter and register it as a custom dimension, because Google says reporting on a custom dimension starts 24 to 48 hours after you create it (Google Analytics Help).
- Judge on orders per visitor and revenue per visitor for everyone in each group over the same days. Keep clicks on the block as a diagnostic only.
- Check the split before you read the result. If you planned an even split and the groups differ by more than chance allows, something assigned visitors unevenly. Microsoft researchers warn that ignoring a sample ratio mismatch can make a bad change look good, or the reverse (Fabijan and colleagues, 2019).
- Run whole weeks and fix the end date first. The holdout test planner turns your daily orders, the smallest lift worth detecting and the share held out into a number of days.
What the result means:
- Pass: orders per visitor are higher where the block shows, by more than chance allows at your traffic, and average order value and refunds are not worse.
- Fail: clicks on the block rise and orders per visitor stay level. The block rerouted shoppers and added no orders.
- Too small to read: the gap sits inside the noise. Extend the test, or keep the answer as a direction and say so.
Sources, 30 September 2026: How Do Recommender Systems Lead to Consumer Purchases? A Causal Mediation Analysis of a Field Experiment (Information Systems Research, 2022); Estimating the Causal Impact of Recommendation Systems from Observational Data (ACM Conference on Economics and Computation, 2015); Measuring the Value of Recommendation Links on Product Demand (Information Systems Research, 2019); How Recommendation Affects Customer Search: A Field Experiment (Information Systems Research, 2025); Shopify Search & Discovery reports and analytics (Shopify Help Center, 2026); Ecommerce Search UX 2026: 8 Search Query Types UX Best Practices (Baymard Institute, updated 2026); About custom dimensions and metrics (Google Analytics Help, 2026); Diagnosing Sample Ratio Mismatch in Online Controlled Experiments (Microsoft Research, 2019). The sizes above come from single retailers and one platform; the experiments are randomised, the browsing-log estimate is observational.
Related answers
Frequently asked questions
Do product recommendations increase sales or just move them?
Both happen. In a randomised test at a fashion retailer, a product's own sales fell per page view while its recommended substitutes sold more, for a net gain across the set. Count orders across the store, not one widget's sales.How do I know if my on-site search is worth improving?
Test it: hold the change back from a random share of visitors and compare orders per visitor. Shopify's Search & Discovery reports list searches with no results and searches with no clicks, which shows where to look first.How long should a recommendation holdout test run?
Long enough to reach the sample your daily orders and smallest worthwhile lift require, in whole weeks, with the end date fixed in advance. The holdout test planner turns those inputs into a number of days.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Ecommerce PlatformEcommerce Platform is software that allows businesses to sell products online. It manages inventory, payments, and customer relationships.
- Google AnalyticsGoogle Analytics is a web analytics service that tracks and reports website traffic.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Mediation AnalysisMediation analysis is a statistical method that explains how a treatment affects an outcome. It separates direct effects from indirect effects through a mediator variable.
- Product PagesProduct Pages are individual pages on an e-commerce site that describe a specific product. They provide detailed information to shoppers.
- Usability TestingUsability Testing evaluates a product by observing real users interacting with it. It provides direct feedback on how people use a system.