A/B testing your store: sample size comes first
A test reaches a verdict only if your traffic can reach its sample size. Do that sum first, fix the end date, read the result once, and check the visitor split before you trust it.
By Joris van Huët, Founder & CEOPublished 5 min read
Run the numbers for your store: the free customer journey credit calculator.
An A/B test reaches a verdict only if your traffic can reach its sample size, so do that sum before you build the variant. Evan Miller's rule of thumb for the sample size is n = 16σ²/δ², where δ is the smallest change you want to detect and σ² is the variance you expect. For illustration: a 2% conversion rate has a variance of 0.02 × 0.98, and detecting a lift to 2.2% (δ = 0.002) takes 16 × 0.0196 ÷ 0.002² = 78,400 visitors per variation, 156,800 in all: about 16 weeks at 10,000 visitors a week.
How many visitors does a test need?
Enough for the smallest change worth acting on to show up. For a conversion rate p, the variance in Miller's formula is p times one minus p:
n = 16 × σ² ÷ δ² σ² = p × (1 − p)
Because δ is squared, a change twice as large needs a quarter of the visitors, so a bolder variant is the cheapest way to get a verdict. If the sum says months, test a bigger change, test where there is more traffic, or use a method that needs no test.
An A/B test compares two versions you already plan to run. Whether the activity works at all is a holdout question: see what each design proves.
Shopify's own defaults show why the sum matters. Shopify's documentation says a new experiment defaults to a 50% control and 50% treatment split and ends 90 days after it starts, and you can change both before you save. For illustration: 90 days is about 13 weeks, so that default end date would cut the 16-week example short.
Why not stop when the dashboard says "significant"?
Because checking repeatedly and stopping at the first significant reading breaks the test. In Miller's worked example, a calculation rather than a field measurement, with a 50% conversion rate, a test stopped at the first 5% significance or after 150 observations and a change that does nothing, the test wrongly finds a significant result 26.1% of the time. In his table, ten peeks turn what a dashboard reports as 1.0% significance into an actual 5%. His advice is to decide the sample size in advance and wait until the experiment is over.
How do you check a test before you trust it?
Five steps, in your test tool or a spreadsheet:
- Write the plan before launch: baseline conversion rate, the smallest lift worth acting on, the sample size from the formula, and an end date on a whole-week boundary. Pass: your traffic reaches the sample by that date. Fail: it does not, so the test can only detect a larger change than you planned for.
- Check the split, not the winner. If you split traffic 50/50 and count a and b visitors, compute (a − b)² ÷ (a + b). NIST's table lists 10.828 as the value a chi-square statistic with one degree of freedom exceeds 0.1% of the time by chance, a conservative cut-off for a split check. Pass: your number is below it. Fail: above it, which is a sample ratio mismatch (SRM). For illustration: 5,200 against 4,800 visitors gives 400² ÷ 10,000 = 16, a fail; 5,050 against 4,950 gives 1, a pass.
- Treat a failed split as a broken test. Microsoft's Experimentation Platform says not to trust the results of a test with an SRM until you diagnose the root cause, and describes one case where a bot-detection filter removed the most engaged visitors from one variation and the result flipped once it was fixed.
- Run whole weeks. Microsoft's Experimentation Platform warns customers that they should usually run experiments for at least a week, to avoid weekday and weekend effects.
- Read the result once, at the end date, next to a guardrail. Shopify lists conversion rate, bounce rate, reached checkout rate and add to cart rate for a theme experiment, and average order value for a launch. If your tool shows only conversion, check order value yourself before you pick a winner.
How common is a broken split? A Microsoft Experimentation Platform article of 14 September 2020, published by Microsoft and not independently audited, says a recent analysis showed that about 6% of Microsoft's A/B tests have an SRM, and that about 10% of LinkedIn's zoomed-in tests used to.
What can a finished test not tell you?
Why the variant won, and whether it will keep winning. Microsoft's Experimentation Platform looked at a month of its experiments with both 7-day and 14-day scorecards and reported that the 14-day result fell outside the 7-day 95% confidence interval 9% of the time, where about 5% would be expected if the 14-day result were the true effect (20 November 2024). A result is a measurement of one period, so write down the dates next to it.
Sources, 30 September 2026: How Not To Run an A/B Test (Evan Miller, 2010); Diagnosing Sample Ratio Mismatch in A/B Testing (Microsoft Research, Experimentation Platform, 2020); External Validity of Online Experiments (Microsoft Research, Experimentation Platform, 2024); Types of rollouts and changes and Rollout analytics (Shopify Help Center, 2026); Critical Values of the Chi-Square Distribution (NIST/SEMATECH e-Handbook of Statistical Methods).
Related answers
Frequently asked questions
How many visitors do I need for an A/B test?
Evan Miller's rule of thumb for the sample size is n = 16σ²/δ², where δ is the smallest change you want to detect and σ² is p × (1 − p) for a conversion rate. A change twice as large needs a quarter of the visitors, so halve the lift you want and you need four times as many.Is it OK to stop an A/B test early when it looks significant?
No. Stopping at the first significant reading inflates false positives. In Evan Miller's example, a change that does nothing is wrongly called significant 26.1% of the time. Fix the sample size and the end date first, then read the result once.What is a sample ratio mismatch in an A/B test?
It means the visitors split differently from the ratio you configured. Microsoft's Experimentation Platform says not to trust the results of a test with one until you diagnose the root cause, because the missing visitors are often the ones the change affected most.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- Confidence IntervalConfidence Interval is a statistical range of values that likely contains the true value of a metric. In marketing analytics, it quantifies uncertainty around estimates, indicating the precision of an outcome or causal effect.
- ConversionConversion is a specific, desired action a user takes in response to a marketing message, such as a purchase or a sign-up.
- Conversion rateConversion Rate is the percentage of website visitors who complete a desired action out of the total number of visitors.
- ExperimentationExperimentation in marketing conducts controlled tests to determine the causal impact of specific actions. This includes A/B testing and other controlled experiments to establish causality.
- ExperimentsExperiments are scientific procedures that test hypotheses or demonstrate facts. In marketing, experiments like A/B tests determine the causal effect of campaign changes, enabling data-driven decisions.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Statistical SignificanceStatistical Significance measures the probability that observed results are not due to random chance. It confirms the reliability of test outcomes.