Your Meta creative test is not an A/B test
A creative test inside Meta compares two ads that were shown to different people, picked by the delivery algorithm. It tells you which ad wins in that system, not whether the creative itself is better.
By Joris van Huët, Founder & CEOPublished 6 min read
Run the numbers for your store: the free break-even ROAS calculator.
A creative test run inside Meta tells you which ad won inside Meta's delivery system. It does not tell you which creative is better. The gap between those two statements decides whether your next ten ads repeat a lesson or repeat an accident.
Two recent ecommerce videos lean on the same habit. One is a funnel teardown that scrolls through a brand's live ads and sorts them by angle (video). Another explains how a marketplace tests new listings with a few impressions and then feeds whichever one starts strongest (video). Both treat the platform's choice as a verdict on the creative. It is a verdict on something narrower.
What a platform test actually compares
In a platform A/B test, each ad is delivered by an algorithm that looks for the people most likely to respond to that particular ad. So ad A and ad B end up in front of different mixes of people. Michael Braun and Eric Schwartz named this divergent delivery in the Journal of Marketing and showed what it does to the result: the comparison mixes the effect of the creative with the effect of who the algorithm chose to show it to.
A 2025 paper by researchers from Meta, Boston University and Northwestern University measured the pattern across 3,204 Lift tests and 181,890 A/B tests (arXiv). The Lift tests, which compare people eligible to see the ads with a control group that sees none, showed no meaningful audience imbalance. The A/B tests showed clear imbalance, which the authors describe as intentional. The delivery system is doing the job it was built for.
What the result is still good for
The A/B result is not worthless. It answers a narrower question honestly: under these settings, on this platform, which ad gets more results for the money. Braun and Schwartz draw the same line in the AMA's summary of the study: the test helps predict which ad will perform best in the same targeted environment, not how the creative works in general.
What a platform creative test can tell you:
- Which of two ads to keep running in this campaign, with these settings.
- Which ad the delivery system finds cheaper results for.
What it cannot tell you:
- Whether the hook, the offer or the format made the difference.
- Whether the winner will also win in email, on your product page or on another platform.
- Whether either ad sold anything that would not have sold anyway.
The last line is the one budgets depend on, and a creative test never answers it. That takes a holdout test.
The budget half of the same problem
Delivery also decides spend. When several ads share a budget, the system moves money toward whichever ad looks best early, on very little data. An ad that starts slowly gets fewer impressions, so it never collects the evidence that might have changed the system's mind. The marketplace in the second video works the same way: a few test impressions, then the early leader gets the rest.
So an ad that barely spent has not lost a test. It was never really in one.
Meta's delivery system is built to find the cheapest results for each ad, not to run a fair experiment for you. Those are different jobs, and only the first one is the platform's.
How to tell what is actually working
Start with the last creative test you ran, and your own numbers.
- Check the spend split. Export the results by ad with spend, impressions and reach. If one ad took most of the budget, you are looking at an allocation outcome, not a comparison.
- Check who each ad reached. If your ad account lets you break results down by age, placement or region, look at whether the two ads found different people. If they did, part of the winner's edge is its audience, which is selection bias under a friendlier name.
- Check what happened after the click. Give each ad its own utm_content value. GA4 reads that parameter into its manual ad content dimensions (Google's dimension reference), so you can compare conversion rate and order value per ad on your own site. A winner whose visitors behave very differently from the loser's was probably shown to different people.
What to test instead
Three designs, by the question each can answer:
- Which ad to run on this platform. Keep using the platform test, and hold the objective, targeting, budget, bidding and placements identical across the cells. The Meta co-authored paper found that configuration choices like these reduce divergent delivery without removing it, and that tests optimised for awareness showed less imbalance than tests optimised for conversions (arXiv). Read the result as a platform decision, not as a lesson about creative.
- What the creative itself does. Take the audience choice away from the algorithm. Split your email list at random and send each angle to one half, or show the two angles as the headline and hero image to randomly split site visitors. The comparison is fair because a coin flip assigned the groups, which is the logic of a randomised controlled trial, rather than a system optimising each arm.
- Whether the ads sell anything at all. Use a design with a no-ad control group: a platform lift test or a geo holdout, set up as in how to run a holdout test on Meta ads. This is the question your budget depends on, and incrementality testing vs A/B testing explains why only this design answers it.
Write the question down first
Before the next test, write one line: which of the three questions it answers. A creative test read as if it were a lift test is how a brand ends up scaling the ad the algorithm liked for its audience, and calling it a lesson about hooks.
The cure is the discipline in vary one element, learn something reusable: one change, one question, one reading. And when a winner does come through, remember the cost of scaling a false winner before you move the whole budget behind it.
Related answers
Frequently asked questions
What is divergent delivery in Meta ads?
It is the delivery algorithm showing each ad in a test to a different mix of people, because it looks for the people most likely to respond to each ad. The result then mixes the effect of the creative with the effect of the audience the algorithm picked for it.Are Meta A/B tests useless for creative decisions?
No. They are a fair guide to which ad to keep running on Meta under the same settings. They are not a guide to why the creative worked, or to how it will perform in email, on your site or on another platform.Why did one of my new ads barely spend?
When ads share a budget, delivery moves spend toward whichever ad looks best early, on little data. Low spend means the ad never got a fair comparison, not that it lost one.How do I test a creative angle fairly?
Assign the audience yourself. Split an email list at random, or randomly split site visitors between two versions of a page. To learn whether the ads sell anything at all, use a lift test or a geo holdout with a no-ad control group.
Go deeper: Causal attribution, explained.
Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.
Keep reading
Terms in this article
- AttributionAttribution identifies user actions that contribute to a desired outcome and assigns value to each. It reveals which marketing touchpoints drive conversions.
- Control GroupControl Group is a segment of an audience intentionally not exposed to a marketing campaign, used to measure the campaign's true causal impact.
- Conversion rateConversion Rate is the percentage of website visitors who complete a desired action out of the total number of visitors.
- Holdout TestA holdout test is an experiment where a portion of the audience does not see a campaign. This measures the campaign's true incremental impact.
- IncrementalityIncrementality measures the true causal impact of a marketing campaign. It quantifies the additional conversions or revenue directly from that activity.
- Incrementality TestingIncrementality Testing measures the additional impact of a marketing campaign. It compares exposed and control groups to determine causal effect.
- Product PageProduct Page is a webpage dedicated to a single product. It includes images, descriptions, pricing, and purchase options.
- Selection BiasSelection Bias occurs when data points selected for analysis do not represent the target population. This leads to distorted findings about marketing campaign impact.