Skip to content

Method, applied

Free gift vs discount: what a holdout test would actually show

In an independent study of $285.3 million in gross sales across 10 Shopify stores, ethercycle found orders carrying a free gift averaged about $173, against roughly $116 for orders using a discount code and $112 at full price. It is a striking gap, and the obvious conclusion is that gifts make people spend more. That conclusion does not follow from the data, for a reason ethercycle names themselves: gift offers are threshold-gated, so the bigger carts qualify by construction. Their report closes the section by saying nobody has run the holdout test. This is what that test would look like, and why it is the only thing that settles the question.

Why the comparison is confounded by construction

A gift-with-purchase promotion nearly always reads spend $75, get a free item. Every cart in the gift group has therefore already crossed $75. Carts below the threshold cannot appear in that group at all, no matter how the shopper behaved. So the gift group is not a sample of customers, it is a sample of large carts, and comparing it to everyone else guarantees a gap even in a world where the gift changed nobody's mind. ethercycle put it plainly in their own words: part of the basket gap is who qualifies, not what the gift causes.

This is selection, and it is the same structural error that makes platform-reported ROAS flattering: the platform counts conversions from people it selected for being likely to convert. The mechanism is identical whether the selector is an ad algorithm or a $75 threshold. Once you see the shape, you find it everywhere in ecommerce reporting, which is precisely the point.

The finding that survives it

Not everything in the report is threshold-bound, and the strongest result is the one that gets least attention. ethercycle report that gift orders skewed heavily toward new customers, between 46.5% and 54.2%, against 30.2% for price-discounted orders. A spend threshold explains why a qualifying cart is large. It does not explain who walks through the door. That difference in customer mix is a genuinely interesting signal and it is not obviously an artefact of the gating.

The margin argument stands on its own too: across the dataset, 37.4% of all orders carried a price cut, and at one store a single evergreen code left live for twelve months gave back $349,000. You do not need a causal claim to act on that. Money handed to people who would have bought anyway is the cleanest definition of marketing debt there is.

How to run the test properly

Three design choices decide whether the answer means anything.

  1. Randomise before the offer, at the visitor level. Split incoming traffic into two arms when they arrive, not when they reach a cart value. A holdout assigned after the threshold is crossed measures nothing, because assignment is then caused by the very behaviour you are trying to explain.
  2. Measure revenue per visitor, not per order. Average order value is the wrong unit here, because the threshold acts directly on it. Revenue per visitor counts everyone in the arm, including the people who bought nothing, so a promotion that lifts a few baskets while suppressing overall purchase rate cannot hide behind a flattering AOV.
  3. Size it before you start, and report the interval. Decide the smallest lift worth acting on and run until the arms can detect it. A lift test that stops when the numbers look good measures patience, not promotions. Report the range, not a single number.

Where a live holdout is not practical, because the promotion is already running everywhere or the traffic is too thin to split, the same question can be approached after the fact with incrementality methods on historical order data, comparing periods and stores that did and did not run the offer while controlling for what else changed. It is weaker evidence than a randomised arm, and it should be labelled as such.

Why this keeps happening

A promotion, a channel, and an ad platform all produce the same kind of report: here is a group that saw the thing, here is what they spent. The number is always real. What is missing is the group that did not see it, and without that comparison the report cannot tell you what the thing caused. Ecommerce runs on the first kind of number and budgets on the second.

Worth saying clearly: ethercycle handled this better than most. They stated the selection caveat under their own chart, labelled the result correlation rather than causation, disclosed that they sell a gift-with-purchase app, and asked for the experiment instead of claiming the win. That is what good practice looks like when the data flatters your own product. The open question they left is a real one, and it is answerable.

Run it on your own store

We measure what marketing actually caused, from a GA4 export, with no pixel and nothing installed. The method: causal attribution and incrementality testing, at €99 one-time or €299/mo.

FAQ

Does a free gift with purchase actually increase average order value?

Nobody has proven it either way with an experiment, and the honest answer today is that we do not know. ethercycle's study of $285.3 million in gross sales found gift orders averaged about $173 against $116 for orders using a discount code, but gift offers are almost always threshold-gated: spend a minimum, get the gift. That means larger carts qualify for the gift by construction, so part of the gap is who qualifies rather than what the gift caused. ethercycle states this caveat themselves. Only a holdout test separates the two.

What is threshold-gating and why does it break the comparison?

A gift-with-purchase offer usually reads spend $75, get a free item. Any cart that receives the gift has therefore already crossed $75, while carts below the threshold never appear in the gift group at all. Comparing the two groups after the fact is comparing big carts to all carts, which guarantees a gap even if the gift changed nobody's behaviour. This is selection bias, and it is the same structural error that makes platform-reported ROAS look better than it is.

How would you actually test whether the gift causes the lift?

Randomise at the visitor level, not the cart level. Split incoming traffic into two groups before anyone sees an offer: one sees the gift promotion, the other sees the store without it. Run until each arm has enough orders to detect the effect size you care about, then compare total revenue per visitor across the whole arm, including everyone who bought nothing. Comparing revenue per visitor rather than per order is what keeps the threshold from re-entering through the back door.