Skip to content

AI product images: measure returns, not clicks

An AI product image is judged twice, by the click now and by the return later. Tie every order to the image that was live when it was placed, and read return rate and refund reasons at a fixed horizon before you call it a win.

By , Founder & CEOPublished 5 min read

Run the numbers for your store: the free return rate calculator.

An AI product image gets judged twice: by the click this week and by the return next month. Most teams read the first verdict and roll the image out before the second one arrives. An image that wins on conversion rate can still lose on margin, and nothing in an ad dashboard will warn you.

AI imagery runs through this week's ecommerce YouTube. One creator walks through designing print-on-demand products with image generators (the video). Another's Depop guide argues that realistic, phone-style AI mockups sell better than supplier photos (the video). Both treat the image as the thing that sells. The measurement question is what happens after the sale.

Why the click and the return disagree

A product image sets an expectation. An image that flatters the product can raise click-through and conversion rate precisely because it promises more than the box delivers. The gap shows up later, in returns, refund requests, complaints that the item is not as pictured, weaker reviews and fewer second orders.

The reports that grade the image only see the first half. An ad platform counts the purchase when it happens. GA4 counts it too, and a refund only reaches GA4 if you send a separate refund event with the order's transaction ID (Google for Developers). By the time the returns arrive, the image has been declared a winner and copied across the catalog.

MetricWhen it arrivesWhat it rewards
Click-through rateHoursThe promise
Conversion rateDaysThe promise
Return rate and refund reasonsWeeksThe product matching the promise
Reviews and repeat purchaseMonthsThe product matching the promise

Five numbers per image version

Split every order by the image that was live when it was placed, then read these for each version:

  1. Conversion rate on the product page.
  2. Return rate within your full return window.
  3. Refund reasons, especially not as described, wrong color and quality.
  4. Review rating for orders placed under each version.
  5. Repeat purchase rate at a horizon you pick before you look.

Then put them together. Net revenue per visitor is revenue minus refunds, divided by visitors, for each version. If the new image wins on conversion rate and loses on net revenue per visitor, it lost. The same logic applies to your ad numbers: returns-adjusted ROAS is the version of the metric that waits for the second verdict.

Mapping orders to versions is easier than it sounds. Write down the date each product's images changed. Every order before that date belongs to the old version, every order after it to the new one, and your returns system already knows which orders came back.

Who grades the image

The tools that generate images and the platforms that run your ads are graded on clicks and purchases. Neither of them sees your returns desk. An image generator's showcase shows the image, not the return rate, and an ad platform has no reason to wait for returns before it reports a winner. Their incentive is the first verdict. Yours is the second.

None of that makes AI images a bad idea. Some will beat your photos on every line, returns included. The point is to find out which, product by product, before the catalog changes.

What to test

  • Split, do not switch. If your store can split traffic on a product page, show real photos to half the visitors and AI images to the other half, with the same copy and price. Change only the image, so the result tells you something about images; see vary one element, learn something reusable.
  • Or split by product. Move a random half of comparable products to AI images and keep the other half on real photos as the control. That is a split test at the catalog level.
  • Run until the return window closes for the last order in the test, not until conversion rate looks significant.
  • Decide the rule before you start: the metric (net revenue per visitor after refunds), the horizon, and the result that sends you back to real photos.
  • Watch the ads too. If AI images run in your ads while real photos sit on the product page, the gap moves to the landing: the ad promises one thing and the page shows another. Compare conversion from those ads with conversion from ads that match the page.

If the AI version wins on every line, including returns, you have a real winner. Scaling one before the return window closes is how you end up paying for a false winner across the whole catalog. Keep the dated record of every test either way; it is worth more than the images, as own the test results, not just the winners argues.

If you already switched

Plenty of stores changed their images months ago and never tested. You can still read it after the fact. Take the products that switched, compare their return rate and refund reasons in the months before and after the switch date, and do the same for products that did not switch over the same months. If only the switched products moved, the images are the likely cause. If everything moved, look at the season, your carrier or your sizing first. It is weaker evidence than a split, and much better than none.

The same caution applies to anything that rebuilds your pages at once. An AI rebuilt page breaks channel attribution for the same reason: when everything changes on one date, nothing can be credited to anything.

Frequently asked questions

  • Do AI product images increase returns?
    They can, when the image promises more than the product delivers. Whether they do in your store is measurable: compare return rate and refund reasons for orders placed under each image version, over your full return window.
  • How long should an image test run?
    Until the return window has closed for the last order in the test, not until conversion rate looks significant. The click verdict arrives in days, and the return verdict arrives weeks later.
  • What metric should decide an image test?
    Net revenue per visitor after refunds, for each image version, with review rating and repeat purchases as tie-breakers. Conversion rate alone rewards an image for promising too much.
  • How do I see refunds in GA4?
    GA4 only shows refunds you send to it, through a refund event carrying the order's transaction ID. Without that event, GA4 keeps counting the original purchase as if it stood.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.