Skip to content

Media mix modeling providers: what to ask before you buy

Public information cannot rank media mix modeling providers by accuracy: real sales data has no known answer to score a model against. Ask five questions instead, then test each provider on an experiment you hold back.

By , Founder & CEOPublished 6 min read

Run the numbers for your store: the free multi-channel budget calculator.

Public information cannot rank media mix modeling providers by accuracy, because real sales data has no known answer to score a model against. A 2026 preprint says MMMs "are rarely validated against ground truth, because ground truth is unobservable in real data", and Google's Meridian documentation says "there is no absolute best solution to MMM". What you can do is put five questions to each provider: how much history it needs, how often it refreshes, how it validates, whether experiments calibrate it, and what range it reports.

Why can't public information rank MMM providers by accuracy?

No ground truth. On real data nobody knows a channel's true effect, so a model cannot be scored against it. The 2026 preprint (Heusch, arXiv; one author, synthetic data) presents a public synthetic dataset with known answers for that reason. Google's documentation adds that fit does not settle it: "Multiple models can have good fit and predictive power yet provide different ROI and optimization results", and "A model with 99% out-of-sample R-squared can still be a poor model for causal inference." A fit statistic on a sales page is not an accuracy ranking.

Observational estimates and experiments disagree. In 15 U.S. advertising experiments at Facebook, observational methods often failed to reproduce the randomized result even with extensive data (Gordon et al., Marketing Science, 2019, peer-reviewed; two of the four authors worked at Facebook). Those were user-level methods, not marketing mix models, and no source cited here tests a provider's MMM against a holdout.

Prices are not like for like. Take four MMM sellers from the comparison table. Recast, Prescient AI and Haus list no public price (the last two as read on 2026-09-26), and Fospha lists Lite at $1,500/mo for $100k to $500k in monthly media spend (as read on 2026-09-08).

What should a provider say about history and refresh?

History. Meridian's documentation gives a rule of thumb of a minimum of two years of weekly data for geo-level models and three years for national-level ones, and calls its guidance "rough and directional". Robyn's analyst's guide says a minimum of two years of weekly data, with 1 independent variable per 10 observations. Ask: how many parameters will the model estimate on my data, and how many weeks per parameter do I have? Pass: a number for your data. Fail: "enough history" with no count. The software guide has the arithmetic.

Refresh. Meridian's documentation says refreshes can be as frequent as you like, suggests quarterly, annually or at the pace of your budget decisions, and warns that "MMM estimates often exhibit high variance", so a small amount of new data can move them. Robyn's guide says refreshing is not always the best approach and a rebuild is sometimes better. Ask: what changes between refreshes, and how do you stop last quarter's advice flipping? Pass: a stated rule. Fail: a new answer each run with no account of why.

How should a provider validate and calibrate?

Validation. Meridian's documentation says the main use of a holdout sample is out-of-sample fit, and that there is "no guarantee that the model with the best out-of-sample model fit is the best model for causal inference". Ask: besides fit, what do you check against something you did not fit? Pass: an experiment or a known-answer test. Fail: fit statistics only.

Calibration. Robyn's guide strongly recommends calibrating with experimental and causal results treated as ground truth, and its Key Features page cites a third-party whitepaper finding that uncalibrated models show a 25% average difference to the ground truth (published by Meta, not audited here). Meridian's documentation says incrementality experiments are "perhaps the strongest basis for formulating your intuition", and that turning them into priors "isn't a precise formula". Ask: which experiments calibrated my model, how old are they, and what do my estimates become without them? Pass: a before-and-after comparison. Fail: no experiments, or no answer.

How do you test providers on your own data?

Hold back one result and let each provider predict it:

  1. Pick one experiment you trust. A holdout or lift test on one channel, with its dates and its measured effect on revenue. With none, run one first (incrementality testing).
  2. Keep it out of calibration. Give each provider the same weekly revenue and spend history, including the test weeks, but not the experiment's result. A model calibrated on a test has already seen the answer.
  3. Ask for the channel's incremental revenue over the test dates, as a range. Meridian's documentation says the most accurate way to assess how much data you need is to run the model and evaluate the width of the credible intervals, so a range is a fair request.
  4. Compare. Pass: the range overlaps the experiment's and sits on one side of break-even, so it tells you to cut or raise. Fail: it misses the experiment's range, or it spans break-even.

For illustration: a test measured 120,000 of incremental revenue, with a range of 80,000 to 160,000, and break-even is 70,000; provider A says 50,000 to 400,000, which overlaps but spans break-even, and provider B says 90,000 to 140,000, which overlaps and sits above it. Experiments are noisy too: across 25 large field experiments, the median confidence interval on return on investment was over 100 percentage points wide (Lewis and Rao, Quarterly Journal of Economics, 2015, peer-reviewed), so judge overlap, not exact match.

Sources, 30 September 2026: A Synthetic Benchmark Dataset with Endogenous Marketing Spend (Heusch, arXiv preprint, 2026); About MMM as a causal inference methodology, Amount of data needed, Collect and organize your data, Refresh the model, Holdout observations and Calibrate treatment priors (Meridian documentation, Google for Developers, 2026); An Analyst's Guide to MMM and Key Features (Robyn documentation, Meta Marketing Science, updated December 2024); A Comparison of Approaches to Advertising Measurement (Gordon et al., Marketing Science, 2019); The Unfavorable Economics of Measuring the Returns to Advertising (Lewis and Rao, Quarterly Journal of Economics, 2015); Recast, Prescient AI pricing, Haus pricing and Fospha pricing (vendor pages).

Frequently asked questions

  • Which media mix modeling company is best?
    No public data can rank them by accuracy. Real sales have no known answer to score a model against, prices are quotes or tiers, and Google's documentation says there is no absolute best solution to MMM. Compare providers on your own data, using one experiment you hold back.
  • How much history does an MMM provider need?
    Google's Meridian documentation gives a rule of thumb of a minimum of two years of weekly data for geo-level models and three years for national-level models; Robyn's guide says a minimum of two years. Ask any provider to state the figure for your data.
  • Should an MMM provider calibrate with experiments?
    Where you have them, yes. Meta's Robyn documentation strongly recommends experimental results as the ground truth, and Google's Meridian documentation calls incrementality experiments perhaps the strongest basis for a channel prior. Ask which experiments calibrated your model and what changes without them.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.