Skip to content

AI tools for marketing analytics: where numbers come from

An AI answer's number comes from a query the tool ran, the model's own arithmetic or an estimate, and only a query can be re-run. Use AI to ask and to draft, and take budget numbers from something you can repeat.

By , Founder & CEOPublished 6 min read

A number in an AI analytics answer comes from one of three places: a query the tool ran, the model's own arithmetic, or an estimate. Only the first leaves something you can re-run. Google's help page says Ask Advisor in Google Analytics "may have some capability limitations and make mistakes", and Google's BigQuery documentation says Gemini for Google Cloud products can "generate output that seems plausible but is factually incorrect". Use AI to ask and to draft, and take the number that moves budget from a report, query or test you can repeat.

Which AI features already sit inside your analytics tools?

Five are documented by their vendors, as the pages read today:

ToolWhat the vendor's page saysWhat to check
Ask Advisor in Google Analytics (Beta)an "agentic conversational experience"; answers draw only on the property's own data; supports "Why" questionswhether a figure came from a report query or the model's own calculation
AI overviews in Google Analyticsa concise "AI-generated overview of the most critical changes to your data since your last visit", also in detail reportsthe report behind each change
Google Analytics MCP server (Experimental)gives an LLM tools such as run_report, which "Runs a Google Analytics report using the Data API"the dimensions, filters and dates the LLM chose
Gemini in BigQuerygenerates SQL from a prompt, and the same prompt can return different syntaxthe SQL itself: run it yourself
Shopify Sidekick"an AI-enabled commerce assistant"; marketing data by UTM or referrer, "but not attribution model comparisons or discount code performance"the report behind each figure

AI does help where the output is a draft you can read: a query you then check, a report you already trust explained in plain words, finished results summarised. None of these tools changes where the number has to come from.

What do benchmarks and experiments show about AI on marketing questions?

Benchmarks score SQL and agents, not your export. AD-Bench (arXiv preprint by Tencent authors, revised 22 June 2026, not peer-reviewed) built 225 test requests from real marketing-analysis requests on a production advertising platform. The best model tested answered 76.9% correctly on the first try overall and 61.4% on the hardest level, graded by an LLM judge that was also one of the models tested, a limitation the authors list. On the Spider 2.0 leaderboard, as read today (a benchmark project's page, with entries submitted by named teams), the top Snowflake-track entry scores 96.70 on 547 examples that include "well-prepared database metadata and documentation". Nothing on that page says how those systems score on your own export.

"Why" questions come in two kinds. Google's page on AI overviews says its AI "works through countless combinations of dimensions and metrics to proactively connect the dots, explaining why these spikes happened", and Ask Advisor lists "Why did my total revenue drop on August 1st?" among its questions. Some of those can be checked against independent records: a tracking break, a stock-out, an outage. "Which channel caused these orders?" is the harder kind. Gordon and co-authors (Marketing Science, 2019, peer-reviewed) compared 15 US advertising experiments at Facebook, with 500 million user-experiment observations and 1.6 billion ad impressions, against observational estimates. They report that "observational methods often fail to produce the same effects as the randomized experiments". A tool that only sees your property's data can describe what changed. Until a holdout backs it, credit for a cause is a hypothesis.

How do you check an AI answer yourself?

Take one answer that matters and run five checks:

  1. Sort each number by source. Ask the tool what produced it: a report, a query, or its own calculation. A number with no source is unverified.
  2. Rebuild it. Open the same report with the same dates, time zone, dimension and filter, or run the SQL yourself and reconcile orders to Shopify for the same days.
  3. Match the definitions. A rate per user and a rate per session are different numbers, and so are purchases and key events.
  4. Ask again later. BigQuery's page says the same prompt can return different syntax, so repeat the question on another day and after an update.
  5. Treat "why" as a hypothesis. Check it against independent records (a campaign change log, stock, outages) and, for channel credit, a holdout.

Pass: the same number, or a gap you can name. Fail: no source, an unexplained gap, or a different answer to the same question. Vendors grade their own assistants, so the accuracy figure that matters for your store is the one you get this way. Also ask where your data goes: Google says Ask Advisor data, "including prompts and generated answers, is not used for training, fine-tuning, or evaluation of AI models", and that chat logs are kept for 55 days. Ask any other tool for the same two facts before you paste order data into it.

Where does a causal read fit?

Later, once the question is which channel caused orders rather than what changed, a causal attribution tool like Causality Engine shows what each channel caused from a GA4 export, and its Pro plan includes AI chat across your historical attribution data.

Sources, 30 September 2026: Ask Advisor in Google Analytics (Beta) (Google Analytics Help, 2026); How Ask Advisor uses data (Google Analytics Help, 2026); AI overviews (Google Analytics Help, 2026); Google Analytics MCP server (Google, 2026); Write SQL queries with Gemini (Google Cloud, 2026); Sidekick (Shopify Help Center, 2026); Marketing reports (Shopify Help Center, 2026); AD-Bench (Hu et al., arXiv, 2026); AD-Bench full text (arXiv, 2026); Spider 2.0 leaderboard (XLANG Lab, read 30 September 2026); A Comparison of Approaches to Advertising Measurement (Gordon et al., Marketing Science, 2019).

Frequently asked questions

  • Can AI analyze my GA4 data accurately?
    It can report what a query returns, and Google's own page says Ask Advisor may make mistakes. Re-run any number you act on in the GA4 report or in SQL. A "why" answer stays a hypothesis until a holdout or another independent record backs it.
  • What is the best AI tool for ecommerce marketing analytics?
    None of the sources cited here ranks them on your data. The vendors' pages describe different jobs: Ask Advisor answers questions inside GA4, the Google Analytics MCP server lets an LLM run reports, Gemini in BigQuery drafts SQL, and Sidekick answers questions in Shopify admin. Pick by the job, then check each number.
  • Do AI analytics tools train on my data?
    Check each vendor's page. Google says Ask Advisor data, including prompts and generated answers, is not used for training, fine-tuning or evaluation of AI models, and that chat logs are kept for 55 days. Ask any other tool for the same two facts.

Go deeper: Causal attribution, explained.

Sixty-second versions of these ideas: Causality Engine on YouTube Shorts.

Keep reading

Terms in this article

Browse the full glossary

Your platforms guess.
We run the math.

Upload a GA4 export and see what each channel caused, next to last-click, in 1–2 minutes. The read is yours to keep.

Free, in your browser: your file is not uploaded. The full read is €99, refundable within 30 days. Prices exclude VAT.
Or book a 30-min call.