AI Search Measurement and Reporting

How to measure AI search — attribution, LLM visibility, share of voice, hallucination monitoring, and self-reporting — and why your analytics show a floor, not a ceiling.

First published: Jun 24, 2026 · Last updated: Jul 22, 2026 · Advanced
demand #4 in Measurement and Reporting#25 in AI Search#324 on the site

AI search is growing fast (9.7x in 12 months), converting well (up to 23x better than organic), and almost invisible in your current analytics — 35–70% of AI visits arrive with no referrer and land in Direct, and GSC still doesn't break out AI Overview clicks. So you stack layers instead of relying on one number: direct attribution (GA4 + Ahrefs), LLM visibility (GSC Gen AI reports, Bing citations), share of voice (Brand Radar and friends), hallucination monitoring, and self-report. The data you have is a floor, not a ceiling — and measuring AI visibility is mostly measuring brand health.

TL;DR — AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. attribution remains incomplete: some visits arrive without usable referrer information and GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. includes AI-feature activity within Web search reporting rather than exposing a separate AI filter. Because no single tool sees it all, you stack layers — direct attribution, LLM visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals., share of voice, hallucinationAn AI hallucination is when a large language model generates output that is confidently stated but factually wrong, made up, or unsupported by its source. It's a side effect of next-token prediction — not a bug that can be fully eliminated. monitoring, and self-report (plus incrementality when you have the volume). Treat what you can measure as a floor. And remember the punchline: the signals that drive LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). visibility can overlap with broader brand signals, so interpret AI-search metrics alongside brand health rather than as an isolated channel.

Evidence for this claim Google reports clicks and impressions from AI Overviews and AI Mode within the Search Console Performance report's Web search type rather than as separate filters. Scope: Current Google Search Console reporting for Google AI search features. Confidence: high · Verified: Google Search Central: AI features and your website Evidence for this claim Analytics attribution depends on available campaign and referrer information, so some sessions can be classified as direct when no usable source is available. Scope: Google Analytics attribution behavior; does not establish a universal percentage for AI referrals. Confidence: high · Verified: Google Analytics Help: Traffic-source dimensions

Three things break at once.

Attribution failure. Most AI-referred traffic arrives with no referrer signal, so it lands in Direct rather than as an AI channel. Estimates of how much AI traffic is invisible this way range from about 35% to 70.6% (the high end is vendor-sourced from Loamly’s 446,405-visit sample — use the range as directional, not gospel). This isn’t a Google conspiracy; it’s how referrers work. As I put it when we tested this at Ahrefs: “Websites have control over what info they send. They can send the full path, just the origin, or nothing — it’s up to them. We report whatever referrer we’re told to report. If they don’t send us one, then it would go in the ‘Direct’ bucket.”

Visibility without clicks. AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. generate impressions and brand exposure without sending traffic. The old impressions → clicks → conversions funnel doesn’t hold when the answer is the destination.

No standard tooling. Until May 2026 there was no native GA4 channel for AI traffic. GSC still does not separate AI Overview clicks from regular organic. Bing only added AI performance metrics in February 2026 as a public preview. And the result of all this friction: only 16% of brands systematically track AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. performance (McKinsey, September 2025). Most of your competitors are measuring nothing — which means any real measurement system is an edge.

The floor-not-ceiling principle

If you only remember one framing, make it this one. Published AI-traffic shares hover around 0.25% of average site traffic (Ahrefs’ study of ~82,000 sites) — but that’s just the measurable slice. Two corrections inflate it:

  1. Dark trafficAI traffic attribution is the practice of correctly identifying and measuring website visits that come from AI tools — ChatGPT, Perplexity, Gemini, Claude, AI Overviews, and AI browsers. It's hard because many of those tools strip the referrer header, so the visits land in your analytics as Direct traffic with no source.. With 35–70% of AI visits referrer-less, the real number is plausibly 2–3x what your analytics show.
  2. The conversion premium. AI traffic converts dramatically better. For Ahrefs, 0.5% of visitors drove 12.1% of signups — a 23x premium, and those visitors browsed ~50% more pages per session with a lower bounce rate. Industry-wide the premium is more like 4–4.4x (Semrush/Adobe), but the direction is consistent.

So the right number to open a stakeholder report with is not 0.25%. It’s the growth rate (9.7x in 12 months) or the conversion premium. Those reframe the stakes; the traffic-share number undersells them.

The five layers of measurement

No single tool sees the whole picture, so you stack partial views. This is Paul DeMott’s 5-layer GEOGenerative Engine Optimization (GEO) is the practice of optimizing content and brand presence so AI-powered search engines and assistants — Google AI Overviews, ChatGPT, Perplexity — cite, recommend, or mention you when generating answers. Google's position is that it's still SEO. framework (Search Engine Land, May 2026), with my own data folded into each layer. Each layer below is also its own deep-dive in this cluster.

Layer 1 — Direct attribution (retrieved vs. mentioned vs. cited)

This is GA4 plus Ahrefs Web Analytics: who actually visited, from which AI source.

  • GA4’s AI Assistant channel (added May 13, 2026) catches referred sessions from ChatGPT, Gemini, Claude, Copilot, Grok, and similar. Useful — but it only sees sessions that arrive with a referrer. The 35–70% that don’t still sit in Direct, no matter how you configure channels. It also excludes Google AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. and AI Mode, which appear as plain Organic Search, didn’t apply retroactively, and uses one of your two custom channel-group slots.
  • Ahrefs Web Analytics has the AI channel built in rather than requiring custom setup, and updates closer to real time.
  • The vocabulary matters here: retrieved ≠ mentioned ≠ cited. A retrieval is your content being fetched; a mention is your brand named in the answer; a citation is your URL linked as a source. Build a report on the wrong one and the whole thing misleads. See Retrieved vs. Mentioned vs. CitedThree distinct states of AI visibility: retrieved (an AI fetched your page as source material), mentioned (your brand appears in the answer text), and cited (your URL is linked as a source). They don't always happen together, and each is measured with a different tool. in AI.

Layer 2 — LLM visibility (the impressions-and-citations layer)

How often you appear in AI answers, click or no click.

  • GSC Gen AI Performance ReportsThe Google Search Console report that shows how your site actually performed in Google Search, built from real impressions and clicks. It reports four metrics — clicks, impressions, average CTR, and average position — and keeps the most recent 16 months of data. (June 2026) show impressions only — no clicks, no CTR, no position — for AI Overviews, AI Mode, and Discover AI features. Real exposure signal, but you can’t turn it into traffic. (More below and in the GSC cluster article.)
  • Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility.’ AI Performance report (Feb 2026 preview) is more transparent: total citations, average cited pages, groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. queries (the phrases Copilot searched internally to find you), and page-level citation activity.
  • Reality check: only 38% of pages cited in Google AI Overviews ranked in the traditional top 10 (down from 76%) — LLM visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals. is not the same as your blue- link rankings.

Layer 3 — AI hallucination monitoring

A wrinkle that has no equivalent in traditional search: AI tools send traffic to pages that don’t exist and describe your brand inaccurately. In my Ahrefs AI traffic page-type analysis, I found 3.6% of Ahrefs’ AI assistant traffic went to non-existent (hallucinated) URLs. Beyond bad links, you want to watch whether models describe your product, pricing, and positioning correctly — a wrong “fact” repeated across models is a measurable reputation problem. See AI Hallucination MonitoringAI hallucination monitoring is the practice of systematically detecting, documenting, and fixing instances where AI search tools and chatbots fabricate or misrepresent facts about your brand — wrong pricing, invented quotes, made-up URLs, and the like..

TIP Add a repeatable known-fact check to the measurement stack

A sampled response can mention a brand and still omit an important fact. Capture both observations, plus the provider and retrieval state, instead of collapsing them into one visibility score.

Run a narrow answer-level fact check with my free AI Brand Visibility Checker Free

  1. Choose a fact that is objectively true and material to how the brand should be described.
  2. Capture the answer, provider, model, retrieval mode, date, mention result, and fact-coverage result.
  3. Trend repeated samples; do not turn unevaluated providers or one missing fact into a universal hallucination rate.
The sample creates two reportable fields—brand mentioned and known fact missing—without pretending the unevaluated providers were tested.

A simulated Llama response mentions the fictional Acme Analytics brand. The known-fact check marks Founded in 2018 as missing. Mistral, ChatGPT Search, Gemini, Claude, and Perplexity are marked not evaluated with a warning not to infer zero rates.

Layer 4 — Self-report (the bridge analytics can’t build)

Add a “How did you first hear about us?” field with AI options to your sales, contact, and post-conversion forms. This is the only layer that captures AI’s top-and-middle-of-funnel influence — the discovery that happened weeks before a referrer-less Direct visit. DeMott reports this surfaces double-digit AI attribution in some pipeline studies. It’s low-tech and it works precisely where the tracking fails. See AI Traffic AttributionAI traffic attribution is the practice of correctly identifying and measuring website visits that come from AI tools — ChatGPT, Perplexity, Gemini, Claude, AI Overviews, and AI browsers. It's hard because many of those tools strip the referrer header, so the visits land in your analytics as Direct traffic with no source. (which folds the dark-traffic problem and the self-report fix together).

Layer 5 — Share of voice (and why it’s a trap on its own)

Share of Voice is the percentage of relevant AI answers in which your brand is mentioned or cited. Tools automate it at scale: Ahrefs Brand Radar (400M+ search-backed prompts across ChatGPT, Perplexity, Gemini, Copilot, AI Overviews, AI Mode, and Grok as of July 2026, up from 350M+ earlier this year — it’s a growing, live-updating index, so treat the exact figure as directional; monthly refresh), Semrush AI Toolkit (100M+ prompts), Profound, BrightEdge, Scrunch. You can also do it manually: run a fixed prompt set across 3+ models monthly and tally mentions.

But heed DeMott’s warning: “Share of Voice is a vanity metric without business connection.” SOV tells you how often you show up, not whether showing up drives awareness, traffic, or pipeline. Always pair it with Layer 1 (traffic) and Layer 4 (self-report). See AI Share of VoiceAI Share of Voice (SoV) measures how often and how prominently a brand appears in AI-generated responses relative to competitors, across a defined pool of relevant prompts. It's a visibility signal, not a traffic or revenue metric. (SoV).

One more honest layer beyond these five: incrementality testing. Difference-in-differences — a high-AI-visibility test cohort vs. a control — over 6–12 months is the only way to prove AI search caused revenue rather than merely correlating with it. It’s the slowest and the most rigorous. Most teams won’t get here for a year; collect the baseline data now so you can.

Platform-specific reporting (the gotchas)

Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.. AI Overview and AI Mode clicks are merged with standard organic in the Performance report — you cannot isolate them, and you shouldn’t claim you can. The June 2026 Gen AI reports add an impressions-only layer on top. The useful signal: rising AI Overview impressions alongside flat or falling clicks in the main report is “The Great Decoupling” — measurable AI Overview cannibalization (Ahrefs saw the blog’s clicks/impressions correlation flip from +0.425 to −0.352, with AIOs tied to a 34.5% CTR reduction).

Bing Webmaster Tools. Ahead of Google on transparency. Its AI Performance report exposes citations separately from organic clicks, and the grounding queries are a genuinely unique window — they tell you what Copilot was actually trying to answer when it pulled your page. Compare your most-cited Bing pages to your top organic pages; gaps are opportunity.

GA4 & Ahrefs Web Analytics. Covered in Layer 1 — both carry the same dark-traffic limitation; Ahrefs is built-in and faster, GA4 is configurable and excludes AIO/AI Mode.

What correlates with LLM visibility (Ahrefs’ 75,000-brand study)

In Ahrefs’ 75,000-brand study by Louise Linehan and Xibeijia Guan, which I discussed in my Evolve 2025 talk, the signals ranked like this by correlation with AI Overview visibility:

  • Branded web mentions — 0.664 (the strongest signal)
  • Branded anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. — 0.527
  • Branded search volume — 0.392
  • Domain Rating — 0.326 (the weakest of the four)

Read that list again. Three of the four are brand signals, not classic technical SEO ones. The implication is the through-line of this whole hub: measuring AI visibility is mostly measuring brand health, and the work that improves AI visibility (mentions, citations, authority) is the same work that improves traditional SEO.

Where to go next

This hub is the map. Each layer is its own deep dive in the cluster:

  • Retrieved vs. Mentioned vs. Cited in AI — the three visibility types you must not conflate, and which ones you can actually measure.
  • LLM Visibility / AI Visibility — impressions, citations, and how to audit your presence in AI answers with tools like Brand Radar.
  • AI Hallucination Monitoring — tracking wrong facts, bad pricing, and the hallucinated-URL traffic wrinkle.
  • AI Traffic Attribution — the dark-traffic problem in full, platform-by-platform referrer behavior, GA4 setup, and the self-report bridge.
  • AI Share of Voice (SoV) — defining it, measuring it, the tools, and how to keep it from being a vanity metric.
  • GA4 for AI TrafficGA4 for AI traffic is the practical configuration work — native AI Assistant channel, a custom channel group with a source regex, Explorations, and a Search Console join — that lets you identify, segment, and report on the visits arriving from AI chatbots that GA4 can actually see. — the actual GA4 configuration: custom channel groups, referrer regex, and Explorations for segmenting AI-platform traffic.
  • AI Crawler Log AnalysisAI crawler log analysis is the practice of pulling raw server or CDN access logs and examining them for requests from AI bots — training crawlers, AI-search indexers, and user-triggered fetchers — to verify with first-party data which bots actually hit your site, whether they're real or spoofed, and what they got. — reading raw server logs to measure AI-bot crawl activity, verify user-agents against real IPs, and spot crawl-vs-render problems.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.