Retrieved vs. Mentioned vs. Cited in AI
Retrieved, mentioned, and cited are three different things AI search can do with your content — each needs a different tool to measure and a different lever to fix.
Retrieved, mentioned, and cited are three separate states, not one. Retrieved = an AI fetched your page (visible only in server logs — and the -User bots are the signal that matters). Mentioned = your brand is in the answer text with no link. Cited = your URL is linked. They don't move together: 99.6% of AI influence is invisible (OtterlyAI), ChatGPT cites only ~15% of what it retrieves, and even being cited rarely earns a click (1% per Pew). No single tool sees all three — you need logs for retrieval, brand monitoring for mentions, and citation tracking for citations.
TL;DR — “Showing up in AI” is really three different things. Retrieved means an AI read your page. Mentioned means your name appeared in the answer. Cited means your link was in the answer. They don’t always happen together — you can be read and never named, or named and never linked — and you need a different tool to spot each one.
Three things, not one
Retrieval, mention, and citation are distinct observable outcomes; no platform guarantees that crawl access produces any of them. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: OpenAI: ChatGPT search OpenAI documents separate crawler user agentsA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. and controls for search versus training contexts. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: OpenAI: Crawlers
When someone says “I want to show up in ChatGPT,” they’re usually picturing one thing: their link in the answer. But there are actually three separate things an AI can do with your content, and people mash them all together under “AI visibility.” That’s where the confusion starts.
- Retrieved — the AI fetched and read your page while it was building the answer. It might use what it learned, it might not. You won’t see this in Google Analytics; the only place it shows up is your server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened..
- Mentioned — your brand or name appears in the text of the answer, but without a link. The AI just “knew about you” — often from its training, not from reading your site right then.
- Cited — your actual URL is in the answer as a source you can click.
Why the difference matters
Here’s the part people miss: these don’t move together. Your page can be read without being named, and named without being linked. One analysis of Microsoft’s groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. that 99.6% of AI influence is invisible — content gets used behind the scenes constantly but almost never turns into a visible citation.
So if your AI tool only counts citations (links), it’s showing you a sliver of the picture. You could be shaping answers all day and that tool would say you’re nowhere.
What this means for you
- Want to know if AIs are reading you? Check your server logs for AI botsAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls..
- Want to know if you’re getting named? Use a brand-monitoring tool that reads AI answers.
- Want to know if you’re getting linked? Use a citation tracker (and Bing now has an official report for its own AI answers).
And one more thing worth internalizing early: even getting cited usually doesn’t get you a click. Pew found only about 1% of people click the cited sources in an AI summary. So “show up in AI” is less about traffic today and more about being in the conversation. Want the full breakdown — the botsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index., the platforms, and how to measure each layer? Switch to the Advanced tab.
TL;DR — Retrieved, mentioned, and cited are three distinct states with three distinct detection methods. Retrieved lives in your server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. — and the
-Usersuffix botsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (ChatGPT-User,Perplexity-User) are live-inference fetchers, a far better citation signal than training crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. likeGPTBot. Mentioned is a brand-name string match in the response (Brand Radar). Cited is a linked URL (Brand Radar Cited Pages, Bing WMT AI Performance). They don’t move together: 99.6% of AI influence is invisible (OtterlyAI), ChatGPT cites ~15% of what it retrieves, ~82.9% of B2B citations come from third-party sites, and even a citation earns a click ~1% of the time. No single tool sees all three; you need a stack.
The three states — and why conflating them costs you
Retrieved means a system fetched or selected the content as candidate context, evidenced by logs or retrieval traces. Mentioned means the answer used the brand, entity, or facts in its prose, evidenced by the answer text. Cited means the interface exposed a source link or citation to the page, evidenced by a visible source URL. A later state does not make every earlier state directly observable.
© Patrick Stox LLC · CC BY 4.0 ·
A model may use retrieved information without displaying a link, so external measurement cannot always prove the internal source path. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: OpenAI: ChatGPT search CrawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. permissions indicate access policy, not guaranteed retrieval, mention, or citation. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: OpenAI: Crawlers
Most “AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.” conversations collapse three different events into one word. They’re not the same thing, and treating them as one is how people end up optimizing for the wrong signal and measuring the wrong number.
Retrieved — an AI’s retrieval system fetched your page as source material while it was generating an answer. This is the retrieval/groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. phase of RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. (retrieval-augmented generation). The model “read” you. You may get no mention and no link out of it. The only place this shows up is your server logs — no response-analysis tool can see it, because nothing about it appears in the answer.
Mentioned — your brand, product, or name appears in the generated text, but with no hyperlink and no source attribution. The model knew about you, often from training data alone, and dropped your name in. “Tools like Ahrefs, Semrush, and Moz can help with keyword research” — that’s three brands mentioned, zero links. Detectable by string-matching the response corpus (Brand Radar).
Cited — your URL is included as an attributed source: a clickable link, a numbered reference, a named source box. This requires the AI to have both retrieved your content and decided it was worth attributing. It’s what everyone calls “AI SEOAI search optimization is the practice of making your brand and content visible, citable, and accurately represented across AI-powered search — Google AI Overviews, ChatGPT, Perplexity, Copilot. It's built on traditional SEO plus a heavier emphasis on off-site brand mentions and content AI systems can cite. success” — and it’s the tip of the iceberg.
The key gap: content can be retrieved without being mentioned, and mentioned without being cited. Each state needs a different measurement tool and a different optimization lever.
The distinction that matters most: training crawl vs. live inference
This is the single most useful thing to understand about detecting retrieval, and it’s the one I see skipped constantly.
AI companies run two different kinds of bots, and they mean completely different things in your logs:
- Training / indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. crawlers —
GPTBot,ClaudeBot,PerplexityBot,Googlebot-Extended. These crawl to build or update a model or a search index. AGPTBothit tells you OpenAI might use your content someday, somewhere. It has no direct connection to any specific citation event. - Live-inference / citation fetchers —
ChatGPT-User,Claude-User,Perplexity-User. These fetch your page in real time, during an actual user conversation, to ground a specific answer.
That -User suffix is the tell. A ChatGPT-User hit in your logs is about as
close as you can get to knowing “OpenAI just fetched this page to potentially cite
it in a live answer.” As Screaming Frog puts it: “For citation bots like
ChatGPT-User or Perplexity-User, every error response is a missed opportunity… if
they can’t access it, your site won’t appear in the response.” A 404 or a CDN
block to one of these bots is a citation you lost in real time.
Don’t conflate the two. “GPTBot crawled me, so ChatGPT will cite me” is a myth — they’re fundamentally different signals.
Why the gap is so wide — the invisible-influence problem
The states diverge far more than most people expect:
- 99.6% of AI influence is invisible. OtterlyAI’s analysis of Bing Webmaster Tools groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data found the overwhelming majority of retrieval never surfaces as a user-visible, attributed citation. This is the most concrete published number on the retrieved-but-not-cited gap.
- ChatGPT cites only ~15% of the pages it retrieves (Leapd / Discovered Labs). Being in the retrieval pool is necessary but nowhere near sufficient.
- Your own site is rarely the source. ~82.9% of B2B citations come from third-party sites (Discovered Labs); McKinsey puts a brand’s own website at just 5–10% of the sources AI references about it. The internet talking about you matters more than your own pages.
- Even being cited rarely earns a click. In Pew’s March 2025 U.S. browsing study, users clicked a standard result in 8% of Google visits with an AI summary (vs. 15% without) and clicked a cited source in 1%. In Semrush’s separate clickstream sample, 92–94% of Google AI Mode sessions did not lead to an external-domain visit.
So the gap compounds: retrieved → mentioned → cited → clicked, and you leak volume at every step.
The counterweight — and the reason this is still worth doing — is quality. At Ahrefs, June 2025 internal data showed AI traffic converting at 23x the rate of traditional organic trafficVisitors from unpaid search results — it compounds without ad spend.: about 0.5% of visitors came from AI but accounted for ~12.1% of signups during the measured 30-day period. Few clicks, but the ones you get are extraordinary.
How each platform handles retrieval and citation
Same content, same query, wildly different attribution outcomes by platform:
- ChatGPT (OpenAI) — base training data plus a retrieval layer when search is
on. ChatGPT Search launched on Bing’s index, but OpenAI also runs its own
crawler (
OAI-SearchBot) for independent fetching and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and hasn’t published the current mix between the two — treat “it’s basically Bing” as a simplification, not a documented fact. Cites ~15% of what it retrieves; ~0.59% citation rate across tested topics — the lowest of the majors. It will happily discuss your brand from training data with no retrieval and no citation — the purest “mentioned but not retrieved or cited.” - Perplexity — runs a real-time search on every query with its own
PerplexityBot. Averages 21.87 citations per response, the highest of any major platform, and is the least likely to mention without citing. The most transparent platform. - Google AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. / AI Mode — your organic index + Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself., with
Gemini generating the answer. Uses query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics. (one user query → many
internal sub-queries), so pages get retrieved for sub-queries that never appear
in the visible answer — dark retrieval by design. Citations attach to specific
text segments via
url_citationannotations. - Bing Copilot — Bing index + GPT layer; averages 6.89 citations per response. As of Feb 2026 it has the first official webmaster tool exposing grounding/citation data (see Official Docs).
- Claude (Anthropic) — frequently answers from training data with no retrieval and 0 citations, but when it does retrieve, it’s among the highest quality and most likely to cite clear, well-structured content. Lowest retrieval frequency, highest retrieval-to-citation ratio.
How to detect each state
Retrieved → server logs. Logs record every bot request, including from AI
crawlers that mostly fetch raw HTML without renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. JavaScript — that varies by
provider, though: Google’s Gemini-based answers reuse GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer.’s renderer, so
don’t assume every AI crawler skips rendering. None of it shows up in GA. Watch
the user agentsA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target., and above all separate -User (live inference) from the base
crawlers (training).
Tools: Screaming Frog Log File Analyser, Botify, Ahrefs Bot Analytics, Cloudflare.
The limitation: logs prove a fetch happened, not that you were cited. And standard
analytics misclassify the majority of AI referrals as direct traffic — logs are
the only place the full retrieval picture is visible.
Mentioned → brand monitoring. Ahrefs Brand Radar string-matches your brand
against a large search-backed prompt corpus across ChatGPT, Gemini, Copilot,
Perplexity, AI Overviews, AI Mode, and Grok. The filter that matters here is
Response contains {your brand}. The catch: it runs a static prompt library and
can’t see the AI-internal “dark queries” that make up most of the real retrieval
surface, and AI answers are wildly inconsistent run to run (SparkToro: <1 in
100 chance of identical brand lists across 100 runs of the same prompt). Treat it
as a directional trend, not a precise count.
Cited → citation tracking. Brand Radar’s Cited Pages (your pages that get
cited) and Cited Domains (which third-party sites get cited when you’re discussed)
reports; Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. AI Performance for official Copilot data; GSC for AI
Overview traffic (it lands under the Web search type). The “found but not cited”
move — Response contains {your brand} AND citation does not contain {your domain}
— isolates exactly the gap: answers shaped by you, with no link to you.
Two inverse gaps worth knowing
- Mentioned, never retrieved or cited. A model names your brand straight from training data, no live fetch, no link. Common for well-known brands.
- Cited, never mentioned (“ghost citations”). Superlines found Gemini cited one domain 182 times in 30 days while saying the brand name zero times — 73% of that brand’s AI presence was citations with no name attached. You’re feeding the answer authoritative signals while the user never associates it with you.
This is why mention tracking and citation tracking have to run in parallel — neither one implies the other.
What to do about the gaps
- Retrieved but not cited — first, stop losing it at the door: no errors or CDN
blocks to
-Userbots. Then improve completeness and structure (clear, direct answers high on the page; question-mirroring headings). In Discovered Labs’ January 2026 analysis, semantic completeness correlated ~0.89 with citation likelihood in Google AI Overviews — their own reported figure, not an independently audited one. - Mentioned but not cited — build more explicit attribution signals on-page, and lean into the bigger lever: get the internet talking about you. Ahrefs’ #1 predictor of appearing in AI Overviews is branded web mentions (correlation 0.664 — well above Domain Rating at 0.326 or backlinks at 0.218), from the 75,000-brand study by Louise Linehan and Xibeijia Guan. Sixteen pages on Zapier mentioning Ahrefs was enough to put us in 1,431 AI responses.
- Cited but no traffic — this is mostly the current reality; optimize for being named alongside the citation (so there’s brand recall even without a click) and for the higher-value query types where AI traffic converts.
That’s the strategic shift in one line: SEO was “optimize your site.” Now it’s “optimize how the internet talks about you.” Optimizing your own pages affects whether you can be retrieved and cited; optimizing the conversation about you affects whether you get mentioned and whether third-party citations multiply your reach.
The measurement stack you actually need
No single tool covers all of this. The honest answer is a stack:
- Logs → the retrieval layer (Screaming Frog, Botify, Ahrefs Bot Analytics).
- Brand monitoring → the mention layer (Brand Radar, manual prompt testing).
- Citation tracking → the cited layer (Brand Radar Cited reports, Bing WMT AI Performance).
- GSC + GA4 → the traffic layer (and remember GA4 hides most AI referrals as direct).
One last caution: AI answers churn hard — roughly 40–60% of cited sources change month to month, and ~70% of AI Overview content changes for the same query. Track trends over many runs, not single snapshots, and don’t over-fit to a number that’s this volatile.
AI summary
A condensed take on the Advanced version:
- Three distinct states, three distinct tools. Retrieved (an AI fetched your page) shows up only in server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened.. Mentioned (your name in the answer, no link) needs brand monitoring. Cited (your URL linked) needs citation tracking. They don’t move together.
- Training crawl ≠ live inference.
GPTBot/ClaudeBot/Googlebot-Extendedcrawl for training/indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — no link to any citation event. The-UserbotsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (ChatGPT-User,Perplexity-User) fetch live, during a conversation — that hit is the closest signal that a page was used for a real answer. An error to a-Userbot is a citation lost in real time. - The gap is huge. 99.6% of AI influence is invisible (OtterlyAI/Bing groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it.); ChatGPT cites ~15% of what it retrieves; ~82.9% of B2B citations come from third-party sites; and even a citation earns a click ~1% of the time (Pew).
- Platforms diverge wildly. Perplexity ~21.87 citations/response (cites almost everything); ChatGPT ~0.59% citation rate (mentions from training without citing); Claude often answers with 0 retrieval; Google uses query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics. (dark retrieval); Bing Copilot ~6.89, now with an official report.
- Two inverse gaps: mentioned-but-never-retrieved (from training data) and ghost citations (cited without the brand named — Superlines: 73% of one brand’s AI presence).
- Fix levers: unblock
-Userbots and improve completeness (retrieved→cited); build branded web mentions, the #1 AIOAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. predictor at 0.664 correlation (mentioned→cited); optimize for named co-citation (cited→recall). - Measure with a stack (logs + brand monitoring + citation tracking + GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results./GA4), track trends not snapshots — AI answers churn 40–60% month to month.
Official documentation
How the engines describe RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. retrieval and citation selection, from the primary sources.
- AI features in Google Search — how AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. and AI Mode use core Search to retrieve and attribute sources.
- AI optimization guide — Google’s stance that standard SEO best practices apply; no special AIO optimizations.
- Grounding with Google Search (Gemini API) — the five-step groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. process and the
url_citationannotations that link a text segment to a source URL; thegoogle_search_callobject logs everything retrieved, while annotations log only what was used.
Bing / Microsoft
- Introducing AI Performance in Bing Webmaster Tools (public preview, Feb 2026) — the first official platform tool to expose retrieval/groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. signals: Citations, Grounding Queries, and Average Cited Pages.
- Bing Webmaster Tools — AI Performance help — metric definitions for the report.
- Search Engine Land coverage of the Bing AI Performance report — context on what it does and doesn’t show.
Quotes from the source
On-the-record statements from the engines on how retrieval and citation selection work. Deep links jump to the quoted passage where the page allows it.
Google — the groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. / retrieval mechanism
- “The model analyzes the prompt and determines if a Google Search can improve the answer.” — Gemini API groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. docs (step one of the documented five-step process). Source
- “If needed, the model automatically generates one or multiple search queries and executes them.” — the query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics. that creates dark retrieval: pages fetched for sub-queries that never appear in the visible answer. Source
- “Each
url_citationannotation links a text segment (defined bystart_indexandend_index) to a source URL.” — citations attach to specific segments; the full retrieved-source list is not exposed to end users. Source - “There are no additional requirements to appear in AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. or AI Mode, nor other special optimizations necessary.” — Google’s official optimization stance. Source
Bing / Microsoft — the AI Performance reportThe Google Search Console report that shows how your site actually performed in Google Search, built from real impressions and clicks. It reports four metrics — clicks, impressions, average CTR, and average position — and keeps the most recent 16 months of data.
- Citations: “Shows the total number of citations that are displayed as sources in AI-generated answers during the selected time frame.” — the Cited state. Source
- Grounding Queries: “Shows the key phrases the AI used when retrieving content that was referenced in AI-generated answers.” — the closest official surface to the Retrieved state. Source
- Average Cited Pages: “Shows the average number of unique pages from your site that are displayed as sources in AI-generated answers per day.” Source
How to check each layer
A pass to confirm which of the three states you’re actually in — each layer needs its own check.
Retrieved? (server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened.)
- Pull server logs (or Cloudflare/CDN logs) and filter for AI user agentsA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target..
- Separate training crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. (
GPTBot,ClaudeBot,PerplexityBot,Googlebot-Extended) from live-inference botsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (ChatGPT-User,Claude-User,Perplexity-User). - Confirm
-Userbots are getting 200s, not 404/403/5xx or CDN blocks — every error to a-Userbot is a citation lost in real time. - Verify the bots are genuine (reverse + forward DNS / published IP ranges) — the user-agent string alone is spoofable.
Mentioned? (brand monitoring)
- Run
Response contains {your brand}in Brand Radar across the platforms. - Spot-check manually: ask brand-related questions on each AI and look for your name appearing without a link.
- Run the same prompt many times — answers are inconsistent run to run; one snapshot isn’t data.
Cited? (citation tracking)
- Check Brand Radar Cited Pages (your URLs) and Cited Domains (third-party sites cited when you’re discussed).
- Check Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. → AI Performance for official Copilot citation counts and groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. queries.
- Check GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. for AI OverviewAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. traffic (under the Web search type).
- Run the found-but-not-cited filter:
Response contains {your brand} AND citation does not contain {your domain}— that’s your shaping-without-credit gap.
Cross-checks
- Watch for ghost citations — cited URLs where your brand name never appears in the text. Mention and citation tracking must run in parallel.
- Confirm you haven’t blocked AI bots in
robots.txt— blocking removes you from all three states at once.
Retrieved vs. Mentioned vs. Cited — cheat sheet
The three states side by side
| Retrieved | Mentioned | Cited | |
|---|---|---|---|
| What it means | An AI fetched your page as source material (RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. retrieval/groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it.) | Your brand/name appears in the answer text — no link | Your URL is an attributed, clickable source |
| How to detect it | Server log analysis | String-match the AI response corpus | Citation reports + linked-URL inspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. |
| Tools | Screaming Frog Log File Analyser, Botify, Ahrefs BotA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. Analytics, Cloudflare | Ahrefs Brand Radar (Response contains {brand}), manual prompts | Brand Radar Cited Pages/Domains, Bing WMT AI Performance, GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. |
| Why it matters | Necessary for citation; the -User bot hit is your earliest real signal | Brand awareness inside the answer, even with no click | The strongest signal — authority + (rare) clicks |
| The catch | A fetch ≠ a citation; invisible to GA | Often from training data with no retrieval; inconsistent run to run | ~82.9% of citations are third-party; ~1% click rate |
Bots in your logs — what each one means
| Bot | Owner | Type | Signal |
|---|---|---|---|
GPTBot | OpenAI | Training crawl | Under consideration for training/indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — not a citation signal |
OAI-SearchBot | OpenAI | Search indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. | Indexed for ChatGPT search retrieval |
ChatGPT-User | OpenAI | Live inference | Fetched in real time during a conversation — closest “about to cite” signal |
ClaudeBot | Anthropic | Training crawl | Crawled for Claude training |
Claude-User | Anthropic | Live inference | Fetched in real time during a Claude conversation |
PerplexityBot | Perplexity | Search indexing | Indexed for Perplexity’s corpus |
Perplexity-User | Perplexity | Live inference | Fetched in real time during a Perplexity query |
Googlebot-Extended | AI retrieval | Fetched for potential AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. / AI Mode use |
Fast facts
- 99.6% of AI influence is invisible (OtterlyAI / Bing groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data).
- ChatGPT cites ~15% of the pages it retrieves; ~0.59% citation rate overall.
- Perplexity averages 21.87 citations/response (most transparent); Bing Copilot 6.89.
- ~82.9% of B2B citations are third-party; a brand’s own site is 5–10% of references (McKinsey).
- 1% of Google visits with an AI summary included a cited-source click (Pew); 92–94% of Google AI Mode sessions did not lead to an external-domain visit in Semrush’s separate clickstream sample.
- The rule: training bot ≠ citation signal;
-Usersuffix bot = live retrieval signal.
Funnel-state measurement mistakes
Calling a crawler hit a citation
A fetch shows retrieval activity, not that the answer named or linked the source. Use answer captures for mentions and citations instead of promoting log evidence into a later funnel state.
Inferring retrieval from a mention
The model can produce a brand name from parametric knowledge or other sources. Mark retrieval only when the relevant log or product evidence supports it.
Treating all bot traffic as user-triggered retrieval
Training, indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., monitoring, and user-triggered agents do different jobs. Segment documented botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. identities and avoid labeling every AI crawlerAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. request as answer-time use.
The Evidence Ladder
| State | Minimum evidence | What it does not prove |
|---|---|---|
| Retrieved | A relevant user-triggered retrieval request in logs or equivalent product evidence | That the brand appeared in the answer |
| Mentioned | Captured answer text names the brand or entity | That a URL was linked or clicked |
| Cited | Captured answer contains a link to the site’s URL | That the user clicked or converted |
| Visited | Analytics or server event records the landing session | That the visit received a complete referrer |
| Converted | First-party event records the defined outcome | That the AI touchpoint deserves sole credit |
Move up the ladder only when evidence for that state exists. Report gaps as gaps rather than filling them with attribution assumptions.
Test yourself: Retrieved, mentioned, and cited
Resources worth your time
My related writing & talks
- GEO? AEO? LLMO? What’s With All This AI Stuff? (Ahrefs Evolve 2025) — my talk where the retrieved/mentioned/cited framework, the branded-mentions correlation, and the hallucinated-URL story come from. Slides · Webinar version.
- Cutting-Edge AEO Strategies with Patrick Stox (Marketing Speak, ep. 539).
- The Complete Guide to AI Visibility — the bigger picture this article sits inside.
- How to Track AI Overviews — the citation-tracking side, hands on.
- AI Visibility Audit — auditing all three layers.
- Brand Radar Methodology and 10 Ways to Use Brand Radar — including the “foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore., but not cited” use case.
- Meet the New Web Crawlers: AI Bots Are Closing in on Search Engine Bots — the botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. landscape behind the retrieval layer.
From others
- Screaming Frog — Monitor AI Bots in the Log File Analyser — the practical
-Uservs. training-bot tutorial. - Botify — Tracking AI Bots with Log File Analysis.
- Peec AI — A Beginner’s Guide to Source Gap Analysis — sources (retrieved) vs. citations (used), framed as a measurable gap.
- Discovered Labs — How Each Platform Cites Sources Differently.
- OtterlyAI — AI Citations Report 2026 — 1M+ data points from Bing groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data; the primary source for the 99.6% invisible-influence stat and the 73% crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index.-issues finding.
- SparkToro — AIs Are Highly Inconsistent When Recommending Brands — the research behind “<1 in 100 chance of identical brand lists across 100 runs.”
- Superlines — AI Search Statistics 2026 — platform-by-platform citation rates and the ghost citation phenomenon (cited without being named).
- BrightEdge — Brand Visibility in ChatGPT and Google AI — a comparison of how ChatGPT and Google AI Mode respond to action-oriented prompts across four industries.
- AiBoost — Log File Analysis for AI Bot Traffic — crawl-to-referral ratios: GPTBot ~1,276 crawls per referral click, ClaudeBot ~23,951.
Stats worth citing
- 99.6% of AI influence is invisible. OtterlyAI’s analysis of Bing WebmasterMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. Tools groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data — content is used extensively behind the scenes but rarely surfaces as a visible citation. The headline number on the retrieved-but-not-cited gap. Source
- Only 51.5% of AI-generated sentences are fully supported by citations (and only 74.5% of citations actually support their sentence) — Stanford, Evaluating Verifiability in Generative SearchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. Engines, EMNLP 2023. Source
- ChatGPT cites only ~15% of the pages it retrieves — Leapd / Discovered Labs. Source
- Users clicked a cited source in 1% of Google visits with an AI summary in Pew’s March 2025 U.S. browsing study; 92–94% of Google AI Mode sessions in Semrush’s separate clickstream sample did not lead to an external domain. And yet Ahrefs’ measured AI traffic converted at 23x the rate of traditional organic trafficVisitors from unpaid search results — it compounds without ad spend. in its June 2025 internal study — few clicks, exceptional ones.
- Perplexity averages 21.87 citations/response (highest, most transparent); Bing Copilot ~6.89; ChatGPT ~0.59% citation rate (lowest of the majors). Source
- ~82.9% of B2B citations come from third-party sources (Discovered Labs); a brand’s own website is just 5–10% of the sources AI references about it (McKinsey). The internet talking about you outweighs your own pages.
- Branded web mentions are the #1 predictor of appearing in AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. (correlation 0.664 — above Domain Rating 0.326 and backlinks 0.218), in Ahrefs’ 75,000-brand study by Louise Linehan and Xibeijia Guan. Source
- Ghost citations: Gemini cited one domain 182 times in 30 days while naming the brand zero times; 73% of that brand’s AI presence was citations without mentions (Superlines). Source
- AI answers churn: ~40–60% of cited sources change month to month, ~70% of AI Overview content changes for the same query, and there’s <1 in 100 chance two runs of the same prompt return identical brand lists (SparkToro). Track trends, not snapshots. Source
Retrieved vs. Mentioned vs. Cited
Three distinct states of AI visibility: retrieved (an AI fetched your page as source material), mentioned (your brand appears in the answer text), and cited (your URL is linked as a source). They don't always happen together, and each is measured with a different tool.
Related: LLM Visibility, AI Share of Voice, Measurement and Reporting
Retrieved vs. Mentioned vs. Cited
When people talk about “AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.,” they usually mean one thing — getting cited. But there are actually three separate states your content can occupy in an AI answer, and conflating them leads to bad measurement and wasted effort.
- Retrieved — an AI’s retrieval system fetched your page as source material while generating a response. This is the groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it./retrieval phase of RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. (retrieval-augmented generation). You may get no mention and no link out of it. It’s detectable in server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened., where live-inference botsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (the
-Usersuffix ones, likeChatGPT-User) are the closest signal you can get to “this page was used for a live answer.” It is invisible to response-analysis tools. - Mentioned — your brand, product, or name appears in the AI’s generated text, but with no link or source attribution. The model “knows about you” — often from training data alone, with no live retrieval at all. Detectable via brand-monitoring tools that string-match the response corpus (e.g. Ahrefs Brand Radar).
- Cited — your page’s URL is included as an attributed source: a clickable link, a numbered reference, or a named source box. It requires the AI to have both retrieved your content and decided it was worth attributing. Detectable via citation-tracking reports (Brand Radar Cited Pages/Domains, Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. AI Performance).
The key gap: content can be retrieved without being mentioned, and mentioned without being cited. OtterlyAI’s analysis of Bing groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. data foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. that 99.6% of AI influence is invisible — used behind the scenes but never surfaced as a visible citation. No single tool covers all three states, which is why measuring AI visibility takes a stack (logs + brand monitoring + citation tracking), not one dashboard.
Related: LLM Visibility, AI Share of Voice, Measurement and Reporting
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 22, 2026.
Editorial summary and recorded change details.Summary
Corrected the attribution of Ahrefs' 75,000-brand AI Overview correlation study.
Change details
-
Attributed the 0.664 branded-mention correlation to researchers Louise Linehan and Xibeijia Guan instead of incorrectly calling it Patrick's own research.
Updated Jul 20, 2026.
Editorial summary and recorded change details.Summary
Corrected this article's own sourceUrl to match the confirmed ai-search measurement-and-reporting taxonomy path.
Change details
-
Updated sourceUrl to /ai-search/measurement-and-reporting/retrieved-mentioned-cited/ after verifying the current taxonomy placement.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 19, 2026.
Editorial summary and recorded change details.Summary
Corrected the ChatGPT Search / Bing framing to match the verified OAI-SearchBot correction made across the corpus tonight, tempered an absolute claim that AI crawlers don't run JavaScript, fixed a duplicated word, and corrected an unattributed semantic-completeness stat to its sourced figure.
Change details
- Before
ChatGPT (OpenAI) — base training data + a Bing-powered retrieval layer when Browse is on.AfterRewrote the Advanced-lens ChatGPT platform bullet to split what OpenAI documents (OAI-SearchBot's independent fetching/indexing) from what's undisclosed (the current Bing/own-index mix), instead of stating ChatGPT runs a flat 'Bing-powered retrieval layer.' - Before
Logs record every bot request, including AI crawlers that don't run JavaScript and never show up in GA.AfterTempered the 'How to detect each state' section's claim that AI crawlers 'don't run JavaScript' — that's not universal; Google's Gemini-based answers reuse Googlebot's renderer, so rendering behavior varies by provider. -
Fixed a duplicated 'At At Ahrefs' typo in the invisible-influence section.
- Before
Semantic completeness correlates ~0.87 with citation likelihood in Google AI Overviews.AfterCorrected the unattributed 'semantic completeness correlates ~0.87' stat to the sourced figure (r=0.89, Discovered Labs' January 2026 analysis) and added attribution/date.
Full comparison unavailable — no prior snapshot was archived for this revision.