LLM Visibility / AI Visibility

What LLM visibility (AI visibility) is, why it isn't the same as organic visibility, and how to measure it across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini.

First published: Jun 24, 2026 · Last updated: Jul 23, 2026 · Advanced
demand #1 in Measurement and Reporting#14 in AI Search#113 on the site

LLM visibility is how often and how prominently your brand shows up inside AI-generated answers. It's not the same as organic visibility — only ~38% of AI Overview citations come from the traditional top 10, and the #1 correlate is branded web mentions (0.664), not Domain Rating (0.326). It splits into three states (retrieved → mentioned → cited) that need different tools, behaves completely differently platform to platform, and is volatile enough that single prompt runs aren't data. Measure it with a defined prompt pool tracked as Share of Model Voice over time — and treat tools like Brand Radar as directional, not exact traffic counts.

TL;DR — LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). visibility is the aggregate of how often and how prominently a brand appears in AI-generated answers (AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., AI Mode, ChatGPT, Perplexity, Copilot, Gemini, Grok). It is not organic visibility — only ~38% of AI Overview citations come from the traditional top 10 (was 76%), and the strongest correlate is branded web mentions (0.664), not Domain Rating (0.326). It splits into three states — retrieved → mentioned → cited — that need different tools; citation-only measurement understates the picture badly. It behaves radically differently per platform, it’s volatile enough that single prompt runs aren’t data points, and the right unit of measurement is a defined prompt pool tracked as Share of Model Voice over time. Tools like Brand Radar give directional indicators, not exact traffic counts.

What LLM visibility is (and isn’t)

Measurements are conditional on prompts, model versions, time, locale, and sampling. Evidence for this claim ChatGPT search can answer with information from the web and provide linked sources. Scope: ChatGPT search; appearance in a sampled answer is product-, query-, locale-, and time-dependent. Confidence: high · Verified: OpenAI: ChatGPT search Provider crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. controls describe access policy, not guaranteed inclusion in answers. Evidence for this claim OpenAI distinguishes crawler controls for search inclusion from controls for model training. Scope: OpenAI's documented crawlers and controls; allowing a crawler does not guarantee retrieval, citation, or answer inclusion. Confidence: high · Verified: OpenAI: Crawlers

LLM visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals. — AI visibility, same thing — is the rolled-up measure of how often and how prominently a brand, domain, or page shows up in the answers AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. systems generate. It’s a category, not a single number, the same way organic visibility is. The difference is the surface it measures: a synthesized, written answer instead of a ranked list of links.

The thing I’d burn into your brain first: LLM visibility is not the same as organic visibility. In mid-2025, ~76% of AI Overview citations came from pages ranking in the organic top 10. By early 2026 that’s down to roughly 38% in our data (and other studies put it even lower). Organic rank used to be a decent proxy for AI presence. It isn’t anymore. These are two different signals, and you need two different measurement programs.

Why the divergence? Because the signals that drive AI visibility aren’t the ones that drive rankings. In Ahrefs’ study of 75,000 brands by Louise Linehan and Xibeijia Guan — findings I later discussed at Ahrefs Evolve 2025 — the order came out like this for Google AI Overview visibility:

SignalCorrelation with AI visibility
Branded web mentions0.664
Branded anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page.0.527
Branded search volume0.392
Domain Rating0.326

The #1 signal is brand presence off your own site. That’s the structural inversion: unlinked mentions — which pass no PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. and barely register in traditional SEO — are the strongest correlate of AI visibility. Which is exactly why I keep saying the job shifted from “optimize your site” to “optimize how the internet talks about you.” It lines up with the earned-media data too: 82–89% of AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. come from third-party sources (Forbes/TechCrunch/WSJ-type outlets), not brand-owned pages. One concrete example from the talk — Zapier had 16 pages mentioning Ahrefs, and those 16 pages were cited across 1,431 Ahrefs AI responses. Other people’s content about you does most of the work.

Three states: retrieved → mentioned → cited → clicked

“AI visibility” hides three genuinely different states. Conflating them is where most measurement goes wrong:

StateWhat it meansHow you see it
RetrievedThe AI’s retrieval system fetched your page as source materialServer logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. (ChatGPT-User, Perplexity-User, Googlebot-Extended)
MentionedYour brand name appears in the generated answer textBrand monitoring / string-match (Brand Radar); manual prompt testing
CitedYour URL is explicitly linked as a sourceBrand Radar Cited Pages/Domains; Bing WMT AI Performance; GSC AI features

And then a fourth, downstream of all three: clicked — the rare case where someone actually visits your site from the answer (GA4 / web analytics).

Most tools only see the cited state. The retrieved and mentioned states are largely invisible to citation-only tools — OtterlyAI’s read of Bing’s data put it bluntly: “99.6% of your AI influence is invisible.” So if your dashboard is counting citations and calling that your AI visibility, you’re understating it by a lot. And the reverse failure exists too — Superlines found ~73% of AI presence was citations without a brand mention (“ghost citations”). Run mention tracking and citation tracking together or you get a false picture either way.

A question underneath all three: is this even a retrieval-driven answer?

Retrieved, mentioned, and cited all assume the AI went and fetched something for this specific query. A meaningful share of AI answers don’t — the model answers from what it already learned during training (its parametric memory), with no live retrieval step at all. See RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. for the mechanics: production systems run a query classifier that decides, per query, whether to search — it isn’t a step that happens on every request.

This matters for measurement because it changes what’s actionable. A retrieval-driven answer can, in principle, be moved by on-page and off-page work — better content gets fetched, better mentions get pulled in. A memory-driven answer can’t be moved that way: the model already “learned” what it knows about your brand at training time, and no page edit changes that until the model is retrained — which happens on the provider’s schedule, not yours. If you’re seeing a stable brand description across many prompt runs with no citations and no sign of retrieval, you may be looking at a memory-driven answer. Track it, but don’t spend editing effort expecting it to move.

There’s no fully reliable way to tell which is which purely by reading the output — it takes deliberately testing for it: comparing a platform’s search-on vs. search-off behavior where that’s exposed, watching whether citations appear at all, or checking whether the answer changes when the underlying source page changes. Treat “no citations at all” and “citations present” as different measurement regimes with different remedies, not two scores on the same scale.

Platform by platform — why visibility varies so much

There is no single “AI visibility” number, because the platforms behave completely differently. They draw from different sources and cite at wildly different rates:

PlatformBehavior
PerplexityHeavy citer — ~21.87 citations per response; cites ~13% of the time
Google AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.Huge volume; RAG + query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics. over the standard index
Bing Copilot~6.89 citations per response
ChatGPTSelective — cites in well under 1% of responses (~0.59%); ~87% of its citations align with Bing’s top organic results
GrokCites frequently — ~27% of the time
Claude / othersOften answers from training data with no retrieval at all

The headline gap: Perplexity cites ~21.87 sources per response while ChatGPT cites in 0.59% of responses. The same content can be highly visible on one platform and effectively invisible on another. And cross-platform overlap is poor — only 7 of the top 50 most-cited domains appear across all three major platforms (AI Overviews, ChatGPT, Perplexity). LLM visibility is really three (or seven) separate visibility profiles. Measure each platform; don’t average them into one number and pretend it means something.

It’s volatile — single runs aren’t data

AI answers churn. Roughly 40–60% of cited sources change month to month, and for the same query a large share of the AI Overview content changes between runs. SparkToro found less than a 1-in-100 chance of getting identical brand lists across 100 runs of the same ChatGPT prompt. The methodological consequence is simple and non-negotiable: running a prompt once is not a data point. You need a defined prompt pool, many runs, and trend windows — not point-in-time snapshots.

How to measure it — the stack mapped to the states

No single tool covers all four states. Build a stack where each layer maps to one:

  • Layer 1 — server logs → Retrieved. The -User bots (ChatGPT-User, Perplexity-User, Googlebot-Extended) are live inference fetches — someone asked a question and the AI went to get your page right then. That’s different from training/indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. bots (GPTBot, ClaudeBot). Every error you serve a -User bot is a missed citation.
  • Layer 2 — brand monitoring → Mentioned. Ahrefs Brand Radar does string-match mentions across the major platforms and has a “found but not cited” filter for isolating influence-without-attribution. Add manual prompt testing for QA.
  • Layer 3 — citation tracking → Cited. Brand Radar’s Cited Pages / Cited Domains reports; Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility.’ AI Performance reportThe Google Search Console report that shows how your site actually performed in Google Search, built from real impressions and clicks. It reports four metrics — clicks, impressions, average CTR, and average position — and keeps the most recent 16 months of data. (the first official platform source of citation data — citations, groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. queries, average cited pages); and Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.’s AI features filter / Gen AI Performance reports (impressions only — no clicks).
  • Layer 4 — web analytics → Clicked. GA4’s AI Assistant channel, plus a custom channel group. Expect a dark-traffic problem: a large share of AI-sourced visits arrive with no referrer and land in Direct, and AI Overview/AI Mode clicks merge into google/organic — so the clicked layer is the least clean.

One caveat I want to be honest about: Brand Radar (and tools like it) report directional indicators, not exact traffic counts. Treat the trend and the relative share seriously; don’t treat any single figure as a precise tally.

TIP Classify a sampled answer without inventing missing-provider data

The pictured run is deliberately narrow: one simulated provider response, retrieval off, and five providers marked not evaluated. That makes it useful for classification practice—not for claiming market-wide visibility.

Check whether a sampled answer mentions the brand and contains a supplied known fact with my free AI Brand Visibility Checker Free

  1. Supply the answer text, target brand, and one fact whose presence can be checked directly.
  2. Record mention and known-fact coverage independently, along with provider and retrieval state.
  3. Repeat the controlled capture across the prompt pool and providers before trending visibility.
This sample demonstrates classification boundaries: mentioned does not mean every known fact is present, and not evaluated does not mean absent.

A simulated Llama answer mentions the fictional Acme Analytics brand by exact match, while the known-fact check marks Founded in 2018 as missing. Five other provider profiles are explicitly not evaluated and warn against inferring zero mention or citation rates.

The prompt-pool method and Share of Model Voice

The unit of measurement is the prompt pool: define ~250–500 high-intent queries, run them across the LLM endpoints on a weekly or monthly cadence, record where your brand appears, and track Share of Model Voice (SOMV) — brand appearances ÷ total tracked prompts × 100. Pull prompts from real demand: sales transcripts, support tickets, Reddit, G2 reviews, PAA, autocomplete. Benchmarks to calibrate against: average mention rate is around 17.2% across relevant prompts (AthenaHQ), and 40–70% is considered strong.

For platform mechanics, Brand Radar runs on a 400M+ search-backed prompt corpus (up from 350M+ earlier this year — it’s a growing, live-updating index, so treat the exact figure as directional), refreshes monthly for ChatGPT/Perplexity/Gemini/Copilot and continuously for AI Overviews/AI Mode, and distinguishes “cited” (linked URLs) from “found but not cited” (string matches without links).

A measurement edge case worth knowing

AI doesn’t just under-report you — sometimes it invents you. We had thousands of visits going to pages on Ahrefs that didn’t actually exist because an AI synthesized plausible-looking URLs. Hallucinated citations are a real dimension of AI-visibility measurement, and most tools won’t filter them out for you — so sanity- check the URLs you’re “cited” on.

What this is and isn’t worth

Be clear-eyed about value. Even being cited rarely produces a click — users clicked a cited source in only 1% of Google visits with an AI summary in Pew’s March 2025 U.S. browsing study, and the vast majority of Google AI Mode sessions end with no site visit. LLM visibility is primarily a brand awareness and authority metric, not a direct-traffic metric. It still matters: a large share of AI Mode users accept the curated shortlist without doing further research, so being in the answer carries weight even without the click.

Where this fits

This is the definition-and-measurement hub for the cluster. For the deep dive on the three states, see Retrieved, Mentioned, CitedThree distinct states of AI visibility: retrieved (an AI fetched your page as source material), mentioned (your brand appears in the answer text), and cited (your URL is linked as a source). They don't always happen together, and each is measured with a different tool.. For the GA4/analytics layer specifically — the dark-traffic and attribution mess — see AI traffic attributionAI traffic attribution is the practice of correctly identifying and measuring website visits that come from AI tools — ChatGPT, Perplexity, Gemini, Claude, AI Overviews, and AI browsers. It's hard because many of those tools strip the referrer header, so the visits land in your analytics as Direct traffic with no source.. For the off-site signal that actually drives this (entities and mentions), see Entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. and the broader AI search hubAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity..

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.