LLM Visibility / AI Visibility

What LLM visibility (AI visibility) is, why it isn't the same as organic visibility, and how to measure it across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini.

First published: Jun 24, 2026 · Last updated: Jul 27, 2026 · Advanced
demand #1 in Measurement and Reporting#22 in AI Search#158 on the site

LLM visibility is how often and how prominently your brand shows up inside AI-generated answers. It's not the same as organic visibility — only ~38% of AI Overview citations come from the traditional top 10, and the #1 correlate is branded web mentions (0.664), not Domain Rating (0.326). It splits into three states (retrieved → mentioned → cited) that need different tools, behaves completely differently platform to platform, and is volatile enough that single prompt runs aren't data. Measure it with a defined prompt pool tracked as Share of Model Voice over time — and treat tools like Brand Radar as directional, not exact traffic counts.

TL;DR — LLM visibility is the aggregate of how often and how prominently a brand appears in AI-generated answers (AI Overviews, AI Mode, ChatGPT, Perplexity, Copilot, Gemini, Grok). It is not organic visibility — only ~38% of AI Overview citations come from the traditional top 10 (was 76%), and the strongest correlate is branded web mentions (0.664), not Domain Rating (0.326). It splits into three states — retrieved → mentioned → cited — that need different tools; citation-only measurement understates the picture badly. It behaves radically differently per platform, it’s volatile enough that single prompt runs aren’t data points, and the right unit of measurement is a defined prompt pool tracked as Share of Model Voice over time. Tools like Brand Radar give directional indicators, not exact traffic counts.

What LLM visibility is (and isn’t)

Measurements are conditional on prompts, model versions, time, locale, and sampling. Evidence for this claim ChatGPT search can answer with information from the web and provide linked sources. Scope: ChatGPT search; appearance in a sampled answer is product-, query-, locale-, and time-dependent. Confidence: high · Verified: OpenAI: ChatGPT search Provider crawler controls describe access policy, not guaranteed inclusion in answers. Evidence for this claim OpenAI distinguishes crawler controls for search inclusion from controls for model training. Scope: OpenAI's documented crawlers and controls; allowing a crawler does not guarantee retrieval, citation, or answer inclusion. Confidence: high · Verified: OpenAI: Crawlers

LLM visibility — AI visibility, same thing — is the rolled-up measure of how often and how prominently a brand, domain, or page shows up in the answers AI search systems generate. It’s a category, not a single number, the same way organic visibility is. The difference is the surface it measures: a synthesized, written answer instead of a ranked list of links.

The thing I’d burn into your brain first: LLM visibility is not the same as organic visibility. In mid-2025, ~76% of AI Overview citations came from pages ranking in the organic top 10. By early 2026 that’s down to roughly 38% in our data (and other studies put it even lower). Organic rank used to be a decent proxy for AI presence. It isn’t anymore. These are two different signals, and you need two different measurement programs.

Why the divergence? Because the signals that drive AI visibility aren’t the ones that drive rankings. In Ahrefs’ study of 75,000 brands by Louise Linehan and Xibeijia Guan — findings I later discussed at Ahrefs Evolve 2025 — the order came out like this for Google AI Overview visibility:

SignalCorrelation with AI visibility
Branded web mentions0.664
Branded anchor text0.527
Branded search volume0.392
Domain Rating0.326

The #1 signal is brand presence off your own site. That’s the structural inversion: unlinked mentions — which pass no PageRank and barely register in traditional SEO — are the strongest correlate of AI visibility. Which is exactly why I keep saying the job shifted from “optimize your site” to “optimize how the internet talks about you.” It lines up with the earned-media data too: 82–89% of AI citations come from third-party sources (Forbes/TechCrunch/WSJ-type outlets), not brand-owned pages. One concrete example from the talk — Zapier had 16 pages mentioning Ahrefs, and those 16 pages were cited across 1,431 Ahrefs AI responses. Other people’s content about you does most of the work.

Three states: retrieved → mentioned → cited → clicked

“AI visibility” hides three genuinely different states. Conflating them is where most measurement goes wrong:

StateWhat it meansHow you see it
RetrievedThe AI’s retrieval system fetched your page as source materialServer logs (ChatGPT-User, Perplexity-User, Googlebot-Extended)
MentionedYour brand name appears in the generated answer textBrand monitoring / string-match (Brand Radar); manual prompt testing
CitedYour URL is explicitly linked as a sourceBrand Radar Cited Pages/Domains; Bing WMT AI Performance; GSC AI features

And then a fourth, downstream of all three: clicked — the rare case where someone actually visits your site from the answer (GA4 / web analytics).

Most tools only see the cited state. The retrieved and mentioned states are largely invisible to citation-only tools — OtterlyAI’s read of Bing’s data put it bluntly: “99.6% of your AI influence is invisible.” So if your dashboard is counting citations and calling that your AI visibility, you’re understating it by a lot. And the reverse failure exists too — Superlines found ~73% of AI presence was citations without a brand mention (“ghost citations”). Run mention tracking and citation tracking together or you get a false picture either way.

A question underneath all three: is this even a retrieval-driven answer?

Retrieved, mentioned, and cited all assume the AI went and fetched something for this specific query. A meaningful share of AI answers don’t — the model answers from what it already learned during training (its parametric memory), with no live retrieval step at all. See RAG for the mechanics: production systems run a query classifier that decides, per query, whether to search — it isn’t a step that happens on every request.

This matters for measurement because it changes what’s actionable. A retrieval-driven answer can, in principle, be moved by on-page and off-page work — better content gets fetched, better mentions get pulled in. A memory-driven answer can’t be moved that way: the model already “learned” what it knows about your brand at training time, and no page edit changes that until the model is retrained — which happens on the provider’s schedule, not yours. If you’re seeing a stable brand description across many prompt runs with no citations and no sign of retrieval, you may be looking at a memory-driven answer. Track it, but don’t spend editing effort expecting it to move.

There’s no fully reliable way to tell which is which purely by reading the output — it takes deliberately testing for it: comparing a platform’s search-on vs. search-off behavior where that’s exposed, watching whether citations appear at all, or checking whether the answer changes when the underlying source page changes. Treat “no citations at all” and “citations present” as different measurement regimes with different remedies, not two scores on the same scale.

Platform by platform — why visibility varies so much

There is no single “AI visibility” number, because the platforms behave completely differently. They draw from different sources and cite at wildly different rates:

PlatformBehavior
PerplexityHeavy citer — ~21.87 citations per response; cites ~13% of the time
Google AI OverviewsHuge volume; RAG + query fan-out over the standard index
Bing Copilot~6.89 citations per response
ChatGPTSelective — cites in well under 1% of responses (~0.59%); ~87% of its citations align with Bing’s top organic results
GrokCites frequently — ~27% of the time
Claude / othersOften answers from training data with no retrieval at all

The headline gap: Perplexity cites ~21.87 sources per response while ChatGPT cites in 0.59% of responses. The same content can be highly visible on one platform and effectively invisible on another. And cross-platform overlap is poor — only 7 of the top 50 most-cited domains appear across all three major platforms (AI Overviews, ChatGPT, Perplexity). LLM visibility is really three (or seven) separate visibility profiles. Measure each platform; don’t average them into one number and pretend it means something.

It’s volatile — single runs aren’t data

AI answers churn. Roughly 40–60% of cited sources change month to month, and for the same query a large share of the AI Overview content changes between runs. SparkToro found less than a 1-in-100 chance of getting identical brand lists across 100 runs of the same ChatGPT prompt. The methodological consequence is simple and non-negotiable: running a prompt once is not a data point. You need a defined prompt pool, many runs, and trend windows — not point-in-time snapshots.

Repeated runs expose the distribution hidden by a snapshot: the same prompt can cite, merely mention, or omit a brand.

Across eight synthetic runs, the prompt 'best audit tools' is cited three times, mentioned three times, and absent twice. 'crawl budget help' is cited twice, mentioned three times, and absent three times. 'schema checker' is cited four times, mentioned twice, and absent twice. The fixture contains no live provider output or customer data.

How to measure it — the stack mapped to the states

No single tool covers all four states. Build a stack where each layer maps to one:

  • Layer 1 — server logs → Retrieved. The -User bots (ChatGPT-User, Perplexity-User, Googlebot-Extended) are live inference fetches — someone asked a question and the AI went to get your page right then. That’s different from training/indexing bots (GPTBot, ClaudeBot). Every error you serve a -User bot is a missed citation.
  • Layer 2 — brand monitoring → Mentioned. Ahrefs Brand Radar does string-match mentions across the major platforms and has a “found but not cited” filter for isolating influence-without-attribution. Add manual prompt testing for QA.
  • Layer 3 — citation tracking → Cited. Brand Radar’s Cited Pages / Cited Domains reports; Bing Webmaster Tools’ AI Performance report (the first official platform source of citation data — citations, grounding queries, average cited pages); and Google Search Console’s AI features filter / Gen AI Performance reports (impressions only — no clicks).
  • Layer 4 — web analytics → Clicked. GA4’s AI Assistant channel, plus a custom channel group. Expect a dark-traffic problem: a large share of AI-sourced visits arrive with no referrer and land in Direct, and AI Overview/AI Mode clicks merge into google/organic — so the clicked layer is the least clean.

One caveat I want to be honest about: Brand Radar (and tools like it) report directional indicators, not exact traffic counts. Treat the trend and the relative share seriously; don’t treat any single figure as a precise tally.

The prompt-pool method and Share of Model Voice

The unit of measurement is the prompt pool: define ~250–500 high-intent queries, run them across the LLM endpoints on a weekly or monthly cadence, record where your brand appears, and track Share of Model Voice (SOMV) — brand appearances ÷ total tracked prompts × 100. Pull prompts from real demand: sales transcripts, support tickets, Reddit, G2 reviews, PAA, autocomplete. Benchmarks to calibrate against: average mention rate is around 17.2% across relevant prompts (AthenaHQ), and 40–70% is considered strong.

For platform mechanics, Brand Radar runs on a 400M+ search-backed prompt corpus (up from 350M+ earlier this year — it’s a growing, live-updating index, so treat the exact figure as directional), refreshes monthly for ChatGPT/Perplexity/Gemini/Copilot and continuously for AI Overviews/AI Mode, and distinguishes “cited” (linked URLs) from “found but not cited” (string matches without links).

A measurement edge case worth knowing

AI doesn’t just under-report you — sometimes it invents you. We had thousands of visits going to pages on Ahrefs that didn’t actually exist because an AI synthesized plausible-looking URLs. Hallucinated citations are a real dimension of AI-visibility measurement, and most tools won’t filter them out for you — so sanity- check the URLs you’re “cited” on.

What this is and isn’t worth

Be clear-eyed about value. Even being cited rarely produces a click — users clicked a cited source in only 1% of Google visits with an AI summary in Pew’s March 2025 U.S. browsing study, and the vast majority of Google AI Mode sessions end with no site visit. LLM visibility is primarily a brand awareness and authority metric, not a direct-traffic metric. It still matters: a large share of AI Mode users accept the curated shortlist without doing further research, so being in the answer carries weight even without the click.

Where this fits

This is the definition-and-measurement hub for the cluster. For the deep dive on the three states, see Retrieved, Mentioned, Cited. For the GA4/analytics layer specifically — the dark-traffic and attribution mess — see AI traffic attribution. For the off-site signal that actually drives this (entities and mentions), see Entity SEO and the broader AI search hub.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin an expert quote first.