Entity SEO
How to make AI systems confidently identify and trust your brand — the Knowledge Graph, entity disambiguation, sameAs schema, and why branded mentions correlate with AI citations more strongly than backlinks do.
1 evidence signal on this page
- Related live toolEntity Coverage Analyzer
Entity SEO makes AI systems confidently identify who or what you are — Ahrefs research found branded mentions correlate with AI Overview citations more strongly than backlinks do, though correlation isn't proof either one causes citation.
TL;DR — An entity is a thing a machine can identify — you, your company, your product. Entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. is making sure search engines and AI tools know exactly who you are, instead of confusing you with something else. The basics: use the same name and description everywhere, consider a Wikidata entry (and Wikipedia if you qualify), and get mentioned on trustworthy sites. None of these are strict requirements — they’re signals that make it easier for AI to confidently identify and cite you.
What’s an entity?
An entity is an identifiable thing; Schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. provides identifiers and relationship properties that can clarify what a page describes. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Schema.org: Thing Structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. is one supporting signal and is not a guaranteed Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. or citation submission mechanism. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Structured data introduction
When you read “Taj Mahal,” you know from context whether I mean the monument, the blues musician, or the casino. A machine doesn’t — unless it has learned these are three separate entities, each with its own identity, type, and set of facts.
An entity is just a distinctly identifiable thing: a person, a place, a company, a product, even a concept. Google stores millions of them, with their relationships, in a database called the Knowledge Graph. Entity SEO is the work of making sure your entity — your brand, you as an author, your products — is in there, clearly, and not muddled up with anything else.
Why AI systems care so much about this
Traditional search matched the words you typed against the words on a page. AI search is different: it tries to understand who and what a query is about, then summarize an answer. To name your brand in that answer, the AI first has to be sure it knows which brand you are.
If your identity is fuzzy — same name as another company, different descriptions on every profile, no outside confirmation — AI hedges and skips you. If it’s crisp and well-confirmed, AI cites you with confidence. That’s the whole game.
Simple things that actually help
- Be consistent. Use the exact same business name, description, address, and phone number everywhere — your site, social profiles, directories, Google Business Profile. Inconsistency confuses the machine.
- Claim a Wikidata entry. Wikidata is a machine-readable database that AI systems and Google read directly. Almost any business can create an entry — no “notability” hurdle like Wikipedia has.
- Aim for Wikipedia eventually. A Wikipedia page is gold for AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals., but you can’t write your own — you need genuine press coverage first so editors deem you notable.
- Get mentioned by others. Being talked about on trustworthy third-party sites matters more than almost anything. AI learns who you are from how often, and where, your name comes up.
- Fill out your Google Business Profile if you’re a local business — it’s a well-documented Google local-data source, though it isn’t established that every AI system relies on it the same way for local answers.
The deeper mechanics — the Knowledge Graph, entity salience, the schema, and the research data behind “mentions beat links” — are in the Advanced tab.
TL;DR — Entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. is practitioner terminology, not a documented Google ranking system — but for AI the emphasis shifts hard toward disambiguation and corroboration. Google says the Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. shed a large share of its entities in a June 2025 cleanup; the exact scale and its effect on any specific AI product aren’t things I can independently verify here, so treat that as directional, not gospel. In Ahrefs research, branded web mentions and YouTube presence correlated with AI-answer visibility far more strongly than backlinks did — correlation, not proof of cause. You also can’t shortcut trust with fake corroboration: an independent experiment seeding fictional experts into hundreds of press articles produced essentially no AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking.. And treat
sameAsschema as most robust when it’s server-rendered — whether a given AI crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. executes JavaScript is provider- and version-specific, so don’t assume none of them do.
Entity SEO is just SEO — with a different emphasis for AI
Start with the bounded definition: “Entity SEO” is practitioner terminology, not the name of a documented Google ranking system. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Schema.org: Thing It’s a useful lens on existing SEO practice — identity clarity, corroboration, and accurate structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. — not a separate discipline or a universal AI optimization layer with its own scoring model. Use those practices as clarity signals, not guaranteed ranking levers. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Structured data introduction
My standing framing hasn’t changed for traditional search: “the entity identification part is more on Google’s end than ours.” You don’t manually “do entity SEO” by stuffing entities into a page. You publish quality content, mark it up with accurate structured data, keep your signals consistent, and Google does the identification.
What has changed is the weighting for AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. When a model writes an answer, it pulls sentences and names, not link equityPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. — so the brand-signal picture looks different from the link-driven world of rankings. The mechanics below are about making your entity unmistakable to that machine.
The Knowledge Graph, and what we can and can’t say about AI
Google launched the Knowledge Graph in 2012 with the pitch “things, not strings” — a model that “understands real-world entities and their relationships to one another.” It was built to solve disambiguation (which “Taj Mahal”?), summarization, and discovery, seeded from Freebase, Wikipedia, and the CIA World Factbook. Those 2012 numbers are historical; Google hasn’t published a current entity count.
Here’s where I want to be careful: it’s tempting to say “the Knowledge Graph now feeds AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., AI Mode, and Gemini, so being in it makes you eligible for AI-powered answers.” Google confirms the Knowledge Graph powers Search features like panels, drawing on hundreds of sources across the open web. But I don’t have product-specific primary documentation tying Knowledge Graph membership directly to eligibility for every AI surface, and I’m not going to assert a cross-product causal claim I can’t back with a citation. Treat “the KG is AI infrastructure” as a plausible industry read, not a confirmed mechanism.
What is publicly reported: journalist Jason Barnard (Kalicube) documented a large Knowledge Graph contraction in June 2025, with steep cuts concentrated in stale event entities and ambiguous “Thing”-category entities, alongside more precise person-entity typing. I’m not going to restate the exact percentages here — they come from one third-party analysis of Google’s data, not a source I can independently reproduce, and I’d rather point you to the original writeup than risk a stale or mistyped figure. The directional read holds up: as Barnard put it, “in the Knowledge Graph, clarity is the only point of entry.” Quality and disambiguation appear to matter more than raw signal volume — see the source for the full breakdown.
The entity signal hierarchy — directionally, not exactly
This is the part most “entity SEO” advice gets wrong: it treats loosely related signals as if they were proven ranking levers. We ran correlation studies at Ahrefs across large samples of brands and AI answers, and the directional pattern is worth knowing even though I’m holding back the exact coefficients here pending independent verification against the original datasets (correlation studies are easy to mis-cite, and a correlation is not causation regardless of the number).
Directionally, across the studies: branded web mentions and YouTube presence/mentions correlated with AI-visibility metrics (like AI Overview brand mentions) noticeably more strongly than backlinks did — brand-mention and YouTube signals outranked link-based signals like Domain Rating and referring domains in the rankings we measured. Branded search volume also behaved as a leading indicator rather than a lagging one. As I summed up in my cross-platform AI brand-visibility correlations study: “Google favors brands and Perplexity seems to show some favoritism as well. What’s surprising is the weakness of ChatGPT here.”
We also foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. meaningfully limited overlap between which sources get cited across Google, ChatGPT, and Perplexity — a small minority of sources showed up across all three simultaneously. Entity SEO is a multi-platform exercise, not a single-engine one. For the exact coefficients and sample sizes, see the linked study and the AI Overview brand-correlation study directly rather than relying on numbers restated here.
Disambiguation is the #1 priority
AI cannot cite an entity it cannot confidently identify. If your brand shares a name with another company, has inconsistent NAP (name/address/phone) data, or has no external verification, you get skipped. The June 2025 cleanup’s whole point was raising the share of unambiguously typed entities — Google is optimizing for the brands it can be sure about.
So disambiguation is the job: a consistent canonical identity, declared once on your
entity home (About page or homepage) via Organization JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. with a stable
@id, and confirmed against external references with sameAs. John Mueller’s one
caveat: don’t point sameAs at a Knowledge Graph ID URL (kg:/m/...) because
“that ID might change” — use stable URLs like Wikidata, Wikipedia, and social
profiles instead.
You can’t fake corroboration — the entities.org experiment
Here’s the most important guardrail in the whole field. Researchers at Entities.org describe AI systems wanting entity consensus — independent, high-confidence sources agreeing about the same claim before they’ll assert it. Below a threshold of independent corroboration, AI hedges; above it, AI asserts.
The headline finding: a controlled experiment seeded fictional experts into a large number of press articles and measured the result across nine models — essentially no AI recommendations resulted. Volume without independence fails; the researchers’ framing is that AI systems can detect synthetic consensus. I’m not restating their exact article count, model-recommendation threshold, or citation-share figures here — those are specific numbers from one third-party research project that I haven’t independently reproduced, so check the original research for the precise counts. What I’m comfortable asserting directionally: self-published content alone rarely carries a brand into an AI answer, and most AI answers appear to lean on third-party sources rather than brand-published ones.
This is the deep reason “entity stacking,” mass press-release seeding, and buy-in-bulk link campaigns don’t work for AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.. There’s no hidden entity score to game — earn the real mentions.
sameAs, Wikidata, and Wikipedia
Neither Wikipedia nor Wikidata is a documented requirement for Google or AI
recognition — Google names Wikipedia as one of many common sources it can draw on
for knowledge panels, not a mandatory input. That said, they’re useful, accessible
external references worth prioritizing in your sameAs list, roughly in this order:
- Wikidata — machine-readable and low barrier to entry (no Wikipedia-style notability gate); several AI systems and Google are reported to draw on it, though I don’t have primary documentation for exactly how every provider weights it.
- Wikipedia — if you’re notable. Third-party research cites it as a large citation source for some AI systems, though I’m not restating the exact percentage here — check the source research for the current figure and its methodology. Wikipedia content is also used for real-time groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. by some systems, not just training. You can’t write your own page — earn the press coverage that establishes notability first.
- LinkedIn / Crunchbase / official registrations, then social profiles.
One cautionary tale from our own team: Ryan Law accidentally led Google to think he
owned the Ahrefs website by adding schema to his personal site with an error in the
sameAs property. Always double-check schema before you push it live.
Structured data: indirect lever, not a citation button
Be clear-eyed about what schema does. An Ahrefs observational study by Louise Linehan and Xibeijia Guan tracked a large set of pages that added JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. over roughly seven months and compared citation volume before and after — see the schema/AI citations study for the dataset and exact percentages. The directional finding: adding schema to an already-cited page did not produce a clear positive lift in AI citation volume across the platforms tracked.
So treat schema as not a reliable citation lever. Its value is entity disambiguation — connecting your brand to its Knowledge Graph identity — and correct indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., not a direct bump in how often AI quotes you. Mueller’s own position lines up: structured data is not a direct ranking factor, though it helps machines understand entities and is genuinely useful for shopping data that’s “basically impossible to read in high fidelity… from a text page.”
For AI specifically, @id on every entity and an @graph connecting Organization,
Person, and Article into one graph is a useful modeling pattern practitioners
recommend for keeping a connected identity consistent — it isn’t a documented Google
or AI-provider requirement, and using it doesn’t guarantee a ranking, panel, or
citation effect. Isolated schema blocks for the same entities work too; a connected
graph is about making maintenance and consistency easier, not unlocking a feature.
Server-render your schema — don’t assume any crawler runs your JS
This is a technical mistake that can silently sink otherwise-good entity work, and it’s worth being precise about what’s actually documented rather than repeating a blanket claim. Google explicitly documents renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. JavaScript-generated structured data for Search. Whether a given AI crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — GPTBot, ClaudeBot, PerplexityBot, or any other named provider — executes JavaScript is provider- and version-specific, and it changes over time; I don’t have current, per-provider primary documentation to assert a universal “AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. don’t run JavaScript” rule, and I’d rather send you to verify the current behavior yourself (a dated fetch test against the provider you care about) than repeat an unverified generalization.
The safe, provider-agnostic move either way: treat server-rendered schema in the
initial HTML as the robust baseline. If a crawler you’re targeting turns out not to
render JavaScript, only server-delivered schema is guaranteed visible to it; if it
does render JS, server-rendering costs you nothing. Check your sameAs and @graph
in the raw HTML response (not just the rendered DOM) so you know what’s actually
being served before any rendering step.
So what do you actually do?
This is practitioner guidance — a way I find useful to organize the work, not a documented AI-system scoring model. I think about it as three layers: identity, relationships, and independently verified corroboration.
Identity. Nail your entity home — one canonical page, Organization/Person
schema server-rendered as the robust baseline, and accurate sameAs. Keep NAP
(name/address/phone) and descriptions consistent everywhere, including your Google
Business Profile if you’re local — a well-documented Google data source, even though
I can’t confirm it’s the primary local source for every AI system.
Relationships. Connect your entities — sameAs to Wikidata (and Wikipedia if
you’re notable), an @graph linking Organization, Person, and Article where it
helps you keep things consistent. In your content, cover the entities and
relationships your readers actually need explained — don’t chase a Cloud Natural
Language salience score as if it were a ranking input; salience describes how a
document-analysis API scores a piece of text, not a documented Google Search or
AI-citation ranking factor.
Corroboration. The layer that moves the needle most in my experience — earn genuine third-party mentions and a real YouTube presence, because those (not links alone) are strong signals of whether AI has independent, corroborated confidence you’re who you say you are. As I’ve said about the AI era, “SEO itself hasn’t drastically changed, but getting good results may now require closer collaboration with other teams like PR and partnerships.”
Turn a missing entity lookup into a concrete identity and corroboration checklist with my free Google Knowledge Graph Explorer Free
- Look up the canonical person, organization, or brand name you use publicly.
- Treat a no-result state as one lookup observation—not proof that every Google system lacks entity understanding.
- Work through the entity-home, accurate sameAs, and independent-corroboration playbook before checking again later.
The result is labeled Not found and says Acme Analytics Test Brand is not in the returned Knowledge Graph results yet. Its playbook recommends using one canonical name, publishing an unambiguous entity home page, adding accurate Person or Organization markup with sameAs links, building corroboration in appropriate independent sources, and rechecking later. This lookup does not prove that every Google system lacks entity understanding and does not predict Knowledge Panel eligibility.
AI summary
A condensed take on the Advanced version:
- Entity = a thing a machine can uniquely identify (brand, person, product). Entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. makes that identification confident and unambiguous so AI can cite you.
- Entity SEO is practitioner terminology, not a named Google ranking system — “the entity identification part is more on Google’s end than ours” — but for AI the weighting shifts toward disambiguation + corroboration.
- The Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. reportedly went through a large clarity cleanup in June 2025 (per third-party reporting, not restated here in exact figures). Google confirms the Knowledge Graph feeds Search features like panels; a direct, cross-product link to every AI surface’s eligibility isn’t something I can back with primary documentation, so it’s treated as a plausible read, not confirmed fact.
- Signal hierarchy for AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals. (directional, not exact): in Ahrefs research, branded web mentions and YouTube presence correlated with AI-visibility metrics noticeably more strongly than backlinks did — correlation, not causation. Exact coefficients are in the linked studies, not restated here pending independent verification.
- You can’t fake it: an independent experiment seeding fictional experts into press articles foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. essentially no resulting AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking.. Independently corroborated sources matter; most AI answers appear to lean on third-party sources rather than brand-published ones.
- Schema is an indirect lever: an observational study of pages adding JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. found no clear positive lift in AI citation volume. Schema’s value is disambiguation and correct indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., not a citation button.
- Don’t overclaim the JS trap: whether a given AI crawlerAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. executes JavaScript is provider-specific and changes over time — there’s no verified universal rule. Server-renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. schema is still the robust baseline regardless.
- Do: entity home + server-rendered
Organization/Personschema as the robust default, accuratesameAsto Wikidata/Wikipedia (useful, not required), consistent NAP + Google Business Profile, and earn real third-party mentions + YouTube presence.
Official documentation
Primary-source documentation on entities, the Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself., and entity structured data.
Google — Knowledge Graph & entities
- Introducing the Knowledge Graph: things, not strings (2012) — Amit Singhal’s launch post; the “things, not strings” framing.
- Knowledge Graph Search API — “lets you find entities in the Google Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself.”; returns schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal., 20+ entity types, KG IDs (
kg:/m/...), and aresultScore. - Cloud Natural Language API — Entity — the official entity definition with
name,type,salience(0–1.0),metadata,mentions, andsentiment. - Entity Salience Task (Google Research, EACL 2014) — the foundational paper establishing entity salience as a distinct NLP task.
Google — entity structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding.
- Organization structured data — no required properties; disambiguate via
sameAs,iso6523Code,naics,legalName. - Profile Page structured data —
ProfilePagewithmainEntity(Person/Organization) for first-hand creator profiles. - How knowledge panels are created — panels are “created automatically… when there is enough information available on the open web.” No manual submission.
Microsoft / Bing
- Bing structured data / markup support — schema for products, events, recipes, articles, videos; Copilot can incorporate knowledge-graph and structured-data signals when available.
Quotes from the source
On-the-record statements from search representatives and Google’s own materials.
Google — the Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself.
- “an intelligent model—in geek-speak, a ‘graph’—that understands real-world entities and their relationships to one another: things, not strings.” — Amit Singhal, then-SVP Engineering, Google (2012). Source
Google Cloud Natural Language — what an entity is
- An entity is “a phrase in the text that is a known entity, such as a person, an organization, or location,” assigned a
saliencescore for its “importance or centrality… to the entire document text.” — Google Cloud Natural Language API reference. Source
John Mueller, Google — sameAs and KG IDs
- Mueller advises against using Knowledge Graph ID URLs for
sameAsbecause “that ID might change” — prefer stable URLs (Wikipedia, Wikidata, social profiles). Coverage
John Mueller, Google — schema and LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4).
- “Some features thrive with structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding.… Pricing, shipping, availability for shopping is basically impossible to read in high fidelity & accurately from a text page, for example.” (Mueller has repeatedly noted structured data is not a direct ranking factor.) Coverage
Patrick Stox, Ahrefs — entity identification is Google’s job
- “The entity identification part is more on Google’s end than on our end.” Source
Entity SEO checklist
A pass to make your entity unmistakable to search engines and AI:
- Entity home chosen (homepage or About page) with complete
OrganizationJSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal.:name,url,logo,description, and a stable@id. - Schema is server-rendered in the raw HTML rather than injected only by Google Tag Manager or client-side JS — the robust baseline regardless of which crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. you believe render JavaScriptMaking sure search engines can crawl, render, and index content that depends on JavaScript., since that’s provider- and version-specific and not something to assume either way.
-
sameAspoints to stable external references (Wikidata, Wikipedia, LinkedIn, Crunchbase, socials) — never a Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. ID URL. - Wikidata entry created/claimed where relevant (low barrier; a useful, not required, reference several systems are reported to draw on).
- Wikipedia pursued only once genuine notability (independent coverage) exists — never self-write a promotional page.
- Descriptions and NAP (name/address/phone) are identical across every profile, directory, and your site.
- Google Business Profile fully completed if you’re a local business — a well-documented Google local-data source worth getting right.
-
Person/ProfilePageschema for authors, withsameAs,jobTitle,affiliation, linked into an@graphwithOrganizationandArticle. - Active plan to earn high-authority third-party mentions (PR, partnerships, editorial) — the strongest AI-visibility lever, not link volume.
- YouTube presence in place — one of the strongest off-site signals we’ve measured in Ahrefs correlation research.
- Schema validated before publish (double-check
sameAs,@id— one typo can misattribute ownership).
Entity signals — cheat sheet
Signals ranked directionally by correlation with AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals. (Ahrefs studies — exact coefficients withheld here pending independent verification; see the linked studies for figures)
| Signal | Relative strength (directional) | Traditional SEO value | AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. value |
|---|---|---|---|
| YouTube presence | Strongest off-site correlate measured | Indirect (referral/brand) | Strongest off-site AI correlate across platforms |
| Branded web mentions | Very strong | Trust / co-citation | Top-ranked predictor of AI OverviewAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. brand mentions |
| Branded anchors | Strong | Anchor relevance | Strong AI correlate |
| Branded search volume | Moderate | Lagging demand indicator | Leading AI visibility indicator |
| Domain Rating | Moderate | Core authority metric | Moderate |
| Referring domains | Moderate | Core ranking signal | Moderate |
| Backlinks | Weakest of this set | Core ranking signal | Notably weaker correlate than mentions |
Mentions vs. links: branded web mentions correlated with AI Overview visibility meaningfully more strongly than backlinks did in the underlying study — correlation, not causation, and the exact ratio is in the source study rather than restated here.
Disambiguation quick rules
- Entity home +
OrganizationJSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. with stable@id. sameAs→ Wikidata, Wikipedia, LinkedIn, socials (useful references, not requirements). Neverkg:/m/...(it changes).- Identical name / description / NAP everywhere.
Hard truths
- Schema ≠ citations. An observational study of pages adding JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. no clear positive lift in AI citationAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. volume. Schema’s job is disambiguation, not citation-getting.
- Can’t fake consensus. An independent experiment seeding fictional experts into press coverage found essentially no resulting AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking.. Independent corroboration matters; self-published content alone rarely carries a brand into an AI answer.
- Don’t assume “AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. don’t run JS” as a universal rule. It’s provider-specific and changes over time. Server-render schema as the robust default regardless.
- Platforms diverge. Cross-platform citation overlap (Google, ChatGPT, Perplexity) is limited — treat entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. as a multi-platform exercise.
The entity authority framework
This is a practitioner framework — a useful way I organize entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. work, not a documented AI-system scoring model or a set of ranking requirements. It borrows a three-part structure practitioners in this space describe: Recognition → Relationships → Corroboration. (Entity authority is “the degree to which search systems recognize your brand as a credible, well-corroborated source on a specific entity” — a practitioner definition, not a Google term.)
1. Recognition — can the machine identify you at all?
The disambiguation layer. Is your identity declared once on your entity home with a
stable @id, and is your name / description / NAP consistent everywhere? Ambiguous
or inconsistent signals plausibly make it harder for AI systems to confidently cite
an entity, though I don’t have a documented mechanism guaranteeing that connection
for every provider.
2. Relationships — does it understand how you connect?
The graph layer. sameAs ties your schema to authoritative external records as a
useful practitioner pattern (not a requirement); @graph linking Organization,
Person (founders/authors), and Article entities is a modeling convenience for
keeping them consistent; consistent co-occurrence ties your brand name to your topic
in the content itself.
3. Corroboration — do independent sources vouch for you? The trust layer, and the one you can’t shortcut. Third-party research suggests AI systems weigh independent, high-confidence sources before asserting a claim confidently, and that self-published content alone rarely carries a brand into an AI answer — so earned media, YouTube presence, and genuine brand demand do real work here. Synthetic consensus (seeded press, bulk links) appears to get detected and discarded rather than rewarded, per the fictional-expert experiment referenced above in the Advanced lens.
The decision rule. Stuck on AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.? Diagnose which layer is failing:
- Not recognized → fix disambiguation (entity home, consistency, Wikidata).
- Recognized but not connected → fix relationships (
sameAs,@graph, topical co-occurrence). - Recognized and connected but not cited → fix corroboration (earn independent mentions).
Don’t add schema and hope. Find the broken layer.
Entity SEO mistakes to avoid
Treating sameAs as proof
Markup can state an identity, but it cannot manufacture third-party corroboration.
Use sameAs to connect profiles that genuinely represent the same entity, then make
the names, descriptions, and facts on those profiles consistent.
Creating identifiers for the wrong entity
A similarly named company, person, or product is not a shortcut. Confirm the entity’s attributes and relationships before linking Wikidata, Wikipedia, social, or directory profiles.
Publishing entity signals only through client-side JavaScript
Whether a given AI crawlerAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. renders JavaScript is provider-specific and not something to assume either way. Put the identity statement, core facts, links, and structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. in server-delivered HTML so you’re covered regardless.
Optimizing mentions without fixing ambiguity
More coverage can reinforce the wrong interpretation when the brand name is shared. Lead with a stable name, category, location or market, official URL, and distinguishing relationships.
Common entity SEO problems
Search results confuse the brand with another entity
Symptom: Panels, summaries, or citations mix namesakes. Likely cause: Weak or conflicting disambiguation signals. Fix: Align the official site, organization markup, authoritative profiles, and third-party descriptions around the same defining facts; then recheck the exact ambiguous query.
A Knowledge Panel exists but contains a wrong fact
Symptom: The panel identifies the right entity but shows outdated or incorrect information. Likely cause: Google selected a conflicting source. Fix: Claim the panel when eligible, correct the first-party source, suggest an edit with evidence, and repair inconsistent corroborating profiles.
Structured data validates but entity understanding does not improve
Symptom: Schema tools pass while the brand remains unrecognized or confused. Likely cause: Validation proves syntax, not real-world identity or authority. Fix: Audit independent mentions and identity consistency rather than adding more properties without evidence.
Prompts for entity audits
Audit these first-party and third-party descriptions of one organization. Extract the
name, aliases, category, location or market, founding facts, people, products, official
URL, and sameAs identifiers from each source. Return a contradiction table, likely
namesake collisions, facts supported by multiple independent sources, and facts that
must not be asserted yet. Do not merge similarly named entities without evidence.
[paste source excerpts and URLs]Review this Organization JSON-LD against the visible page content and the supplied
official profiles. Flag unsupported claims, wrong entity links, duplicate identifiers,
and useful missing disambiguation properties. Return corrected JSON-LD using only facts
present in the inputs, followed by a verification checklist.
[paste JSON-LD, page excerpt, and profile list] Entity signal extraction snippets
List sameAs values in Chrome DevTools
Run in the Console on the entity’s official page:
[...document.querySelectorAll('script[type="application/ld+json"]')].flatMap(el => {
try { const data = JSON.parse(el.textContent); return (Array.isArray(data) ? data : [data]).flatMap(item => item.sameAs || []); }
catch { return []; }
})Find identity links in a crawl extraction
Use this case-insensitive regular expression against JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. source. Capture group 1
contains the sameAs array for review:
"sameAs"\s*:\s*(\[[\s\S]*?\])Compare server HTML with the rendered DOM
Fetch the response before relying on browser output:
curl -sS https://example.com/about/ | grep -Eio 'application/ld\+json|sameAs|Organization'The check is deliberately simple: if identity signals appear only after renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., move them into the server-delivered document.
Tools for entity SEO
- Entity Coverage Analyzer: Inventory the entities and relationships a page states clearly, then find important gaps.
- Google Knowledge Graph Explorer: Check whether Google resolves a name to a distinct entity and inspect identifiers without treating a result as a guarantee of a panel.
- Schema Markup Validator: Validate Organization,
Person, Product, and
sameAsmarkup after confirming the underlying facts. - Search results and Knowledge PanelsThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself.: Test branded and ambiguous queries in clean sessions to see which entity interpretation surfaces.
- A crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. with custom extraction: Compare entity markup, names, official URLs, and profile links across templates at scale.
Validate an entity SEO change
Test server-visible identity signals
Test to run: Fetch the page without JavaScript and inspect the identity statement, official links, and JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal.. Expected result: The defining facts and structured data are present in the response. Failure interpretation: AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. that do not render cannot receive the change. Monitoring window: Immediate. Rollback trigger: A release moves core identity signals behind client renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM..
Test identifier consistency
Test to run: Compare the official site’s sameAs values and key facts with every
linked profile. Expected result: Each URL represents the same entity and agrees on
the stable defining facts. Failure interpretation: The graph contains conflicting
or false edges. Monitoring window: Immediate after profile or schema changes.
Rollback trigger: Any identifier resolves to a namesake or unrelated entity.
Test ambiguous-query disambiguation
Test to run: Recheck the specific branded query that previously mixed entities, recording the result and cited sources. Expected result: The intended entity is distinguishable by category, market, URL, and relationships. Failure interpretation: Corroboration remains insufficient or contradictory. Monitoring window: After recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. and source refresh; not immediate. Rollback trigger: New copy increases confusion or asserts a fact the source set does not support.
Measure entity clarity
Entity consistency rate
Metric: Share of audited first- and third-party profiles agreeing on the entity’s stable defining facts. What it tells you: Whether systems receive one coherent identity. How to pull it: Maintain a source inventory and compare name, category, official URL, location or market, and key relationships. Benchmark / realistic range: Target no material contradictions; the number of profiles needed depends on the entity. Cadence: Quarterly and after rebrands.
Ambiguous-query accuracy
Metric: Share of a fixed branded-query set that resolves to the intended entity without namesake mixing. What it tells you: Whether disambiguation works in the surfaces that matter. How to pull it: Run a documented query set across Search and selected AI systems, saving outputs and citations. Benchmark / realistic range: Establish a baseline per query; systems vary and no universal percentage is defensible. Cadence: Monthly.
Corroborated fact coverage
Metric: Important entity facts supported by the official source plus at least one appropriate independent source. What it tells you: Which claims have evidence beyond self-assertion. How to pull it: A fact-to-source matrix. Benchmark / realistic range: Prioritize complete support for identity-defining facts rather than maximizing every optional property. Cadence: Quarterly.
Test yourself: Entity SEO
Resources worth your time
My related writing & research (Ahrefs)
- AI Overview brand-visibility correlations (75K brands) — the source study for the mentions-vs.-backlinks correlation numbers referenced above.
- AI brand-visibility correlations across all three platforms — the YouTube-signal correlation and platform divergence, with exact figures.
- Do mentions on highly linked pages influence AI mentions? — the Brand-Radar study behind “Google favors brands.”
- Semantic SEO (Despina Gavoyannis, reviewed by me) — the topical-authority side of entity work.
My speaking
- GEO / AEO / LLMO — What’s With All This AI Stuff (Ahrefs Evolve 2025) — my deck on AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity., the data, and what changes. (Standing disclaimer applies: this is my understanding of these systems, not gospel.)
From others
- Schema markup and AI citations (Louise Linehan & Xibeijia Guan, Ahrefs) — the JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. study showing schema isn’t a citation lever, with the dataset and exact percentages.
- Entity SEO: Stop Overcomplicating Things (Ahrefs, Si Quan Ong) — the “entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. is just SEO” position; a colleague’s guide quoting Patrick.
- Google Knowledge Graph explained and Schema markup guide (Ahrefs, Despina Gavoyannis and colleagues) — the foundations (and Ryan Law’s
sameAscautionary tale). - Google’s great clarity cleanup: the Knowledge Graph and the AI future — Jason Barnard (Kalicube) on the June 2025 Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. contraction, with the exact figures.
- Breaking content & SEO silos to build entity authority in AI search — Lang Ploszek (Victorious) on Recognition / Relationships / Corroboration.
- Entity consensus research — the “volume without independence fails” experiment on synthetic consensus, with the study’s methodology and figures.
- Entity authority and AI search visibility (Search Engine Land, Benu Aggarwal) — practical breakdown of how entity authority translates into AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. presence.
- Schema markup and AI search — no hype (Search Engine Land) — sober look at what structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. actually does (and doesn’t do) for AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking..
- What we know about the impact of Wikipedia on ChatGPT search results (ALLMO) — Wikipedia’s reported share of ChatGPT citations and Wikimedia’s enterprise licensing deals with AI providers, with the source figures.
- 2025 AI citation & LLM visibility report (Digital Bloom) — platform-by-platform citation patterns: Google AI OverviewAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. vs. ChatGPT vs. Perplexity, and cross-platform overlap findings.
- sameAs versus knowsAbout in schema.org (Will Scott) — technical distinction between declaring identity equivalence vs. topical expertise in schema.
Entity SEO
Entity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence.
Related: Structured Data, AI Search Optimization
Entity SEO
An entity is any distinctly identifiable thing — tangible (a person, place, organization, or product) or intangible (a concept, event, or color) — that can be uniquely told apart from everything else. Entities have names, types, attributes, and relationships to one another. Google’s Cloud Natural Language API defines an entity as “a phrase in the text that is a known entity, such as a person, an organization, or location” and gives it a salience score (0–1.0) for how central it is to a document.
Entity SEO shifts the question from “what keywords should I rank for?” to “does the machine know with confidence who or what I am, what I do, and why I’m trustworthy?” In practice that means consistent naming and descriptions, an “entity home” (your About page or homepage) with complete Organization/Person structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding., sameAs links to authoritative references like Wikidata and Wikipedia, and — most of all — earned third-party mentions that corroborate your identity.
For traditional search, entity SEO is really just good SEO; as Patrick puts it, “the entity identification part is more on Google’s end than ours.” For AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. the emphasis changes: in Ahrefs research, branded web mentions correlated with AI OverviewAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. visibility noticeably more strongly than backlinks did (a correlation, not proof of cause — see the source study for exact figures). And you can’t fake your way in — an independent experiment that seeded fictional experts across a large set of press articles foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. essentially no resulting AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking., because the corroboration wasn’t genuinely independent.
Related: Structured Data, AI Search Optimization
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 21, 2026.
Editorial summary and recorded change details.Summary
Corrected a research misattribution: the schema/AI-citations observational study is by Louise Linehan and Xibeijia Guan, not Patrick — reworded the body prose that said 'we ran'/'I'm holding back' and moved the study out of 'My related writing & research (Ahrefs)' into 'From others' with author attribution.
Change details
-
Reworded the 'Structured data: indirect lever' paragraph to attribute the JSON-LD/AI-citations observational study to Louise Linehan & Xibeijia Guan (Ahrefs) instead of the first-person 'we ran an observational study'/'I'm holding back the exact page count' framing.
-
Moved the 'Schema markup and AI citations' link from 'My related writing & research (Ahrefs)' to 'From others', attributed to Louise Linehan & Xibeijia Guan.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 21, 2026.
Editorial summary and recorded change details.Summary
Separated Ahrefs articles by other authors from Patrick's own writing in the resources lens.
Change details
-
Moved the 'Entity SEO' guide (Si Quan Ong) and the Knowledge Graph/Schema markup guides (Despina Gavoyannis and colleagues) from 'My related writing & research (Ahrefs)' to 'From others', since colleagues authored them rather than Patrick.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 17, 2026.
Editorial summary and recorded change details.Summary
Bounded a series of unverified statistics and universal platform claims per an editorial research-correction addendum — kept the practical guidance, removed or qualified figures and claims that outran the available evidence.
Change details
-
Removed exact figures for the Knowledge Graph contraction, AI-visibility correlation coefficients, the synthetic-consensus experiment, Wikipedia's citation share, and the schema/AI-citation study, replacing them with directional summaries and links to the original sources for anyone who wants the precise numbers.
-
Corrected the 'AI crawlers don't run JavaScript' claim (repeated across the Advanced lens, checklist, cheat sheet, and anti-patterns) to note that JavaScript rendering is provider- and version-specific, not a universal rule — server-rendering schema stays the recommended default regardless.
-
Removed the claims that Knowledge Graph membership guarantees AI-answer eligibility, that Wikipedia/Wikidata/sameAs/an entity home are required, and that Google Business Profile is every AI system's main local source — reframed each as a useful signal, not a requirement.
-
Relabeled the Recognition/Relationships/Corroboration structure in the 'So what do you actually do?' section and the Frameworks lens as practitioner guidance, not a documented AI-system scoring model.
Full comparison unavailable — no prior snapshot was archived for this revision.