Entity SEO

How to make AI systems confidently identify and trust your brand — the Knowledge Graph, entity disambiguation, sameAs schema, and why branded mentions correlate with AI citations more strongly than backlinks do.

First published: Jun 24, 2026 · Last updated: Jul 21, 2026 · Advanced
demand #5 in Optimization#15 in AI Search#151 on the site
1 evidence signal on this page

Entity SEO makes AI systems confidently identify who or what you are — Ahrefs research found branded mentions correlate with AI Overview citations more strongly than backlinks do, though correlation isn't proof either one causes citation.

TL;DR — Entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence. is practitioner terminology, not a documented Google ranking system — but for AI the emphasis shifts hard toward disambiguation and corroboration. Google says the Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. shed a large share of its entities in a June 2025 cleanup; the exact scale and its effect on any specific AI product aren’t things I can independently verify here, so treat that as directional, not gospel. In Ahrefs research, branded web mentions and YouTube presence correlated with AI-answer visibility far more strongly than backlinks did — correlation, not proof of cause. You also can’t shortcut trust with fake corroboration: an independent experiment seeding fictional experts into hundreds of press articles produced essentially no AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking.. And treat sameAs schema as most robust when it’s server-rendered — whether a given AI crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. executes JavaScript is provider- and version-specific, so don’t assume none of them do.

Entity SEO is just SEO — with a different emphasis for AI

Start with the bounded definition: “Entity SEO” is practitioner terminology, not the name of a documented Google ranking system. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Schema.org: Thing It’s a useful lens on existing SEO practice — identity clarity, corroboration, and accurate structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. — not a separate discipline or a universal AI optimization layer with its own scoring model. Use those practices as clarity signals, not guaranteed ranking levers. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Structured data introduction

My standing framing hasn’t changed for traditional search: “the entity identification part is more on Google’s end than ours.” You don’t manually “do entity SEO” by stuffing entities into a page. You publish quality content, mark it up with accurate structured data, keep your signals consistent, and Google does the identification.

What has changed is the weighting for AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. When a model writes an answer, it pulls sentences and names, not link equityPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. — so the brand-signal picture looks different from the link-driven world of rankings. The mechanics below are about making your entity unmistakable to that machine.

The Knowledge Graph, and what we can and can’t say about AI

Google launched the Knowledge Graph in 2012 with the pitch “things, not strings” — a model that “understands real-world entities and their relationships to one another.” It was built to solve disambiguation (which “Taj Mahal”?), summarization, and discovery, seeded from Freebase, Wikipedia, and the CIA World Factbook. Those 2012 numbers are historical; Google hasn’t published a current entity count.

Here’s where I want to be careful: it’s tempting to say “the Knowledge Graph now feeds AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., AI Mode, and Gemini, so being in it makes you eligible for AI-powered answers.” Google confirms the Knowledge Graph powers Search features like panels, drawing on hundreds of sources across the open web. But I don’t have product-specific primary documentation tying Knowledge Graph membership directly to eligibility for every AI surface, and I’m not going to assert a cross-product causal claim I can’t back with a citation. Treat “the KG is AI infrastructure” as a plausible industry read, not a confirmed mechanism.

What is publicly reported: journalist Jason Barnard (Kalicube) documented a large Knowledge Graph contraction in June 2025, with steep cuts concentrated in stale event entities and ambiguous “Thing”-category entities, alongside more precise person-entity typing. I’m not going to restate the exact percentages here — they come from one third-party analysis of Google’s data, not a source I can independently reproduce, and I’d rather point you to the original writeup than risk a stale or mistyped figure. The directional read holds up: as Barnard put it, “in the Knowledge Graph, clarity is the only point of entry.” Quality and disambiguation appear to matter more than raw signal volume — see the source for the full breakdown.

The entity signal hierarchy — directionally, not exactly

This is the part most “entity SEO” advice gets wrong: it treats loosely related signals as if they were proven ranking levers. We ran correlation studies at Ahrefs across large samples of brands and AI answers, and the directional pattern is worth knowing even though I’m holding back the exact coefficients here pending independent verification against the original datasets (correlation studies are easy to mis-cite, and a correlation is not causation regardless of the number).

Directionally, across the studies: branded web mentions and YouTube presence/mentions correlated with AI-visibility metrics (like AI Overview brand mentions) noticeably more strongly than backlinks did — brand-mention and YouTube signals outranked link-based signals like Domain Rating and referring domains in the rankings we measured. Branded search volume also behaved as a leading indicator rather than a lagging one. As I summed up in my cross-platform AI brand-visibility correlations study: “Google favors brands and Perplexity seems to show some favoritism as well. What’s surprising is the weakness of ChatGPT here.”

We also foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. meaningfully limited overlap between which sources get cited across Google, ChatGPT, and Perplexity — a small minority of sources showed up across all three simultaneously. Entity SEO is a multi-platform exercise, not a single-engine one. For the exact coefficients and sample sizes, see the linked study and the AI Overview brand-correlation study directly rather than relying on numbers restated here.

Disambiguation is the #1 priority

AI cannot cite an entity it cannot confidently identify. If your brand shares a name with another company, has inconsistent NAP (name/address/phone) data, or has no external verification, you get skipped. The June 2025 cleanup’s whole point was raising the share of unambiguously typed entities — Google is optimizing for the brands it can be sure about.

So disambiguation is the job: a consistent canonical identity, declared once on your entity home (About page or homepage) via Organization JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. with a stable @id, and confirmed against external references with sameAs. John Mueller’s one caveat: don’t point sameAs at a Knowledge Graph ID URL (kg:/m/...) because “that ID might change” — use stable URLs like Wikidata, Wikipedia, and social profiles instead.

You can’t fake corroboration — the entities.org experiment

Here’s the most important guardrail in the whole field. Researchers at Entities.org describe AI systems wanting entity consensus — independent, high-confidence sources agreeing about the same claim before they’ll assert it. Below a threshold of independent corroboration, AI hedges; above it, AI asserts.

The headline finding: a controlled experiment seeded fictional experts into a large number of press articles and measured the result across nine models — essentially no AI recommendations resulted. Volume without independence fails; the researchers’ framing is that AI systems can detect synthetic consensus. I’m not restating their exact article count, model-recommendation threshold, or citation-share figures here — those are specific numbers from one third-party research project that I haven’t independently reproduced, so check the original research for the precise counts. What I’m comfortable asserting directionally: self-published content alone rarely carries a brand into an AI answer, and most AI answers appear to lean on third-party sources rather than brand-published ones.

This is the deep reason “entity stacking,” mass press-release seeding, and buy-in-bulk link campaigns don’t work for AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.. There’s no hidden entity score to game — earn the real mentions.

sameAs, Wikidata, and Wikipedia

Neither Wikipedia nor Wikidata is a documented requirement for Google or AI recognition — Google names Wikipedia as one of many common sources it can draw on for knowledge panels, not a mandatory input. That said, they’re useful, accessible external references worth prioritizing in your sameAs list, roughly in this order:

  1. Wikidata — machine-readable and low barrier to entry (no Wikipedia-style notability gate); several AI systems and Google are reported to draw on it, though I don’t have primary documentation for exactly how every provider weights it.
  2. Wikipedia — if you’re notable. Third-party research cites it as a large citation source for some AI systems, though I’m not restating the exact percentage here — check the source research for the current figure and its methodology. Wikipedia content is also used for real-time groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. by some systems, not just training. You can’t write your own page — earn the press coverage that establishes notability first.
  3. LinkedIn / Crunchbase / official registrations, then social profiles.

One cautionary tale from our own team: Ryan Law accidentally led Google to think he owned the Ahrefs website by adding schema to his personal site with an error in the sameAs property. Always double-check schema before you push it live.

Structured data: indirect lever, not a citation button

Be clear-eyed about what schema does. An Ahrefs observational study by Louise Linehan and Xibeijia Guan tracked a large set of pages that added JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. over roughly seven months and compared citation volume before and after — see the schema/AI citations study for the dataset and exact percentages. The directional finding: adding schema to an already-cited page did not produce a clear positive lift in AI citation volume across the platforms tracked.

So treat schema as not a reliable citation lever. Its value is entity disambiguation — connecting your brand to its Knowledge Graph identity — and correct indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., not a direct bump in how often AI quotes you. Mueller’s own position lines up: structured data is not a direct ranking factor, though it helps machines understand entities and is genuinely useful for shopping data that’s “basically impossible to read in high fidelity… from a text page.”

For AI specifically, @id on every entity and an @graph connecting Organization, Person, and Article into one graph is a useful modeling pattern practitioners recommend for keeping a connected identity consistent — it isn’t a documented Google or AI-provider requirement, and using it doesn’t guarantee a ranking, panel, or citation effect. Isolated schema blocks for the same entities work too; a connected graph is about making maintenance and consistency easier, not unlocking a feature.

Server-render your schema — don’t assume any crawler runs your JS

This is a technical mistake that can silently sink otherwise-good entity work, and it’s worth being precise about what’s actually documented rather than repeating a blanket claim. Google explicitly documents renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. JavaScript-generated structured data for Search. Whether a given AI crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — GPTBot, ClaudeBot, PerplexityBot, or any other named provider — executes JavaScript is provider- and version-specific, and it changes over time; I don’t have current, per-provider primary documentation to assert a universal “AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. don’t run JavaScript” rule, and I’d rather send you to verify the current behavior yourself (a dated fetch test against the provider you care about) than repeat an unverified generalization.

The safe, provider-agnostic move either way: treat server-rendered schema in the initial HTML as the robust baseline. If a crawler you’re targeting turns out not to render JavaScript, only server-delivered schema is guaranteed visible to it; if it does render JS, server-rendering costs you nothing. Check your sameAs and @graph in the raw HTML response (not just the rendered DOM) so you know what’s actually being served before any rendering step.

So what do you actually do?

This is practitioner guidance — a way I find useful to organize the work, not a documented AI-system scoring model. I think about it as three layers: identity, relationships, and independently verified corroboration.

Identity. Nail your entity home — one canonical page, Organization/Person schema server-rendered as the robust baseline, and accurate sameAs. Keep NAP (name/address/phone) and descriptions consistent everywhere, including your Google Business Profile if you’re local — a well-documented Google data source, even though I can’t confirm it’s the primary local source for every AI system.

Relationships. Connect your entities — sameAs to Wikidata (and Wikipedia if you’re notable), an @graph linking Organization, Person, and Article where it helps you keep things consistent. In your content, cover the entities and relationships your readers actually need explained — don’t chase a Cloud Natural Language salience score as if it were a ranking input; salience describes how a document-analysis API scores a piece of text, not a documented Google Search or AI-citation ranking factor.

Corroboration. The layer that moves the needle most in my experience — earn genuine third-party mentions and a real YouTube presence, because those (not links alone) are strong signals of whether AI has independent, corroborated confidence you’re who you say you are. As I’ve said about the AI era, “SEO itself hasn’t drastically changed, but getting good results may now require closer collaboration with other teams like PR and partnerships.”

TIP

Turn a missing entity lookup into a concrete identity and corroboration checklist with my free Google Knowledge Graph Explorer Free

  1. Look up the canonical person, organization, or brand name you use publicly.
  2. Treat a no-result state as one lookup observation—not proof that every Google system lacks entity understanding.
  3. Work through the entity-home, accurate sameAs, and independent-corroboration playbook before checking again later.
A no-result lookup is a useful starting state, not a promise that schema or any single source will create a Knowledge Panel.

The result is labeled Not found and says Acme Analytics Test Brand is not in the returned Knowledge Graph results yet. Its playbook recommends using one canonical name, publishing an unambiguous entity home page, adding accurate Person or Organization markup with sameAs links, building corroboration in appropriate independent sources, and rechecking later. This lookup does not prove that every Google system lacks entity understanding and does not predict Knowledge Panel eligibility.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.