Schema Markup for AI
Does schema markup help you show up in AI search? The evidence says it's not a citation lever — it's entity infrastructure. Here's what actually matters.
1 evidence signal on this page
- Related live toolSchema Markup Validator
Schema markup doesn't directly increase AI citations — a 1,885-page study found no significant lift — but it's infrastructure for entity disambiguation that helps AI systems understand who you are.
TL;DR — Schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. is code you add to a page that labels what your content means so machines don’t have to guess. You’ll hear that adding it boosts your odds of getting cited in AIThree distinct states of AI visibility: retrieved (an AI fetched your page as source material), mentioned (your brand appears in the answer text), and cited (your URL is linked as a source). They don't always happen together, and each is measured with a different tool. Overviews or ChatGPT. The best evidence says it doesn’t do that directly. What it does do is help AI systems understand who you are and what your pages are — useful, just not the magic switch it’s sold as.
What schema markup is
When you read a page, you can tell a phone number from a price from an author’s name just by looking. A machine can’t — it sees text. Schema markup (also called structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding.) is extra code you add to a page that spells it out: “this is the author,” “this is the published date,” “this is the price.” It uses a shared vocabulary from a site called schema.org so every search engine and AI system reads it the same way.
Most of the time it’s written in a format called JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. — a little block of code that sits in the page’s HTML and describes the content without changing how the page looks to a visitor. It still has to describe content readers can actually see on the page. Evidence for this claim Google requires structured data to represent visible page content and does not guarantee feature display. Scope: Google structured-data policies, not claims about every AI system. Confidence: high · Verified: Google Search Central: Structured data guidelines
The AI hype vs. what’s real
Here’s the pitch you’ve probably heard: add schema markup and you’ll get cited more in AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. It sounds plausible — structured data is “AI-readable,” so surely the AI rewards it, right?
The data says no. When my colleagues at Ahrefs tracked 1,885 pages that added schema and measured their AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. before and after, the citations basically didn’t move on any platform. And Google flat-out says there’s no special schema you need to add to show up in AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.. Evidence for this claim Google says AI Overviews and AI Mode require no special schema.org markup beyond normal Search eligibility and best practices. Scope: Google AI search features; structured data can still support ordinary Search features. Confidence: high · Verified: Google Search Central: AI features and your website So if a tool or agency is promising that bolting on schema will get you into AI answers, be skeptical.
What schema actually does
It’s still worth doing — just for the right reasons:
- It tells AI systems who you are. The single most useful piece is
OrganizationandPersonmarkup with a property calledsameAs, which links your brand or author identity to your Wikipedia, Wikidata, or LinkedIn pages. That helps AI systems tell your “Apple” from the fruit, and your author from someone else with the same name. - It powers rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show.. Star ratings, product cards, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. — those visual extras in regular search still come from schema.
- It helps search engines understand your content so it can be categorized and surfaced correctly.
So: keep using schema. Just don’t expect it to be the lever that gets you into the AI answer. Want the deep version — the studies, the Google vs. Bing split, and the exact types that matter? Switch to the Advanced tab.
TL;DR — Google says there’s “no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add” for AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. or AI Mode. Evidence for this claim Google says AI Overviews and AI Mode require no special schema.org markup beyond normal Search eligibility and best practices. Scope: Google AI search features; structured data can still support ordinary Search features. Confidence: high · Verified: Google Search Central: AI features and your website Bing’s Fabrice Canel is the one rep to confirm schema helps their LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4).. The Ahrefs 1,885-page study foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. adding JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. moved AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. a statistically insignificant amount on every platform (−4.6% AIO, +2.4% AI Mode, +2.2% ChatGPT). The most plausible mechanism is entity disambiguation via
sameAsfeeding Google’s Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. — not a proven citation lever, since the Knowledge-Graph-to-AI-Overview chain hasn’t been isolated in any study. Watch out for two traps: feature eligibility can change while a Schema.org type remains valid, and client-side markup may not be processed consistently by every crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
Regardless of delivery method, the markup must match content visible on the page. Evidence for this claim Google requires structured data to represent visible page content and does not guarantee feature display. Scope: Google structured-data policies, not claims about every AI system. Confidence: high · Verified: Google Search Central: Structured data guidelines
Start from the official position, because it’s clearer than the hype
Google has been unusually blunt here. Its AI features guidance says: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.” The only structured-data recommendation Google attaches to AI features is “making sure your structured data matches the visible text on the page” — which has always been policy, not an AI-specific tactic.
That doesn’t mean schema is useless to Google. It means Google separates two jobs structured data does: it drives rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. (the appearance enhancement) and it helps Google understand content and entities — “information about the people, books, or companies that are included in the markup.” The second job feeds the Knowledge Graph, and the Knowledge Graph feeds AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.. So there’s a plausible indirect path. There is no documented direct one.
Bing is the one platform that says schema helps its LLMs
This is the cleanest affirmation in the whole space. At SMX Munich in March 2025, Fabrice Canel, Principal Product Manager at Microsoft Bing, confirmed that schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. helps Microsoft’s LLMs understand content. He’s the only AI-search rep to say it on the record about an LLM specifically — not just about crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor.. He paired it with a push on freshness: “Gen AIs value fresh content in particular, partly as a reference check of their LLM training data. Use the API at indexnowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it..org to push that information as it’s published or updated.”
So when people ask “do LLMs even read schema?”, the honest answer is platform-split: Bing/Copilot, yes, by their own statement; Google, only indirectly via the Knowledge Graph; ChatGPT and Perplexity, unconfirmed either way.
The evidence on citations: the Ahrefs 1,885-page study
This is the most rigorous test I’m aware of. The Ahrefs team (Louise Linehan and Xibeijia Guan, reviewed by Ryan Law, published May 2026) identified 1,885 pages that added JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. schema between August 2025 and March 2026, matched them against ~4,000 control pages with similar prior citation levels, and measured citations 30 days before and after — across four statistical tests (t-test, difference-in-differences, event study, sensitivity analysis).
| Platform | Effect after adding schema | Verdict |
|---|---|---|
| Google AI Overviews | −4.6% | Statistically insignificant |
| Google AI Mode | +2.4% | Indistinguishable from zero |
| ChatGPT | +2.2% | Indistinguishable from zero |
The headline from the study: “Adding schema produced no major uplift in citations on any platform.” Earlier, a Search/Atlas analysis in December 2024 found the same shape of result — no consistent correlation between schema coverage and citation rates.
The one honest caveat. The study only measured pages already being cited — every treated page had 100+ AI Overview citations before schema was added. So it answers “does adding schema lift citations on pages already in the consideration set?” (no) but it can’t answer “does schema help a page get into that set in the first place?” — that’s still open. The correlation people cite — AI-cited pages are ~3× more likely to have JSON-LD — almost certainly reflects that technically strong sites use schema and publish good content, not that schema is the cause.
The most plausible mechanism: entity disambiguation
If schema doesn’t move citations, why do I still tell people to invest in it for AI
search? Because of the one job it’s actually documented to do: help resolve
entities — connecting Organization and Person markup to authoritative
external IDs so a knowledge graph can tell your company or author apart from a
namesake.
sameAs is the property that does this. It’s the most direct way to point at
authoritative external identifiers:
- Organization
sameAs(priority order): Wikidata Q-number → Wikipedia → LinkedIn company page → Crunchbase → GitHub (for tech companies). - Person
sameAs: LinkedIn → Wikidata (if an entry exists) → ORCID (academic authors) → X → GitHub.
This is the part of schema I’m most confident changes something — it feeds Google’s Knowledge Graph, which Google documents as an input to AI Overviews. But I want to be precise about what’s confirmed and what isn’t: schema’s role in entity resolution for the Knowledge Graph is documented; the Knowledge-Graph-to-AI- Overview-citation chain is a plausible hypothesis, not something any study has isolated and confirmed. Treat it as the best-supported bet, not a guarantee.
One frequently cited example: SchemaApp, a schema vendor, reported that Wells Fargo corrected AI hallucinationsAn AI hallucination is when a large language model generates output that is confidently stated but factually wrong, made up, or unsupported by its source. It's a side effect of next-token prediction — not a bug that can be fully eliminated. about its business hours by adding schema to location pages. That’s a vendor-published case study I haven’t been able to verify independently — treat it as an anecdote consistent with the entity-disambiguation mechanism, not confirmed proof of it.
Which types are worth implementing
There’s no confirmed, universal ranking of schema types by AI-visibility value — no study isolates that. What follows is what each type documentedly does; treat “helps AI” as the plausible entity/rich-result path described above, not a scored priority order:
| Type | What it does |
|---|---|
Organization + sameAs | Anchors your brand as an entity, with external IDs a knowledge graph can resolve. The strongest-evidenced piece of this list, because entity resolution is the one documented mechanism. |
Person + sameAs | Author identity and expertise attribution; avoids disambiguation errors. Same mechanism as above, applied to authors. |
Article / NewsArticle | Declares content type, author, datePublished, dateModified — Google documents freshness/authorship as part of ordinary content understanding. |
BreadcrumbList | Communicates site hierarchy for ordinary Search understanding. |
FAQPage | Rich result deprecated (see below); the type is still valid. Whether it helps AI systems parse Q&A passages is unconfirmed either way. |
A note on FAQPage: Google deprecated the FAQ rich result on May 7, 2026. The
visual enhancement in the SERP is gone, but the FAQPage schema type itself is
still valid — “the markup can stay on your pages without causing problems” — and
may still help AI systems identify question-and-answer passages. Don’t rush to strip
it out; removing it buys you nothing.
The trap that quietly kills schema for AI: JavaScript
Here’s the implementation detail that catches people. AI-crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. is provider- and version-specific. If your schema is injected client-side, the classic example being Google Tag Manager, then GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. (which renders) may eventually see it, but any AI crawler that uses only the raw HTML will miss it. For broad AI-search coverage, server-render schema into the HTML response. If entity disambiguation is the reason you’re doing this, JS-injected schema defeats the purpose.
Use JSON-LD (Google, Bing, and schema.org all prefer it over Microdata/RDFa),
keep it server-side, and connect your types with @graph + @id so your
Organization, Person, and Article reference each other instead of duplicating
data — that connected-graph pattern is the one most associated with Knowledge Graph
recognition.
The bottom line
Schema is infrastructure, not a citation cheat code. Implement Organization and
Person with real sameAs links, put Article and BreadcrumbList on your
templates, render it all server-side, and make it match your visible content. Then
stop expecting it to be the thing that gets you into AI answers — that work happens
in your content, your brand mentions, and your authority, which I cover across the
rest of the AI search optimizationAI search optimization is the practice of making your brand and content visible, citable, and accurately represented across AI-powered search — Google AI Overviews, ChatGPT, Perplexity, Copilot. It's built on traditional SEO plus a heavier emphasis on off-site brand mentions and content AI systems can cite. cluster.
Build and validate the machine-readable layer for page agreement and entity clarity—not as a citation promise—with my free Schema Generator Free
- Choose the page type and enter only facts that are visible and accurate on the page.
- Resolve invalid identifiers and review recommended author, image, and date fields for completeness.
- Test the deployed server-delivered JSON-LD; valid schema supports understanding but does not guarantee an AI citation.
The Article Schema form reports one issue to fix and zero of four recommended fields completed. The canonical URL is not-a-valid-url and is marked invalid because it lacks a valid absolute URL. The headline exceeds 110 characters, while image, publication date, and author are missing. The generated JSON-LD shows the invalid identifier in mainEntityOfPage.
AI summary
A condensed take on the Advanced version:
- Schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. ≠ AI citationAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. lever. The Ahrefs May 2026 study (1,885 pages, four statistical tests) foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. adding JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. moved citations a statistically insignificant amount everywhere: −4.6% Google AIOAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., +2.4% AI Mode, +2.2% ChatGPT.
- Google’s official line: “There’s also no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add” for AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. or AI Mode. The only ask is that schema match visible content.
- Bing is the exception. Fabrice Canel (Microsoft) confirmed schema helps Bing’s LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). — the only AI-platform rep to say so on the record.
- The most plausible mechanism is entity disambiguation.
OrganizationandPersonwithsameAs(Wikidata, Wikipedia, LinkedIn) feed Google’s Knowledge Graph — documented. Whether that graph inclusion drives AI Overview citations is a reasonable hypothesis, not something any study has isolated. - Correlation ≠ causation. AI-cited pages are ~3× more likely to have schema — but that reflects technically strong sites using both schema and good content.
- Two traps: FAQ rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. were deprecated May 7, 2026 (the
FAQPagetype still works); and AI-crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. varies by provider, so client-side / Tag Manager schema can be missed — server-render it to maximize coverage. - What to actually do:
Organization+Person+sameAs,Article,BreadcrumbList, JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal., server-side, matching visible text.
Official documentation
Primary-source documentation from the search engines and schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor..
- AI Features and Your Website — the source of the “no special schema.org structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add” statement.
- Intro to How Structured Data Markup Works — what structured data does for understanding and rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show..
- General Structured Data Guidelines — the “markup must match visible content” policy.
- FAQPage Structured Data — the type’s documentation, post-deprecation.
- Speakable (BETA) — marking audio-suitable sections; still beta, US English only.
Bing / Microsoft
- Marking Up Your Site with Structured Data — Bing Webmaster Help — Bing’s structured data guidance; JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. recommended, must match page context.
schema.org
- schema.org — the shared vocabulary itself.
- Organization · Person · sameAs — the entity-disambiguation core.
Quotes from the source
On-the-record statements from Google and Microsoft on schema and AI.
Google — no special schema for AI features
- “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add.” — Google Search Central, AI features documentation. Source
Google — what structured data does
- “Google uses structured data that it finds on the web to understand the content of the page, as well as to gather information about the web and the world in general, such as information about the people, books, or companies that are included in the markup.” — Google Search Central. Source
Google — FAQ deprecation
- “FAQPage as a Schema.org type is still valid, and the markup can stay on your pages without causing problems.” — as reported by Search Engine Journal on the May 2026 FAQ rich-result deprecation. Coverage
Fabrice Canel, Principal Product Manager, Microsoft Bing (SMX Munich, March 2025)
- Confirmed on stage that schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. helps Microsoft’s LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). understand content — the only AI-platform rep to state this about an LLM specifically. Coverage
- “Gen AIs value fresh contentContent freshness is how recent or up-to-date a page is — by its original publish date, its last substantive revision, or the currency of the facts inside it. It only helps rankings when the query itself benefits from recent results (Query Deserves Freshness), and cosmetic date changes with no real update don't count. in particular, partly as a reference check of their LLM training data. Use the API at indexnowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it..org to push that information as it’s published or updated.” Coverage
Schema-for-AI implementation checklist
A pass to make sure your structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. actually does its entity-disambiguation job — and doesn’t get silently dropped by AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls.:
-
Organizationschema sitewide, with a populatedsameAsarray (Wikidata, Wikipedia, LinkedIn at minimum), plusname,url,logo. -
Personschema on author/bio pages, withsameAs(LinkedIn, Wikidata if one exists) andknowsAboutfor topical authoritySemantic search is meaning-based retrieval — matching what a user means, not just the words they typed. Search engines detect entities, expand synonyms, infer intent, and rank by conceptual relevance, which is why keyword stuffing lost its power and topical depth gained it.. -
Articleon content pages —author(linked by@idto the Person),datePublished,dateModified,headline,publisher(linked to the Organization). -
BreadcrumbListon interior pages to signal hierarchy / cluster. - Schema is server-rendered, not injected client-side via Tag Manager —
confirm it’s in the raw HTML (View Source /
curl) so coverage does not depend on any AI crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.’s provider-specific JavaScript support. - Markup matches visible content — no fields describing things not on the page (Google policy; mismatches risk a manual action).
- Format is JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal., ideally connected with
@graph+@id. -
FAQPageleft in place if you have it — the rich result is gone (May 2026) but the type is still valid; removing it gains nothing. - Validated in the Rich ResultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. Test and schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. validator.
- Expectations set: schema is entity infrastructure, not a direct AI citation lever — don’t bolt it onto already-cited pages expecting a lift.
Schema types — cheat sheet
Type → what it declares → what’s actually documented
No study confirms a universal priority order for AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals. across these types — read the right column as “documented function,” not a ranking.
| Type | What it declares | What’s documented |
|---|---|---|
Organization + sameAs | Your brand entity, linked to external IDs | Feeds Google’s Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. — the one mechanism with a documented (if indirect) path to AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. |
Person + sameAs | An author entity and external identities | Same entity-resolution mechanism, applied to authors |
Article / NewsArticle | Content type, author, publish/modified dates | Part of Google’s ordinary content-understanding and freshness signals |
BreadcrumbList | Site hierarchy | Ordinary Search context/hierarchy signal |
FAQPage | Q&A structure (rich result deprecated May 2026) | Type still valid, causes no harm; AI-parsing benefit unconfirmed |
Product + Offer | Price, availability, ratings | Documented for rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. and comparison shopping; AI-agent use is plausible, not confirmed by any source here |
HowTo | Step-by-step procedures | Documented for rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show.; deterministic AI extraction is a reasonable inference, not a confirmed study finding |
Fast facts
- Ahrefs study (1,885 pages, May 2026): −4.6% AIO / +2.4% AI Mode / +2.2% ChatGPT — all statistically insignificant.
- Google: “no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add” for AI Overviews / AI Mode.
- Bing: Fabrice Canel confirmed schema helps Bing’s LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). — the one platform that has.
- FAQ rich result: deprecated May 7, 2026;
FAQPagetype still valid. - AI-crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. varies by provider → server-render schema so Tag Manager execution is not a coverage dependency.
- Format: JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal., ideally
@graph+@id. Must match visible content.
Schema-for-AI mistakes to avoid
Promising a citation boost
Schema helps machines interpret entities and page meaning; it is not a switch that forces an AI answer to cite the page.
Marking up facts users cannot verify
Structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. should represent visible, accurate page content. Do not add awards, reviews, authors, prices, or relationships solely because the vocabulary permits them.
Linking the wrong entity with sameAs
A wrong identifier is worse than a missing one. Confirm that every URL represents the same person, organization, product, or place.
Injecting all markup after client rendering
Search engines may render it, but many AI crawlersAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. will not. Put important JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. in the server-delivered HTML when possible.
Common schema-for-AI problems
The markup validates but describes a different entity
Symptom: Syntax passes while names, URLs, or identifiers point to a namesake.
Likely cause: Entity resolution was skipped. Fix: Compare the visible page,
canonical URLHow search engines pick one canonical URL among duplicates and consolidate signals onto it., official profiles, and sameAs values before correcting the graph.
Structured data is missing from raw HTML
Symptom: Browser tools find JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. but curl does not. Likely cause: A tag
manager or client component injects it after load. Fix: render stable entity and
page markup on the server, then compare raw and rendered output again.
Rich-results tools report ineligible markup
Symptom: Schema is valid vocabulary but not eligible for a Google rich result. Likely cause: Schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. validity and Google feature requirements are different checks. Fix: Use the general validator for vocabulary and the Rich ResultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. Test for supported search features; do not equate feature eligibility with AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking..
Frameworks for schema that supports machine understanding
Entity first, syntax second
- Identify: Name the real person, organization, product, place, or creative work.
- Corroborate: Confirm stable facts and authoritative identifiers.
- Model: Choose the most specific appropriate type and relationships.
- Expose: Match visible page content and serve important markup in HTML.
- Validate: Check both Schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structure and any feature-specific requirements.
- Monitor: Recheck when templates or entity facts change.
Separate three outcomes
- Valid markup: The graph follows the vocabulary and syntax.
- Feature eligibility: A search engine may consider the page for a supported rich result when all additional rules are met.
- AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.: A model may mention or cite the page based on a much broader retrieval and trust system.
Do not report one outcome as proof of another.
Schema extraction snippets
List JSON-LD blocks in Chrome DevTools
Run in the Console:
[...document.querySelectorAll('script[type="application/ld+json"]')].map((el, index) => ({ index, json: JSON.parse(el.textContent) }))Check whether JSON-LD exists before rendering
Run in a terminal:
curl -sS https://example.com/page/ | grep -n 'application/ld+json'If the rendered DOM contains markup but this response does not, the implementation is client-dependent.
Extract declared types from a crawl
Use this regular expression against JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. source. Capture group 1 returns a string
or array value following @type:
"@type"\s*:\s*("[^"]+"|\[[^\]]+\])Parse JSON for production analysis; regex is only a quick extraction aid.
Tools for schema and AI understanding
- Schema Markup Validator: Check JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. structure, vocabulary, and graph issues without promising a search feature.
- Schema Markup Generator: Create a starting graph from facts you have already verified against the visible page.
- Rich-Result Eligibility Checker: Separate Google’s supported rich-result requirements from general schema validity.
- Entity Coverage Analyzer: Review whether the page clearly establishes the entities and relationships the markup describes.
- Raw-response and rendered-DOM comparison: Confirm important JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. is available without depending on client-side execution.
Validate a schema change
Test graph validity and page agreement
Test to run: Validate the JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal., then compare every important property with the visible page. Expected result: The graph parses and all stated facts are accurate and user-verifiable. Failure interpretation: Syntax is broken or markup overstates the content. Monitoring window: Immediate. Rollback trigger: A release asserts a wrong entity, unsupported fact, or invalid graph.
Test server delivery
Test to run: Fetch the page without JavaScript and locate the intended JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal.. Expected result: Stable entity and page markup appears in the response. Failure interpretation: Non-renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. cannot receive the change. Monitoring window: Immediate after deployment. Rollback trigger: Previously server-visible markup becomes client-only.
Test entity identifiers
Test to run: Open each sameAs, url, and external identifier and verify the
entity match. Expected result: Every edge resolves to the same intended entity.
Failure interpretation: The graph connects a namesake or obsolete profile.
Monitoring window: Immediate and after rebrands. Rollback trigger: Any identity
edge points to an unrelated entity.
Test yourself: Schema Markup for AI
Resources worth your time
Related Ahrefs writing
- We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved. — Louise Linehan & Xibeijia Guan, May 2026. The controlled study at the center of this article.
- Schema Markup: What It Is & How to Implement It — the implementation walkthrough (Despina Gavoyannis, Chris Haines).
- What is Structured Data? — the glossary primer.
Official references
- schema.org — the vocabulary.
- Google — AI Features and Your Website — Google’s stated position on schema and AI.
From others
- How schema markup fits into AI search — without the hype — Search Engine Land — also references the December 2024 Search/Atlas study finding no consistent citation correlation.
- What 2025 Revealed About AI Search and the Future of Schema Markup — SchemaApp — the Wells Fargo entity-correction case.
- Microsoft Bing/Copilot use schema for its LLMs — Search Engine Land — coverage of Fabrice Canel’s SMX Munich confirmation.
- Schema Helps Microsoft’s LLMs (Copilot) Understand Your Content — Search Engine Roundtable — Barry Schwartz’s write-up of the Bing/Copilot schema confirmation.
- Google Confirms: Structured Data Still Essential in AI Search Era — Search Engine Journal — coverage of Google Search Central Live Madrid 2026 guidance.
- Schema Markup Has No Meaningful Impact on AI Citations — Stan Ventures — independent summary of the Ahrefs findings.
- Structured Data in the AI Search Era — BrightEdge — enterprise perspective on schema in the AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. transition.
Stats worth citing
- Adding schema barely moved AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking.. The Ahrefs study of 1,885 pages that added JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. (vs. ~4,000 controls, four statistical tests) measured −4.6% on Google AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., +2.4% on AI Mode, and +2.2% on ChatGPT — all statistically insignificant. “Adding schema produced no major uplift in citations on any platform.” Source
- The “already cited” caveat. Every treated page in that study already had 100+ AI Overview citations before adding schema — so the study tests pages already in the consideration set, not whether schema helps a page enter it. Source
- Correlation, not causation. Across Ahrefs’ dataset, AI-cited pages are nearly 3× more likely to carry JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. than non-cited pages — a reflection of technically strong sites, not a causal schema effect. Source
- An earlier signal pointed the same way. A Search/Atlas analysis (December 2024) foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. no consistent correlation between schema coverage and AI citation rates. Coverage
Schema Markup for AI
Schema markup (structured data) is machine-readable code — usually JSON-LD — that labels what your content means using the schema.org vocabulary. For AI search it's infrastructure for entity disambiguation, not a direct citation lever: controlled studies found no meaningful uplift in AI citations from adding it.
Related: AI Search Optimization, AI Search, Entity SEO, Structured Data
Schema Markup for AI
Schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. (also called structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding.) is machine-readable code added to a page — typically as JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. inside a <script type="application/ld+json"> tag — that uses the vocabulary at schema.org to describe what a page is about in a form search engines and AI systems can parse without ambiguity. Where a human reads a paragraph and infers it’s a recipe, a machine reads a Recipe type with recipeIngredient and cookTime properties and knows.
For AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity., the honest framing is that schema is plumbing, not a magic switch. Google states plainly that there’s “no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured data that you need to add” to appear in AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. or AI Mode, and the most rigorous test to date — Ahrefs’ May 2026 study of 1,885 pages — foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. that adding schema produced no statistically significant change in citations on Google AIO, AI Mode, or ChatGPT. The correlation people point to (AI-cited pages are ~3× more likely to have schema) reflects that technically strong sites tend to use both schema and produce good content, not that schema causes citations.
Where schema earns its place is entity disambiguation. Organization and Person markup with the sameAs property — links to your Wikidata, Wikipedia, and LinkedIn identities — helps anchor your entity in Google’s Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself., which Google documents as an input to AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.. That’s the most plausible mechanism for schema mattering to AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. — it’s documented at the entity-resolution step, but the Knowledge-Graph-to-AI-citation link itself hasn’t been isolated by any study, so treat it as a hypothesis, not a guarantee. Microsoft’s Fabrice Canel is the one AI-platform rep to explicitly confirm (at SMX Munich, March 2025) that schema helps Bing’s LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). understand content. One practical catch: AI-crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. JavaScript support is provider- and version-specific rather than uniform, so schema injected client-side (e.g. via Tag Manager) risks being invisible to crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. that only fetch raw HTML — server-renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. it is the safer default.
Related: AI Search Optimization, AI Search, Entity SEO, Structured Data
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 22, 2026.
Editorial summary and recorded change details.Summary
Moved schema evidence notes to the exact claims they support.
Change details
-
Placed Google's no-special-schema citation and the visible-content requirement beside their respective claims in both explanatory lenses.
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Tightened claims that overstated schema's AI mechanism as proven: the entity-disambiguation-to-Knowledge-Graph-to-AI-Overview chain is now labeled a plausible hypothesis rather than a 'real, defensible mechanism' or something schema 'demonstrably' does. The Wells Fargo case study is now attributed as an unverified vendor case, and the schema-type tables no longer imply a universal AI-visibility priority order.
Change details
-
Reworded the advanced TL;DR, the entity-disambiguation section, and the ai-summary lens to describe schema's Knowledge Graph role as documented but its citation effect as an unconfirmed hypothesis.
-
Attributed the Wells Fargo entity-correction case (article and glossary) to SchemaApp as a vendor-published, independently unverified case study rather than an established fact.
-
Removed 'single best schema investment for AI search' and similar superlative rankings from the 'Which types are worth implementing' table and the schema-types cheat sheet; reframed both around documented function instead of an AI-visibility priority order.
-
Fixed a categorical claim in the glossary entry ('AI crawlers like GPTBot and ClaudeBot don't execute JavaScript') to match this run's provider-/version-specific hedge, consistent with the same correction on json-ld and nested-schema.
Full comparison unavailable — no prior snapshot was archived for this revision.