Schema Markup for AI

Does schema markup help you show up in AI search? The evidence says it's not a citation lever — it's entity infrastructure. Here's what actually matters.

First published: Jun 24, 2026 · Last updated: Jul 22, 2026 · Advanced
demand #11 in Optimization#21 in AI Search#249 on the site
1 evidence signal on this page

Schema markup doesn't directly increase AI citations — a 1,885-page study found no significant lift — but it's infrastructure for entity disambiguation that helps AI systems understand who you are.

TL;DR — Google says there’s “no special schema.orgSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. that you need to add” for AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. or AI Mode. Evidence for this claim Google says AI Overviews and AI Mode require no special schema.org markup beyond normal Search eligibility and best practices. Scope: Google AI search features; structured data can still support ordinary Search features. Confidence: high · Verified: Google Search Central: AI features and your website Bing’s Fabrice Canel is the one rep to confirm schema helps their LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4).. The Ahrefs 1,885-page study foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. adding JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. moved AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. a statistically insignificant amount on every platform (−4.6% AIO, +2.4% AI Mode, +2.2% ChatGPT). The most plausible mechanism is entity disambiguation via sameAs feeding Google’s Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. — not a proven citation lever, since the Knowledge-Graph-to-AI-Overview chain hasn’t been isolated in any study. Watch out for two traps: feature eligibility can change while a Schema.org type remains valid, and client-side markup may not be processed consistently by every crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..

Regardless of delivery method, the markup must match content visible on the page. Evidence for this claim Google requires structured data to represent visible page content and does not guarantee feature display. Scope: Google structured-data policies, not claims about every AI system. Confidence: high · Verified: Google Search Central: Structured data guidelines

Start from the official position, because it’s clearer than the hype

Google has been unusually blunt here. Its AI features guidance says: “You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.” The only structured-data recommendation Google attaches to AI features is “making sure your structured data matches the visible text on the page” — which has always been policy, not an AI-specific tactic.

That doesn’t mean schema is useless to Google. It means Google separates two jobs structured data does: it drives rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. (the appearance enhancement) and it helps Google understand content and entities — “information about the people, books, or companies that are included in the markup.” The second job feeds the Knowledge Graph, and the Knowledge Graph feeds AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.. So there’s a plausible indirect path. There is no documented direct one.

Bing is the one platform that says schema helps its LLMs

This is the cleanest affirmation in the whole space. At SMX Munich in March 2025, Fabrice Canel, Principal Product Manager at Microsoft Bing, confirmed that schema markupSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor. helps Microsoft’s LLMs understand content. He’s the only AI-search rep to say it on the record about an LLM specifically — not just about crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor.. He paired it with a push on freshness: “Gen AIs value fresh content in particular, partly as a reference check of their LLM training data. Use the API at indexnowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it..org to push that information as it’s published or updated.”

So when people ask “do LLMs even read schema?”, the honest answer is platform-split: Bing/Copilot, yes, by their own statement; Google, only indirectly via the Knowledge Graph; ChatGPT and Perplexity, unconfirmed either way.

The evidence on citations: the Ahrefs 1,885-page study

This is the most rigorous test I’m aware of. The Ahrefs team (Louise Linehan and Xibeijia Guan, reviewed by Ryan Law, published May 2026) identified 1,885 pages that added JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. schema between August 2025 and March 2026, matched them against ~4,000 control pages with similar prior citation levels, and measured citations 30 days before and after — across four statistical tests (t-test, difference-in-differences, event study, sensitivity analysis).

PlatformEffect after adding schemaVerdict
Google AI Overviews−4.6%Statistically insignificant
Google AI Mode+2.4%Indistinguishable from zero
ChatGPT+2.2%Indistinguishable from zero

The headline from the study: “Adding schema produced no major uplift in citations on any platform.” Earlier, a Search/Atlas analysis in December 2024 found the same shape of result — no consistent correlation between schema coverage and citation rates.

The one honest caveat. The study only measured pages already being cited — every treated page had 100+ AI Overview citations before schema was added. So it answers “does adding schema lift citations on pages already in the consideration set?” (no) but it can’t answer “does schema help a page get into that set in the first place?” — that’s still open. The correlation people cite — AI-cited pages are ~3× more likely to have JSON-LD — almost certainly reflects that technically strong sites use schema and publish good content, not that schema is the cause.

The most plausible mechanism: entity disambiguation

If schema doesn’t move citations, why do I still tell people to invest in it for AI search? Because of the one job it’s actually documented to do: help resolve entities — connecting Organization and Person markup to authoritative external IDs so a knowledge graph can tell your company or author apart from a namesake.

sameAs is the property that does this. It’s the most direct way to point at authoritative external identifiers:

  • Organization sameAs (priority order): Wikidata Q-number → Wikipedia → LinkedIn company page → Crunchbase → GitHub (for tech companies).
  • Person sameAs: LinkedIn → Wikidata (if an entry exists) → ORCID (academic authors) → X → GitHub.

This is the part of schema I’m most confident changes something — it feeds Google’s Knowledge Graph, which Google documents as an input to AI Overviews. But I want to be precise about what’s confirmed and what isn’t: schema’s role in entity resolution for the Knowledge Graph is documented; the Knowledge-Graph-to-AI- Overview-citation chain is a plausible hypothesis, not something any study has isolated and confirmed. Treat it as the best-supported bet, not a guarantee.

One frequently cited example: SchemaApp, a schema vendor, reported that Wells Fargo corrected AI hallucinationsAn AI hallucination is when a large language model generates output that is confidently stated but factually wrong, made up, or unsupported by its source. It's a side effect of next-token prediction — not a bug that can be fully eliminated. about its business hours by adding schema to location pages. That’s a vendor-published case study I haven’t been able to verify independently — treat it as an anecdote consistent with the entity-disambiguation mechanism, not confirmed proof of it.

Which types are worth implementing

There’s no confirmed, universal ranking of schema types by AI-visibility value — no study isolates that. What follows is what each type documentedly does; treat “helps AI” as the plausible entity/rich-result path described above, not a scored priority order:

TypeWhat it does
Organization + sameAsAnchors your brand as an entity, with external IDs a knowledge graph can resolve. The strongest-evidenced piece of this list, because entity resolution is the one documented mechanism.
Person + sameAsAuthor identity and expertise attribution; avoids disambiguation errors. Same mechanism as above, applied to authors.
Article / NewsArticleDeclares content type, author, datePublished, dateModified — Google documents freshness/authorship as part of ordinary content understanding.
BreadcrumbListCommunicates site hierarchy for ordinary Search understanding.
FAQPageRich result deprecated (see below); the type is still valid. Whether it helps AI systems parse Q&A passages is unconfirmed either way.

A note on FAQPage: Google deprecated the FAQ rich result on May 7, 2026. The visual enhancement in the SERP is gone, but the FAQPage schema type itself is still valid — “the markup can stay on your pages without causing problems” — and may still help AI systems identify question-and-answer passages. Don’t rush to strip it out; removing it buys you nothing.

The trap that quietly kills schema for AI: JavaScript

Here’s the implementation detail that catches people. AI-crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. is provider- and version-specific. If your schema is injected client-side, the classic example being Google Tag Manager, then GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. (which renders) may eventually see it, but any AI crawler that uses only the raw HTML will miss it. For broad AI-search coverage, server-render schema into the HTML response. If entity disambiguation is the reason you’re doing this, JS-injected schema defeats the purpose.

Use JSON-LD (Google, Bing, and schema.org all prefer it over Microdata/RDFa), keep it server-side, and connect your types with @graph + @id so your Organization, Person, and Article reference each other instead of duplicating data — that connected-graph pattern is the one most associated with Knowledge Graph recognition.

The bottom line

Schema is infrastructure, not a citation cheat code. Implement Organization and Person with real sameAs links, put Article and BreadcrumbList on your templates, render it all server-side, and make it match your visible content. Then stop expecting it to be the thing that gets you into AI answers — that work happens in your content, your brand mentions, and your authority, which I cover across the rest of the AI search optimizationAI search optimization is the practice of making your brand and content visible, citable, and accurately represented across AI-powered search — Google AI Overviews, ChatGPT, Perplexity, Copilot. It's built on traditional SEO plus a heavier emphasis on off-site brand mentions and content AI systems can cite. cluster.

TIP

Build and validate the machine-readable layer for page agreement and entity clarity—not as a citation promise—with my free Schema Generator Free

  1. Choose the page type and enter only facts that are visible and accurate on the page.
  2. Resolve invalid identifiers and review recommended author, image, and date fields for completeness.
  3. Test the deployed server-delivered JSON-LD; valid schema supports understanding but does not guarantee an AI citation.
Validation separates graph quality from citation claims: fix invalid identifiers and incomplete useful fields because they weaken the data, not because passing guarantees AI visibility.

The Article Schema form reports one issue to fix and zero of four recommended fields completed. The canonical URL is not-a-valid-url and is marked invalid because it lacks a valid absolute URL. The headline exceeds 110 characters, while image, publication date, and author are missing. The generated JSON-LD shows the invalid identifier in mainEntityOfPage.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.