Guide : llms.txt

llms.txt is a proposed Markdown fichier at /llms.txt pour guiding AI systems. Google ignores it, 97% of fichiers obtenir zero requêtes, and Claude Code is the réel reader.

Première publication : 24 juin 2026 · Dernière mise à jour : 3 août 2026 · Advanced
Langues
1 indice probant sur cette page

llms.txt is a proposal by Jeremy Howard (Réponse.AI) pour helping AI agents navigate sites — 97% of publié fichiers obtenir zero requêtes, Google ignores it, and Claude Code is the principal consumer, pas search bots.

The specification describes a voluntary LLM-friendly site summary; it ne fait pas establish robot d’exploration compliance or indexation behavior. Evidence for this claim llms.txt is a community proposal for a Markdown file at /llms.txt that offers LLM-friendly site information; it is not a web standard. Scope: The proposal's own specification and stated purpose. Confidence: high · Verified: llms.txt proposal OpenAI currently documents named bots and independent robots.txt settings à la place. Evidence for this claim OpenAI documents robots.txt controls for its declared crawlers and does not list llms.txt as a crawler control. Scope: OpenAI's published crawler controls; absence is not proof that no internal system ever fetches the file. Confidence: high · Verified: OpenAI: Crawlers

TL;DR — llms.txt is a proposal (Jeremy Howard, Réponse.AI, Sept 3, 2024), pas an adopted standard. It’s a curated Markdown index at /llms.txt, with an optional full-content /llms-full.txt. The bots personnes hope are reading it pour AI search — OAI-SearchBot, PerplexityBot — barely register; the dominant consumer is Claude Code and autre coding agents. Google ignores it (Illyes confirmed; Mueller comparé it to the keywords meta tag). In Ahrefs’ 137K-site study, 97% of fichiers got zero requêtes, and SE Ranking trouvé que removing llms.txt from a citation model improved its accuracy. It’s worth ajout pour developer docs consumed by coding agents; pour GEO/AEO on a normal site, the données doesn’t prise en charge it. There’s aussi a réel prompt-injection risk.

Ce que llms.txt en réalité is

llms.txt is a proposal — I vouloir to lead with que word parce que it’s the whole story. There’s aucun RFC, aucun W3C blessing, aucun IETF traiter. It’s un well-reasoned suggestion from Jeremy Howard (co-founder of Réponse.AI and fast.ai), publié September 3, 2024, que caught on hard in the developer-documentation world and got grafted onto SEO by a community hungry pour an AI-visibility shortcut.

The design problem it targets is legitimate. As the spec puts it, “Grand language models increasingly rely on website information, but face a critical limitation: context windows are aussi petit to handle la plupart websites in leur entirety.” And converting HTML to clean, LLM-friendly text is, in the spec’s words, “difficult and imprecise.” Howard’s fix is to let le site author pre-flatten le contenu ils vouloir models to voir into curated Markdown.

The format

/llms.txt is an ordered Markdown document:

  1. H1 (requis) — project or site nom.
  2. Blockquote — a short summary with the clé information.
  3. Optional prose/listes — paragraphs or bullets, but aucun headings ici.
  4. H2-delimited lien listes- [name](url): optional notes.
  5. An ## Optional section — secondary liens a model peut skip quand context is tight.

A minimal exemple, from the spec:

# FastHTML

> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX,
> and fastcore's FT...

## Docs

- [FastHTML quick start](url): A brief overview of many features

## Optional

- [Starlette documentation](url): A subset useful for FastHTML development.

Two companion conventions exist:

  • /llms-full.txt — the entier site content flattened into un Markdown fichier. Anthropic, Perplexity, and Stripe publish les deux ce and the short formulaire.
  • The .md URL convention — offering page.md versions of individual pages so they’re LLM-ready sans the complet dump.

A newer variant skips the static fichiers entirely: edge-served Markdown via content negotiation. Cloudflare’s Markdown pour Agents converts quelconque HTML page to clean Markdown on the fly quand a client sends Accept: text/markdown — aucun per-page .md fichiers, aucun /llms-full.txt to regenerate, aucun template changements, parce que the CDN fait the conversion at the edge (Cloudflare cites ~80% fewer tokens pour the agent). It’s a paid fonctionnalité — Pro plan and up, pas the free tier — so on ce site it’s un Cloudflare upgrade away from being switched on plutôt que something I’ve hand-built. Si it’s worth switching on is the même ouvrir question que hangs over llms.txt itself: serving robots d’exploration a parallel Markdown copy of every page is convenient, but it’s un autre representation to garder in sync, and Bing’s Fabrice Canel has flagged the doubled explorer charger and cloaking-adjacent risk of edge-served alternate content. Vérifier votre logs pour réel agent demand avant flipping it.

llms.txt vs robots.txt — ils ne sont pas the même chose

The unique la plupart courant misconception is que llms.txt is “robots.txt for AI.” It isn’t, and conflating les leads to bad decisions.

robots.txtllms.txt
ObjectifAccès contrôler — ce que robots d’exploration peut récupérerContent guidance — what’s utile to lire
FormatDirectives (User-agent:, Disallow:)Markdown liens + descriptions
Enforced?Yes, by compliant robots d’explorationAucun — advisory seulement
TimingExplorer temps, avant fetchingInference temps, assembling context
A standard?Yes, decades old, universalAucun — a proposal
Peut it block?YesAucun

robots.txt is a réel, enforced standard. llms.txt is a suggestion box que la plupart systems aren’t reading yet. Si vous vouloir to contrôler AI robot d’exploration accès — qui is a separate, legitimate goal — que lives in robots.txt and AI robots d’exploration, pas ici.

Who en réalité reads it

Ce is the section que matters, parce que it’s où the données diverges hardest from the hype. In Ahrefs’ study of 137 210 domains (May 2026 trafic):

  • 97% of publié llms.txt fichiers reçu zero requêtes. Seulement ~3% (à propos de 1 100 domains) saw quelconque trafic at tout.
  • Of the fichiers que did obtenir requested: 96% of requêtes came from bots, 4% from humans.
  • Of the bot requêtes, 77% came from non-AI outils — SEO auditors (Ahrefs itself among les), anonymous robots d’exploration, tech-profiling bots. The SEO irony writes itself.
  • Seulement ~19,5% came from AI outils, and une fois vous break que bas it obtient worse pour the GEO theory:
    • AI agents & infrastructure (coding agents): ~10,5%
    • Training robots d’exploration (GPTBot led at 4,51%): ~5,3%
    • AI assistants: ~2,5%
    • AI retrieval bots — the ones que construire citation indexes: 1,1%.

Lire que dernier line twice. The bots la plupart personnes ajouter llms.txt pour — OAI-SearchBot, PerplexityBot — “barely registered.” The dominant AI consumer is Claude Code, Anthropic’s coding agent, qui the study trouvé “outfetched every AI retrieval bot, assistant, and training robot d’exploration except GPTBot.” And GPTBot is a training robot d’exploration que crawls everything regardless — its presence isn’t evidence of intentional llms.txt parsing.

Autre independent measurements land in the même placer. OtterlyAI watched un site pour 90 days: of 62 100+ AI bot visits, exactly 84 hit /llms.txt — à propos de 0,1%, performing “3x worse than average pages.” A separate analysis of 515M+ LLM bot events over 90 days trouvé 408 requêtes total targeting /llms.txt — “statistically negligible.”

The takeaway is the developer-docs-vs-GEO split: llms.txt fonctionne pour ce que Howard construit it pour (coding agents navigating API docs), and it largely doesn’t obtenir touched by the search-citation bots the SEO community veut it to reach.

Ce que the providers have said

Google — aucun, and aucun plans. Gary Illyes confirmed at Search Central Live (July 2025) que Google doesn’t prise en charge llms.txt and isn’t planning to. John Mueller, the most-quoted voice ici, comparé it to the keywords meta tag (April 2025) and appelé it a “temporary crutch, perhaps to save some tokens” pour AI coding outils — explicitly pas a search-visibility mechanism. Google’s May 2026 guidance stated machine-readable fichiers comme llms.txt aren’t necessary pour AI Overviews or AI Mode, qui “continue to rely on traditional SEO signals.” (Google did briefly publish an llms.txt on its dev docs on December 3, 2025, alors pulled it the même day — plus tard attributed to a CMS mettre à jour, pas a strategy shift.)

OpenAI — nothing. Aucun announcement que ChatGPT, GPTBot, or OAI-SearchBot parse it. Server logs contradict meaningful usage; OAI-SearchBot “barely registered.” OpenAI points personnes to robots.txt pour robot d’exploration contrôler.

Anthropic — the interesting un. Anthropic publishes les deux llms.txt and llms-full.txt at docs.anthropic.com/llms.txt, and Claude Code demonstrably récupère autre sites’ llms.txt fichiers — it’s the largest AI consumer in the données. But that’s a coding-tool story. There’s aucun confirmation que Claude.ai’s search or citation couche honors llms.txt. “Anthropic’s coding agent reads it” and “Anthropic’s search index trusts it” are différent claims, and seulement the premier is pris en charge.

Perplexity — claims yes, données dit rarely. Perplexity has stated it retrieves llms.txt and uses it to “prioritize page selection.” But proactive fetching is the catch: PerplexityBot “barely registered” in la requête données, and there’s “almost zero activity from PerplexityBot requesting llms.txt fichiers proactively.” Paste an llms.txt URL into Perplexity and it reads it fine; autonomous, proactive utiliser is the gap entre policy and pratique.

Bing/Microsoft, Apple, Meta — aucun public position. Treat as unconfirmed.

Fait it déplacer AI citations?

Aucun demonstrated benefit. SE Ranking modeled 300 000 domains with an XGBoost regression + SHAP analysis and trouvé que removing llms.txt from the model improved its accuracy — leur conclusion: “LLMs.txt doesn’t sembler to directement impact AI citation frequency. Au moins pas yet.” A Moteur de recherche Land 10-site, 180-day study saw 8 of 10 sites montrer aucun measurable modifier, and the two “winners” had confounding changements (PR coverage, nouveau FAQ pages, technical fixes) vous pouvez’t separate from the llms.txt itself.

The meta-keywords comparison — fair or pas?

Mueller’s keywords-meta-tag analogy holds on the points que matter: les deux are site-owner self-descriptions, neither is verified, and les deux are ouvrir to gaming/cloaking — vous pourrait montrer un chose in llms.txt and un autre on lune page. Que manipulability is exactly pourquoi a serious AI search product won’t trust a self-description over the réel pages. The counter-argument (Carolyn Shelby) is que llms.txt au moins points to réel URLs que have to deliver — it’s “a spotlight, pas a wish liste.” Les deux are correct, and the practical result is the même: AI search systems don’t lean on it.

Quand it’s worth ajout — and quand to skip it

Ajouter it si vous run developer documentation or an API référence whose principal audience is AI coding agents (Claude Code, Cursor, Copilot, Codeium). That’s the utiliser cas it was construit pour, and it fonctionne.

Skip it — or au moins don’t expect AI-visibility renvoie — pour content sites, media, ecommerce, and ordinary business sites. The cost is ~20 minutes; the GEO benefit today is near-zero; and que 20 minutes is meilleur spent on content structure, FAQ coverage, and the AI search optimization fundamentals que en réalité déplacer AI visibility. The élevé adoption headlines (BuiltWith’s 844K+ figure) are inflated by platform auto-deployment — Mintlify rolled llms.txt out to tout its hosted docs sites at une fois. Adoption n’est pas usage; 97% of fichiers obtenir zero requêtes.

The security angle

Ce is underreported and réel: parce que AI agents are designed to trust what’s in llms.txt, the fichier is a natural attack surface. The Ahrefs study flagged bad actors probing llms.txt fichiers pour prompt-injection vulnerabilities. Si vous do publish un, treat it as security-sensitive: garder it in version contrôler, alert on unauthorized changements, and jamais put anything in it vous wouldn’t montrer publicly.

Implementing it (si vous decide to)

  • Placement: /llms.txt at the domain root, served as text/plain or text/markdown, HTTP 200.
  • Curate, don’t dump: 20–50 of votre la plupart important liens. Si vous liste everything, a model pourrait as bien explorer directement — curation is the entier point.
  • Optionally serve /llms-full.txt pour contexts que vouloir tout votre content.
  • Validate: charger the fichier into Claude, ChatGPT, or Perplexity and demander it to summarize votre site; alors vérifier server logs après a few weeks pour réel requêtes.
  • Generators exist pour VitePress (vitepress-plugin-llms), Docusaurus (docusaurus-plugin-llms), WordPress (“Website LLMs.txt”), Drupal, plus a Python CLI (llms_txt2ctx).

L’essentiel: low cost, currently near-zero GEO benefit, genuinely utile pour the coding-agent utiliser cas it was born pour. Vérifier votre propre server logs avant vous decide — si you’re seeing Claude-Code/Cursor trafic, llms.txt may aider les; si you’re chasing OAI-SearchBot and PerplexityBot, it isn’t the lever.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.