Hướng dẫn về llms.txt

llms.txt là một proposed Markdown file tại /llms.txt cho guiding AI các hệ thống. Google bỏ qua điều này, 97% of files nhận zero các yêu cầu, và Claude Code là đó real reader.

Xuất bản lần đầu: 24 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

llms.txt là một proposal by Jeremy Howard (Câu trả lời.AI) cho helping AI agents navigate các trang — 97% of published files nhận zero các yêu cầu, Google bỏ qua điều này, và Claude Code là đó chính consumer, không tìm kiếm bots.

đặc tả mô tả voluntary LLM-friendly trang web summary; nó không establish crawler compliance hoặc lập chỉ mục behavior. Evidence for this claim llms.txt is a community proposal for a Markdown file at /llms.txt that offers LLM-friendly site information; it is not a web standard. Scope: The proposal's own specification and stated purpose. Confidence: high · Verified: llms.txt proposal OpenAI hiện tại documents named bots và independent robots.txt settings thay vì. Evidence for this claim OpenAI documents robots.txt controls for its declared crawlers and does not list llms.txt as a crawler control. Scope: OpenAI's published crawler controls; absence is not proof that no internal system ever fetches the file. Confidence: high · Verified: OpenAI: Crawlers

Tóm tắt — llms.txt là proposal (Jeremy Howard, Câu trả lời.AI, Sept 3, 2024), không adopted tiêu chuẩn. nó curated Markdown chỉ mục tại /llms.txt, với tùy chọn đầy đủ-nội dung /llms-full.txt. bots mọi người hope là reading nó cho AI tìm kiếm — OAI-SearchBot, PerplexityBot — barely register; dominant consumer là Claude Code và khác coding agents. Google bỏ qua nó (Illyes confirmed; Mueller compared nó để từ khóa meta tag). Trong Ahrefs’ 137K-trang web nghiên cứu, 97% của files đã nhận zero các yêu cầu, và SE Xếp hạng tìm thấy đó removing llms.txt từ citation model improved của nó độ chính xác. nó worth thêm cho nhà phát triển tài liệu consumed by coding agents; cho GEO/AEO on thông thường trang web, dữ liệu không hỗ trợ nó. có cũng thực prompt-injection risk.

Điều gì llms.txt thực ra là

llms.txt là proposal — I muốn để lead với đó word vì nó toàn bộ story. có không RFC, không W3C blessing, không IETF xử lý. nó một well-reasoned suggestion từ Jeremy Howard (co-founder của Câu trả lời.AI và fast.ai), published September 3, 2024, đó caught on hard trong nhà phát triển-tài liệu world và đã nhận grafted onto SEO by community hungry cho AI-visibility shortcut.

Đó design vấn đề điều này targets là legitimate. As đó spec diễn đạt điều này, “Large language models increasingly rely on website information, but face a critical limitation: context windows are too small to handle most websites in their entirety.” (bản dịch) «Lớn language models increasingly rely on website information, nhưng face một cốt yếu limitation: context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety.» Và converting HTML để sạch, LLM-friendly text là, trong đó spec words, “difficult and imprecise.” (bản dịch) «difficult và imprecise.» Howard cách sửa là để let đó site tác giả pre-flatten đó nội dung they muốn models để see vào curated Markdown.

format

/llms.txt là ordered Markdown document:

  1. H1 (bắt buộc) — project hoặc trang web name.
  2. Blockquote — ngắn summary với Điểm mấu chốt information.
  3. Tùy chọn prose/lists — paragraphs hoặc bullets, nhưng không headings ở đây.
  4. H2-delimited link lists- [name](url): optional notes.
  5. ** ## Optional section** — phụ links model có thể skip Khi context là tight.

minimal ví dụ, từ spec:

# FastHTML

> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX,
> and fastcore's FT...

## Docs

- [FastHTML quick start](url): A brief overview of many features

## Optional

- [Starlette documentation](url): A subset useful for FastHTML development.

Hai companion conventions exist:

  • /llms-full.txt — đểàn bộ trang web nội dung flattened vào một Markdown file. Anthropic, Perplexity, và Stripe publish cả hai điều này và ngắn form.
  • ** .md URL convention** — offering page.md versions của riêng lẻ các trang so họ’re LLM-ready không có đầy đủ dump.

newer variant skips static files hoàn toàn: edge-phân phối Markdown qua nội dung negotiation. Cloudflare’s Markdown cho Agents converts bất kỳ HTML trang để sạch Markdown on fly Khi client gửi Accept: text/markdown — không theo-trang .md files, không /llms-full.txt để regenerate, không template thay đổi, vì CDN làm conversion tại edge (Cloudflare cites ~80% ít hơn tokens cho agent). nó paid feature — Pro plan và lên, không free tier — so on điều này trang web nó một Cloudflare upgrade away từ là switched on thay vì điều gì đó I’ve hand-được xây dựng. Liệu nó worth switching on là giống nhau open câu hỏi đó hangs over llms.txt itself: serving các crawler parallel Markdown copy của mỗi trang là convenient, nhưng nó một representation để giữ trong sync, và Bing Fabrice Canel có flagged doubled crawl load và cloaking-liền kề risk của edge-phân phối alternate nội dung. kiểm tra của bạn nhật ký cho thực agent demand trước khi flipping nó.

llms.txt so với robots.txt — họ không phải giống nhau điều

Đó single hầu hết phổ biến misconception là đó llms.txt là “robots.txt for AI.” (bản dịch) «robots.txt cho AI.» Điều này không, và conflating them dẫn đến bad decisions.

robots.txtllms.txt
PurposeAccess control — Điều gì các crawler có thể fetchnội dung hướng dẫn — Điều gì hữu ích để đọc
FormatDirectives (User-agent:, Disallow:)Markdown links + các mô tả
Enforced?Có, by compliant các crawlerKhông — advisory chỉ
TimingCrawl time, trước khi fetchingInference time, assembling context
tiêu chuẩn?Có, decades old, universalKhông — proposal
có thể nó block?Không

robots.txt là thực, enforced tiêu chuẩn. llms.txt là suggestion box đó phần lớn các hệ thống không phải reading tuy vậy. nếu bạn muốn để control AI crawler access — mà là tách biệt, legitimate goal — đó lives trong robots.txt và AI các crawler, không ở đây.

Ai thực ra đọc nó

Đây là section đó matters, vì nó nơi dữ liệu diverges hardest từ hype. Trong Ahrefs’ nghiên cứu của 137 210 domains (có thể 2026 traffic):

  • 97% của published llms.txt files đã nhận zero các yêu cầu. chỉ ~3% (về 1 100 domains) saw bất kỳ traffic tại all.
  • của files đó đã làm nhận requested: 96% của các yêu cầu nghĩ ra từ bots, 4% từ humans.
  • của bot các yêu cầu, 77% nghĩ ra từ non-AI tools — SEO auditors (Ahrefs itself among them), anonymous các crawler, tech-profiling bots. SEO irony ghi itself.
  • chỉ ~19,5% nghĩ ra từ AI tools, và sau khi bạn break đó xuống nó nhận tệ hơn cho GEO theory:
  • AI agents & infrastructure (coding agents): ~10,5%
  • Training các crawler (GPTBot led tại 4,51%): ~5,3%
  • AI assistants: ~2,5%
  • AI retrieval bots — ones đó xây dựng citation indexes: 1,1%.

Đọc đó cuối cùng line twice. Đó bots hầu hết mọi người thêm llms.txt cho — OAI-SearchBot, PerplexityBot — “barely registered.” (bản dịch) «barely registered.» Đó dominant AI consumer là Claude Code, Anthropic coding agent, mà đó nghiên cứu được tìm thấy “outfetched every AI retrieval bot, assistant, and training crawler except GPTBot.” (bản dịch) «outfetched mỗi AI retrieval bot, assistant, và training crawler except GPTBot.» Và GPTBot là một training crawler đó crawl mọi thứ regardless — của nó presence không evidence of intentional llms.txt phân tích cú pháp.

Other independent measurements land trong đó giống nhau place. OtterlyAI watched một site cho 90 days: of 62 100+ AI bot visits, chính xác 84 hit /llms.txt — về 0,1%, performing “3x worse than average pages.” (bản dịch) «3x tệ hơn average các trang.» MỘT tách biệt analysis of 515M+ LLM bot events over 90 days được tìm thấy 408 các yêu cầu total targeting /llms.txt — “statistically negligible.” (bản dịch) «không đáng kể về mặt thống kê.»

takeaway là nhà phát triển-tài liệu-so với-GEO split: llms.txt hoạt động cho Điều gì Howard được xây dựng nó cho (coding agents navigating API tài liệu), và nó largely không nhận touched by tìm kiếm-citation bots SEO community wants nó để reach.

Điều gì providers có đã nói

Google — không, và không plans. Gary Illyes confirmed tại Tìm kiếm Central Trực tiếp (July 2025) đó Google không hỗ trợ llms.txt và không planning để. John Mueller, đó hầu hết-quoted voice ở đây, compared điều này để đó từ khóa meta tag (April 2025) và called điều này một “temporary crutch, perhaps to save some tokens” (bản dịch) «tạm thời crutch, perhaps để save some tokens» cho AI coding tools — explicitly không một tìm kiếm-visibility mechanism. Google Có thể 2026 hướng dẫn stated machine-readable files như llms.txt không necessary cho AI Overviews hoặc AI Chế độ, mà “continue to rely on traditional SEO signals.” (bản dịch) «continue để rely on truyền thống SEO các tín hiệu.» (Google đã làm briefly publish an llms.txt on của nó dev tài liệu on December 3, 2025, thì pulled điều này đó giống nhau day — sau đó attributed để một CMS cập nhật, không một strategy shift.)

OpenAI — không có gì. Không announcement đó ChatGPT, GPTBot, hoặc OAI-SearchBot parse điều này. Máy chủ logs contradict có ý nghĩa usage; OAI-SearchBot “barely registered.” (bản dịch) «barely registered.» OpenAI points mọi người để robots.txt cho crawler control.

Anthropic — đó interesting một. Anthropic publishes cả hai llms.txtllms-full.txt tại docs.anthropic.com/llms.txt, và Claude Code demonstrably fetches other các trang’ llms.txt files — đây là đó largest AI consumer trong đó dữ liệu. Nhưng đó là một coding-tool story. có không xác nhận đó Claude.ai tìm kiếm hoặc citation layer honors llms.txt. “Anthropic’s coding agent reads it” (bản dịch) «Anthropic coding agent đọc điều này» và “Anthropic’s search index trusts it” (bản dịch) «Anthropic tìm kiếm chỉ mục trusts điều này» là khác nhau claims, và chỉ đó đầu tiên là supported.

Perplexity — claims có, dữ liệu says rarely. Perplexity có stated điều này retrieves llms.txt và dùng điều này để “prioritize page selection.” (bản dịch) «prioritize trang selection.» Nhưng proactive fetching là đó catch: PerplexityBot “barely registered” (bản dịch) «barely registered» trong đó yêu cầu dữ liệu, và có “almost zero activity from PerplexityBot requesting llms.txt files proactively.” (bản dịch) «gần như zero activity từ PerplexityBot requesting llms.txt files chủ động.» Paste an llms.txt URL vào Perplexity và điều này đọc điều này fine; autonomous, proactive dùng là đó khoảng trống giữa policy và practice.

Bing/Microsoft, Apple, Meta — không công khai position. Treat as unconfirmed.

Làm nó move AI citations?

Không demonstrated benefit. SE Xếp hạng modeled 300 000 domains với an XGBoost regression + SHAP analysis và được tìm thấy đó removing llms.txt từ đó model improved của nó độ chính xác — của họ conclusion: “LLMs.txt doesn’t seem to directly impact AI citation frequency. At least not yet.” (bản dịch) «LLMs.txt không seem để trực tiếp impact AI citation frequency. Ít nhất chưa.» MỘT Search Engine Land 10-site, 180-day nghiên cứu saw 8 of 10 các trang cho thấy không measurable thay đổi, và đó hai “winners” đã có confounding thay đổi (PR coverage, new FAQ các trang, kỹ thuật các cách sửa) bạn không thể tách biệt từ đó llms.txt itself.

meta-từ khóa so sánh — fair hoặc không?

Mueller từ khóa-meta-tag analogy holds on đó points đó quan trọng: cả hai là site-owner self-các mô tả, neither là verified, và cả hai là open để gaming/cloaking — bạn có thể cho thấy một điều trong llms.txt và một sản phẩm khác on đó trang. Đó manipulability là chính xác vì sao một serious AI tìm kiếm sản phẩm sẽ không trust một self-mô tả over đó thực tế các trang. Đó counter-argument (Carolyn Shelby) là đó llms.txt ít nhất points để real URLs đó có để deliver — đây là “a spotlight, not a wish list.” (bản dịch) «một spotlight, không một wish list.» Cả hai là right, và đó practical kết quả là đó giống nhau: AI tìm kiếm các hệ thống không lean on điều này.

Khi nó worth thêm — và Khi để skip nó

Thêm nó nếu bạn chạy nhà phát triển tài liệu hoặc API reference whose chính audience là AI coding agents (Claude Code, Cursor, Copilot, Codeium). đó sử dụng case nó là được xây dựng cho, và nó hoạt động.

Skip nó — hoặc ít nhất không expect AI-visibility trả về — cho nội dung các trang, media, ecommerce, và ordinary business các trang. cost là ~20 minutes; GEO benefit hôm nay là near-zero; và đó 20 minutes là tốt hơn spent on nội dung structure, FAQ coverage, và AI tìm kiếm optimization fundamentals đó thực ra move AI visibility. cao adoption headlines (BuiltWith 844K+ hình) là inflated by nền tảng auto-deployment — Mintlify rolled llms.txt out để all của nó hosted tài liệu các trang tại sau khi. Adoption không phải usage; 97% của files nhận zero các yêu cầu.

security angle

Đây là underreported và thực: vì AI agents là designed để trust Điều gì trong llms.txt, file là natural attack surface. Ahrefs nghiên cứu flagged bad actors probing llms.txt files cho prompt-injection vulnerabilities. nếu bạn làm publish một, treat nó as security-sensitive: giữ nó trong version control, alert on Không được ủy quyền thay đổi, và không bao giờ put bất cứ điều gì trong nó bạn sẽ không hiển thị publicly.

Implementing nó (nếu bạn quyết định để)

  • Placement: /llms.txt tại đó domain root, phân phối as text/plain hoặc text/markdown, HTTP 200.
  • Curate, không dump: 20–50 of của bạn hầu hết quan trọng links. Nếu bạn list mọi thứ, một model có thể as well crawl trực tiếp — curation là đó entire point.
  • Tùy chọn serve /llms-full.txt cho contexts đó muốn all nội dung của bạn.
  • Validate: load đó file vào Claude, ChatGPT, hoặc Perplexity và ask điều này để summarize trang web của bạn; thì kiểm tra máy chủ logs sau vài weeks cho real các yêu cầu.
  • Generators exist cho VitePress (vitepress-plugin-llms), Docusaurus (docusaurus-plugin-llms), WordPress (“Website LLMs.txt” (bản dịch) «Website LLMs.txt»), Drupal, plus một Python CLI (llms_txt2ctx).

bottom line: thấp cost, hiện tại near-zero GEO benefit, genuinely hữu ích cho coding-agent sử dụng case nó là born cho. kiểm tra của bạn own máy chủ nhật ký trước khi bạn quyết định — nếu bạn’re seeing Claude-Code/Cursor traffic, llms.txt có thể help them; nếu bạn’re chasing OAI-SearchBot và PerplexityBot, nó không phải lever.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.