Hướng dẫn về llms.txt
llms.txt là một proposed Markdown file tại /llms.txt cho guiding AI các hệ thống. Google bỏ qua điều này, 97% of files nhận zero các yêu cầu, và Claude Code là đó real reader.
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này
- Công cụ trực tuyến liên quanllms.txt Generator + Validator
llms.txt là một proposal by Jeremy Howard (Câu trả lời.AI) cho helping AI agents navigate các trang — 97% of published files nhận zero các yêu cầu, Google bỏ qua điều này, và Claude Code là đó chính consumer, không tìm kiếm bots.
llms.txt là community proposal cho Markdown file tại /llms.txt, không adopted web tiêu chuẩn. Evidence for this claim llms.txt is a community proposal for a Markdown file at /llms.txt that offers LLM-friendly site information; it is not a web standard. Scope: The proposal's own specification and stated purpose. Confidence: high · Verified: llms.txt proposal OpenAI’s published crawler controls sử dụng robots.txt và không document llms.txt as control. Evidence for this claim OpenAI documents robots.txt controls for its declared crawlers and does not list llms.txt as a crawler control. Scope: OpenAI's published crawler controls; absence is not proof that no internal system ever fetches the file. Confidence: high · Verified: OpenAI: Crawlers
Tóm tắt — llms.txt là Markdown file bạn put tại
/llms.txtđó hands AI các hệ thống ngắn, curated list của bạn phần lớn quan trọng các trang. nhà phát triển named Jeremy Howard proposed nó trong 2024. nó nice ý tưởng, nhưng trên thực tế gần như không ai đọc nó — Google bỏ qua nó outright, và main điều đó làm đọc nó là AI coding assistants, không tìm kiếm bots phần lớn mọi người là hoping để reach. nó costs little để thêm và probably sẽ không move của bạn AI visibility.
Điều gì llms.txt là
llms.txt là đơn giản Markdown file bạn place tại root của bạn trang web —
yoursite.com/llms.txt. Bên trong, bạn ghi của bạn trang web name, một-line
mô tả, và sau đó ngắn list của links để của bạn phần lớn quan trọng các trang, mỗi
với nhanh note về Điều gì nó covers. ý tưởng là đó AI hệ thống có thể đọc
đó một file và immediately understand Điều gì của bạn trang web là và nơi good
stuff lives, thay vì có để crawl và parse mỗi trang.
nó là proposed by Jeremy Howard — person behind fast.ai ( popular deep learning course) và Câu trả lời.AI — lại trong September 2024.
Điều gì nó supposed để làm so với. Điều gì nó thực ra làm
Đó pitch: AI tools có limited “context windows” (bản dịch) «context windows» (they có thể chỉ đọc so nhiều tại khi), và chuyển thành messy HTML vào sạch text là hard. So vì sao không let đó site owner hand đó model một tidy, pre-được viết map?
reality là nhiều hơn sobering. major AI tìm kiếm providers haven’t adopted nó. Google có flatly đã nói nó không sử dụng nó. và Khi Ahrefs studied 137 000 các trang, 97% của published llms.txt files đã nhận zero các yêu cầu trong month — as trong, không có gì fetched them tại all.
Đó một place điều này genuinely nhận dùng là AI coding assistants — tools như Claude Code đó nhà phát triển dùng để ghi software. Những đọc llms.txt để tìm của họ way khoảng kỹ thuật tài liệu quickly. đó là đó original dùng case, và điều này hoạt động. đây là chỉ very khác nhau từ “this will get my business cited in ChatGPT.” (bản dịch) «này sẽ nhận my business cited trong ChatGPT.»
là nó như robots.txt?
MỘT lot of mọi người assume llms.txt là “robots.txt for AI.” (bản dịch) «robots.txt cho AI.» Điều này không.
- robots.txt controls access — nó tells các crawler Điều gì họ’re được phép để fetch, và well-behaved bots obey nó.
- llms.txt offers hướng dẫn — nó suggests Điều gì worth reading, và không có gì là obligated để listen. nó có thể’t block bất cứ điều gì.
nếu của bạn goal là để giữ AI bots out, llms.txt làm không có gì — đó job cho robots.txt và của bạn CDN.
nên bạn bother?
nếu bạn chạy nhà phát triển tài liệu đó coding agents đọc, sure — nó cheap và nó helps them. nếu bạn’re thông thường business hoping llms.txt sẽ boost của bạn visibility trong AI các câu trả lời, honest câu trả lời hôm nay là: không expect kết quả. Spend đó time on clear nội dung và structure thay vì.
Muốn spec, provider-by-provider breakdown, adoption dữ liệu, và security angle? Chuyển để Nâng cao tab.
đặc tả mô tả voluntary LLM-friendly trang web summary; nó không establish crawler compliance hoặc lập chỉ mục behavior. Evidence for this claim llms.txt is a community proposal for a Markdown file at /llms.txt that offers LLM-friendly site information; it is not a web standard. Scope: The proposal's own specification and stated purpose. Confidence: high · Verified: llms.txt proposal OpenAI hiện tại documents named bots và independent robots.txt settings thay vì. Evidence for this claim OpenAI documents robots.txt controls for its declared crawlers and does not list llms.txt as a crawler control. Scope: OpenAI's published crawler controls; absence is not proof that no internal system ever fetches the file. Confidence: high · Verified: OpenAI: Crawlers
Tóm tắt — llms.txt là proposal (Jeremy Howard, Câu trả lời.AI, Sept 3, 2024), không adopted tiêu chuẩn. nó curated Markdown chỉ mục tại
/llms.txt, với tùy chọn đầy đủ-nội dung/llms-full.txt. bots mọi người hope là reading nó cho AI tìm kiếm — OAI-SearchBot, PerplexityBot — barely register; dominant consumer là Claude Code và khác coding agents. Google bỏ qua nó (Illyes confirmed; Mueller compared nó để từ khóa meta tag). Trong Ahrefs’ 137K-trang web nghiên cứu, 97% của files đã nhận zero các yêu cầu, và SE Xếp hạng tìm thấy đó removing llms.txt từ citation model improved của nó độ chính xác. nó worth thêm cho nhà phát triển tài liệu consumed by coding agents; cho GEO/AEO on thông thường trang web, dữ liệu không hỗ trợ nó. có cũng thực prompt-injection risk.
Điều gì llms.txt thực ra là
llms.txt là proposal — I muốn để lead với đó word vì nó toàn bộ story. có không RFC, không W3C blessing, không IETF xử lý. nó một well-reasoned suggestion từ Jeremy Howard (co-founder của Câu trả lời.AI và fast.ai), published September 3, 2024, đó caught on hard trong nhà phát triển-tài liệu world và đã nhận grafted onto SEO by community hungry cho AI-visibility shortcut.
Đó design vấn đề điều này targets là legitimate. As đó spec diễn đạt điều này, “Large language models increasingly rely on website information, but face a critical limitation: context windows are too small to handle most websites in their entirety.” (bản dịch) «Lớn language models increasingly rely on website information, nhưng face một cốt yếu limitation: context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety.» Và converting HTML để sạch, LLM-friendly text là, trong đó spec words, “difficult and imprecise.” (bản dịch) «difficult và imprecise.» Howard cách sửa là để let đó site tác giả pre-flatten đó nội dung they muốn models để see vào curated Markdown.
format
/llms.txt là ordered Markdown document:
- H1 (bắt buộc) — project hoặc trang web name.
- Blockquote — ngắn summary với Điểm mấu chốt information.
- Tùy chọn prose/lists — paragraphs hoặc bullets, nhưng không headings ở đây.
- H2-delimited link lists —
- [name](url): optional notes. - **
## Optionalsection** — phụ links model có thể skip Khi context là tight.
minimal ví dụ, từ spec:
# FastHTML
> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX,
> and fastcore's FT...
## Docs
- [FastHTML quick start](url): A brief overview of many features
## Optional
- [Starlette documentation](url): A subset useful for FastHTML development.Hai companion conventions exist:
/llms-full.txt— đểàn bộ trang web nội dung flattened vào một Markdown file. Anthropic, Perplexity, và Stripe publish cả hai điều này và ngắn form.- **
.mdURL convention** — offeringpage.mdversions của riêng lẻ các trang so họ’re LLM-ready không có đầy đủ dump.
newer variant skips static files hoàn toàn: edge-phân phối Markdown qua nội dung
negotiation. Cloudflare’s Markdown cho Agents
converts bất kỳ HTML trang để sạch Markdown on fly Khi client gửi
Accept: text/markdown — không theo-trang .md files, không /llms-full.txt để regenerate,
không template thay đổi, vì CDN làm conversion tại edge (Cloudflare
cites ~80% ít hơn tokens cho agent). nó paid feature — Pro plan và lên, không
free tier — so on điều này trang web nó một Cloudflare upgrade away từ là switched
on thay vì điều gì đó I’ve hand-được xây dựng. Liệu nó worth switching on là
giống nhau open câu hỏi đó hangs over llms.txt itself: serving các crawler parallel
Markdown copy của mỗi trang là convenient, nhưng nó một representation để giữ trong
sync, và Bing Fabrice Canel có flagged doubled crawl load và
cloaking-liền kề risk của edge-phân phối alternate nội dung. kiểm tra của bạn nhật ký cho thực
agent demand trước khi flipping nó.
llms.txt so với robots.txt — họ không phải giống nhau điều
Đó single hầu hết phổ biến misconception là đó llms.txt là “robots.txt for AI.” (bản dịch) «robots.txt cho AI.» Điều này không, và conflating them dẫn đến bad decisions.
| robots.txt | llms.txt | |
|---|---|---|
| Purpose | Access control — Điều gì các crawler có thể fetch | nội dung hướng dẫn — Điều gì hữu ích để đọc |
| Format | Directives (User-agent:, Disallow:) | Markdown links + các mô tả |
| Enforced? | Có, by compliant các crawler | Không — advisory chỉ |
| Timing | Crawl time, trước khi fetching | Inference time, assembling context |
| tiêu chuẩn? | Có, decades old, universal | Không — proposal |
| có thể nó block? | Có | Không |
robots.txt là thực, enforced tiêu chuẩn. llms.txt là suggestion box đó phần lớn các hệ thống không phải reading tuy vậy. nếu bạn muốn để control AI crawler access — mà là tách biệt, legitimate goal — đó lives trong robots.txt và AI các crawler, không ở đây.
Ai thực ra đọc nó
Đây là section đó matters, vì nó nơi dữ liệu diverges hardest từ hype. Trong Ahrefs’ nghiên cứu của 137 210 domains (có thể 2026 traffic):
- 97% của published llms.txt files đã nhận zero các yêu cầu. chỉ ~3% (về 1 100 domains) saw bất kỳ traffic tại all.
- của files đó đã làm nhận requested: 96% của các yêu cầu nghĩ ra từ bots, 4% từ humans.
- của bot các yêu cầu, 77% nghĩ ra từ non-AI tools — SEO auditors (Ahrefs itself among them), anonymous các crawler, tech-profiling bots. SEO irony ghi itself.
- chỉ ~19,5% nghĩ ra từ AI tools, và sau khi bạn break đó xuống nó nhận tệ hơn cho GEO theory:
- AI agents & infrastructure (coding agents): ~10,5%
- Training các crawler (GPTBot led tại 4,51%): ~5,3%
- AI assistants: ~2,5%
- AI retrieval bots — ones đó xây dựng citation indexes: 1,1%.
Đọc đó cuối cùng line twice. Đó bots hầu hết mọi người thêm llms.txt cho — OAI-SearchBot, PerplexityBot — “barely registered.” (bản dịch) «barely registered.» Đó dominant AI consumer là Claude Code, Anthropic coding agent, mà đó nghiên cứu được tìm thấy “outfetched every AI retrieval bot, assistant, and training crawler except GPTBot.” (bản dịch) «outfetched mỗi AI retrieval bot, assistant, và training crawler except GPTBot.» Và GPTBot là một training crawler đó crawl mọi thứ regardless — của nó presence không evidence of intentional llms.txt phân tích cú pháp.
Other independent measurements land trong đó giống nhau place. OtterlyAI watched một site
cho 90 days: of 62 100+ AI bot visits, chính xác 84 hit /llms.txt — về 0,1%,
performing “3x worse than average pages.” (bản dịch) «3x tệ hơn average các trang.» MỘT tách biệt analysis of 515M+ LLM bot
events over 90 days được tìm thấy 408 các yêu cầu total targeting /llms.txt —
“statistically negligible.” (bản dịch) «không đáng kể về mặt thống kê.»
takeaway là nhà phát triển-tài liệu-so với-GEO split: llms.txt hoạt động cho Điều gì Howard được xây dựng nó cho (coding agents navigating API tài liệu), và nó largely không nhận touched by tìm kiếm-citation bots SEO community wants nó để reach.
Điều gì providers có đã nói
Google — không, và không plans. Gary Illyes confirmed tại Tìm kiếm Central Trực tiếp (July 2025) đó Google không hỗ trợ llms.txt và không planning để. John Mueller, đó hầu hết-quoted voice ở đây, compared điều này để đó từ khóa meta tag (April 2025) và called điều này một “temporary crutch, perhaps to save some tokens” (bản dịch) «tạm thời crutch, perhaps để save some tokens» cho AI coding tools — explicitly không một tìm kiếm-visibility mechanism. Google Có thể 2026 hướng dẫn stated machine-readable files như llms.txt không necessary cho AI Overviews hoặc AI Chế độ, mà “continue to rely on traditional SEO signals.” (bản dịch) «continue để rely on truyền thống SEO các tín hiệu.» (Google đã làm briefly publish an llms.txt on của nó dev tài liệu on December 3, 2025, thì pulled điều này đó giống nhau day — sau đó attributed để một CMS cập nhật, không một strategy shift.)
OpenAI — không có gì. Không announcement đó ChatGPT, GPTBot, hoặc OAI-SearchBot parse điều này. Máy chủ logs contradict có ý nghĩa usage; OAI-SearchBot “barely registered.” (bản dịch) «barely registered.» OpenAI points mọi người để robots.txt cho crawler control.
Anthropic — đó interesting một. Anthropic publishes cả hai llms.txt và
llms-full.txt tại docs.anthropic.com/llms.txt, và Claude Code demonstrably
fetches other các trang’ llms.txt files — đây là đó largest AI consumer trong đó dữ liệu. Nhưng
đó là một coding-tool story. có không xác nhận đó Claude.ai tìm kiếm hoặc
citation layer honors llms.txt. “Anthropic’s coding agent reads it” (bản dịch) «Anthropic coding agent đọc điều này» và “Anthropic’s
search index trusts it” (bản dịch) «Anthropic tìm kiếm chỉ mục trusts điều này» là khác nhau claims, và chỉ đó đầu tiên là supported.
Perplexity — claims có, dữ liệu says rarely. Perplexity có stated điều này retrieves llms.txt và dùng điều này để “prioritize page selection.” (bản dịch) «prioritize trang selection.» Nhưng proactive fetching là đó catch: PerplexityBot “barely registered” (bản dịch) «barely registered» trong đó yêu cầu dữ liệu, và có “almost zero activity from PerplexityBot requesting llms.txt files proactively.” (bản dịch) «gần như zero activity từ PerplexityBot requesting llms.txt files chủ động.» Paste an llms.txt URL vào Perplexity và điều này đọc điều này fine; autonomous, proactive dùng là đó khoảng trống giữa policy và practice.
Bing/Microsoft, Apple, Meta — không công khai position. Treat as unconfirmed.
Làm nó move AI citations?
Không demonstrated benefit. SE Xếp hạng modeled 300 000 domains với an XGBoost regression + SHAP analysis và được tìm thấy đó removing llms.txt từ đó model improved của nó độ chính xác — của họ conclusion: “LLMs.txt doesn’t seem to directly impact AI citation frequency. At least not yet.” (bản dịch) «LLMs.txt không seem để trực tiếp impact AI citation frequency. Ít nhất chưa.» MỘT Search Engine Land 10-site, 180-day nghiên cứu saw 8 of 10 các trang cho thấy không measurable thay đổi, và đó hai “winners” đã có confounding thay đổi (PR coverage, new FAQ các trang, kỹ thuật các cách sửa) bạn không thể tách biệt từ đó llms.txt itself.
meta-từ khóa so sánh — fair hoặc không?
Mueller từ khóa-meta-tag analogy holds on đó points đó quan trọng: cả hai là site-owner self-các mô tả, neither là verified, và cả hai là open để gaming/cloaking — bạn có thể cho thấy một điều trong llms.txt và một sản phẩm khác on đó trang. Đó manipulability là chính xác vì sao một serious AI tìm kiếm sản phẩm sẽ không trust một self-mô tả over đó thực tế các trang. Đó counter-argument (Carolyn Shelby) là đó llms.txt ít nhất points để real URLs đó có để deliver — đây là “a spotlight, not a wish list.” (bản dịch) «một spotlight, không một wish list.» Cả hai là right, và đó practical kết quả là đó giống nhau: AI tìm kiếm các hệ thống không lean on điều này.
Khi nó worth thêm — và Khi để skip nó
Thêm nó nếu bạn chạy nhà phát triển tài liệu hoặc API reference whose chính audience là AI coding agents (Claude Code, Cursor, Copilot, Codeium). đó sử dụng case nó là được xây dựng cho, và nó hoạt động.
Skip nó — hoặc ít nhất không expect AI-visibility trả về — cho nội dung các trang, media, ecommerce, và ordinary business các trang. cost là ~20 minutes; GEO benefit hôm nay là near-zero; và đó 20 minutes là tốt hơn spent on nội dung structure, FAQ coverage, và AI tìm kiếm optimization fundamentals đó thực ra move AI visibility. cao adoption headlines (BuiltWith 844K+ hình) là inflated by nền tảng auto-deployment — Mintlify rolled llms.txt out để all của nó hosted tài liệu các trang tại sau khi. Adoption không phải usage; 97% của files nhận zero các yêu cầu.
security angle
Đây là underreported và thực: vì AI agents là designed để trust Điều gì trong llms.txt, file là natural attack surface. Ahrefs nghiên cứu flagged bad actors probing llms.txt files cho prompt-injection vulnerabilities. nếu bạn làm publish một, treat nó as security-sensitive: giữ nó trong version control, alert on Không được ủy quyền thay đổi, và không bao giờ put bất cứ điều gì trong nó bạn sẽ không hiển thị publicly.
Implementing nó (nếu bạn quyết định để)
- Placement:
/llms.txttại đó domain root, phân phối astext/plainhoặctext/markdown, HTTP 200. - Curate, không dump: 20–50 of của bạn hầu hết quan trọng links. Nếu bạn list mọi thứ, một model có thể as well crawl trực tiếp — curation là đó entire point.
- Tùy chọn serve
/llms-full.txtcho contexts đó muốn all nội dung của bạn. - Validate: load đó file vào Claude, ChatGPT, hoặc Perplexity và ask điều này để summarize trang web của bạn; thì kiểm tra máy chủ logs sau vài weeks cho real các yêu cầu.
- Generators exist cho VitePress (
vitepress-plugin-llms), Docusaurus (docusaurus-plugin-llms), WordPress (“Website LLMs.txt” (bản dịch) «Website LLMs.txt»), Drupal, plus một Python CLI (llms_txt2ctx).
bottom line: thấp cost, hiện tại near-zero GEO benefit, genuinely hữu ích cho coding-agent sử dụng case nó là born cho. kiểm tra của bạn own máy chủ nhật ký trước khi bạn quyết định — nếu bạn’re seeing Claude-Code/Cursor traffic, llms.txt có thể help them; nếu bạn’re chasing OAI-SearchBot và PerplexityBot, nó không phải lever.
AI summary
condensed take on Nâng cao version:
- llms.txt là một proposal, không một tiêu chuẩn — Jeremy Howard (Câu trả lời.AI/fast.ai),
Sept 3, 2024. MỘT curated Markdown chỉ mục tại
/llms.txt; tùy chọn đầy đủ-nội dung/llms-full.txt. - Điều này không phải robots.txt cho AI. robots.txt controls access và là enforced; llms.txt là advisory và không thể block bất cứ điều gì.
- Gần như không ai đọc điều này. Ahrefs (137K các trang): 97% of files đã nhận zero các yêu cầu. Of files đó đã là requested, 77% of bot hits đã là non-AI tools (SEO auditors), và AI retrieval bots — đó citation-builders — đã là chỉ 1,1%.
- Claude Code là đó dominant AI consumer, không tìm kiếm bots. OAI-SearchBot và PerplexityBot “barely registered.” (bản dịch) «barely registered.»
- Google bỏ qua điều này. Illyes confirmed không hỗ trợ/plans; Mueller compared điều này để đó từ khóa meta tag. OpenAI: không có gì. Perplexity: claims hỗ trợ, nhưng proactive fetching là minimal. Anthropic publishes của nó own + Claude Code đọc điều này, nhưng tìm kiếm-layer dùng là unconfirmed.
- Không citation benefit shown. SE Xếp hạng (300K domains): removing llms.txt từ đó model improved độ chính xác.
- Worth điều này cho nhà phát triển tài liệu consumed by coding agents; skip điều này as một GEO play cho thông thường các trang. Prompt-injection risk — treat điều này as security-sensitive.
Tài liệu chính thức
Chính-nguồn material — spec, proposal, và providers’ own files và statements.
** spec & proposal**
- llmstxt.org — đặc tả: file format,
/llms.txtso với/llms-full.txt, và inference-time rationale. - gốc proposal (Jeremy Howard, Sept 3, 2024) — Câu trả lời.AI announcement post.
- AnswerDotAI/llms-txt on GitHub —
repo, tooling, và
llms_txt2ctxCLI.
Providers’ own files
- Anthropic llms.txt — trực tiếp ví dụ
(Anthropic cũng publishes
llms-full.txt). - Perplexity llms.txt — một trực tiếp ví dụ.
Google chính thức position
- Google có stated machine-readable files như llms.txt là không necessary cho AI Overviews hoặc AI Chế độ, mà rely on truyền thống SEO các tín hiệu (có thể 2026 hướng dẫn); Gary Illyes confirmed không hỗ trợ và không plans tại Tìm kiếm Central Trực tiếp (July 2025). See công cụ tìm kiếm Journal coverage của John Mueller statement cho phần lớn-cited articulation.
Quotes từ nguồn
On—record statements on llms.txt — từ Google, spec creator, và researchers ai measured nó.
John Mueller, Google Search Advocate (April 17, 2025)
- “AFAIK none of the AI services have said they’re using LLMs.TXT (and you can tell when you look at your server logs that they don’t even check for it). To me, it’s comparable to the keywords meta tag – this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)” (bản dịch) «AFAIK none of đó AI services có đã nói họ là dùng LLMs.TXT (và bạn có thể tell khi bạn xem máy chủ của bạn logs đó they không ngay cả kiểm tra cho điều này). Để me, đây là comparable để đó từ khóa meta tag – này là điều gì một site-owner claims của họ site là về … (Là đó site thực sự như đó? well, bạn có thể kiểm tra điều này. Tại đó point, vì sao không chỉ kiểm tra đó site trực tiếp?)» Coverage
Gary Illyes, Google Search Central (Tìm kiếm Central Trực tiếp, July 2025)
- Confirmed Google không hỗ trợ llms.txt và có không plans để. (Reported từ event; không chính transcript URL identified — xác nhận trước khi treating as cuối.)
Jeremy Howard, creator (Câu trả lời.AI)
- “A proposal to standardise on using an
/llms.txtfile to provide information to help LLMs use a website at inference time.” (bản dịch) «MỘT proposal để standardise on dùng an/llms.txtfile để cung cấp information để help LLMs dùng một website tại inference time.» llmstxt.org - Từ đó spec, on đó cốt lõi vấn đề: “Large language models increasingly rely on website information, but face a critical limitation: context windows are too small to handle most websites in their entirety.” (bản dịch) «Lớn language models increasingly rely on website information, nhưng face một cốt yếu limitation: context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety.» llmstxt.org
Brett Tabke, Pubcon / WebmasterWorld (March 2025)
- “we just don’t need people thinking they are different from any other spider.” (bản dịch) «we chỉ không cần mọi người thinking they là khác nhau từ bất kỳ other spider.» Search Engine Land
Carolyn Shelby, công cụ tìm kiếm Land (July 9, 2025 — dissenting view)
- “Llms.txt curates a list of real URLs, and the content has to exist – and deliver – when the model gets there.” (bản dịch) «Llms.txt curates một list of real URLs, và đó nội dung có để exist – và deliver – khi đó model nhận ở đó.» và “Think of llms.txt like a treasure map for AI systems – one you draw yourself. It’s not a wish list. It’s a spotlight.” (bản dịch) «Think of llms.txt như một treasure map cho AI các hệ thống – một bạn draw yourself. đây là không một wish list. đây là một spotlight.» Search Engine Land
nên bạn implement llms.txt? — decision checklist
Hoạt động top để bottom. đầu tiên “có” đó fits thường settles nó.
- là của bạn chính audience AI coding agents (Claude Code, Cursor, Copilot, Codeium) reading nhà phát triển tài liệu / API reference? → Có, thêm nó. Đây là sử dụng case nó là được xây dựng cho và nó hoạt động.
- là bạn thêm nó purely để improve AI tìm kiếm citations cho nội dung, media, ecommerce, hoặc chung business trang web? → không expect kết quả. dữ liệu hiển thị không citation benefit; spend time on nội dung structure thay vì.
- có bạn checked của bạn máy chủ nhật ký đầu tiên? Look cho
Claude-Code/Cursortraffic (llms.txt có thể help them) so vớiOAI-SearchBot/PerplexityBot(nó sẽ không move đó needle). - là bạn treating nó as robots.txt cho AI? → Dừng. nó chặn không có gì. sử dụng robots.txt + của bạn CDN cho access control.
nếu bạn làm publish một:
- Place nó tại
/llms.txt(root), phục vụtext/plain/text/markdown, HTTP 200. - Curate 20–50 của bạn phần lớn quan trọng các URL — không dump sitemap.
- giữ nó trong version control và alert on Không được ủy quyền thay đổi (prompt-injection risk — agents là được xây dựng để trust nó).
- Put không có gì trong nó bạn sẽ không hiển thị publicly.
- Tùy chọn phục vụ
/llms-full.txtcho đầy đủ-nội dung contexts. - Validate: load URL vào Claude/ChatGPT/Perplexity, ask cho trang web summary; recheck máy chủ nhật ký sau khi một vài weeks cho thực các yêu cầu.
Provider hỗ trợ — bảng tra nhanh
Ai thực ra hỗ trợ llms.txt (as của mid-2026)
| Provider | Tìm kiếm/citation hỗ trợ? | Reality |
|---|---|---|
| Google (Gemini / AI Overviews) | Không | Illyes confirmed không hỗ trợ/plans; Mueller compared điều này để đó từ khóa meta tag; AI Overviews rely on truyền thống các tín hiệu |
| OpenAI (ChatGPT / OAI-SearchBot) | Không có ý nghĩa hỗ trợ | Không announcement; OAI-SearchBot “barely registered” (bản dịch) «barely registered» trong yêu cầu dữ liệu |
| Anthropic (Claude Code) | Có — cho đó coding agent | Claude Code là đó top AI consumer of llms.txt; Claude.ai tìm kiếm-layer dùng unconfirmed |
| Perplexity | Claims có, yếu trong thực tế | Trạng thái điều này prioritizes trang selection, nhưng proactive PerplexityBot fetching là near-zero |
| Apple (Applebot) | Không statement | Không công khai position, không evidence of crawling |
| Bing / Microsoft | Không statement | Không công khai position được tìm thấy |
Fast facts
- Proposed: Jeremy Howard (Câu trả lời.AI/fast.ai), Sept 3, 2024. proposal, không adopted tiêu chuẩn.
- Files:
/llms.txt(ngắn, curated) và tùy chọn/llms-full.txt(đầy đủ nội dung flattened để Markdown). - Format: H1 name → blockquote summary → H2 link lists (
- [name](url): note) →## Optionalsection. - so với robots.txt: robots.txt = access control, enforced. llms.txt = hướng dẫn, advisory, chặn không có gì.
- ** number:** 97% của published files nhận zero các yêu cầu (Ahrefs, 137K các trang).
- Dominant AI reader: Claude Code, không tìm kiếm bots.
- Citation benefit: none demonstrated (SE Xếp hạng: removing nó improved model độ chính xác).
- Security: prompt-injection risk — version-control nó, alert on thay đổi.
llms.txt mistakes để tránh
Treating file as access-control tiêu chuẩn
llms.txt không cho phép hoặc block crawling. sử dụng robots.txt, authentication, và
khác thực access controls cho đó job.
Promising AI-visibility lift
Provider hỗ trợ và observed fetching là limited. Mô tả file as tùy chọn navigation aid, không xếp hạng hoặc citation lever.
Xuất bản sensitive hoặc riêng tư các URL
file là công khai và easy để discover. bao gồm chỉ các tài nguyên dự kiến cho công khai consumption; không bao giờ sử dụng obscurity as protection.
Generating uncurated sitemap trong Markdown
proposal là phần lớn hữu ích as ngắn, human-readable hướng dẫn để quan trọng các tài nguyên. huge dump adds maintenance cost không có clarifying mà các trang quan trọng.
phổ biến llms.txt các vấn đề
file trả về HTML thay vì Markdown
Symptom: /llms.txt hiển thị application shell, branded 404, hoặc chuyển hướng đích.
có khả năng nguyên nhân: catch-all route hoặc CDN rule intercepts text path. khắc phục: phục vụ
file trực tiếp tại root với thành công phản hồi và đơn giản text hoặc Markdown
nội dung, sau đó yêu cầu nó không có cookies.
Links trong file fail hoặc chuyển hướng repeatedly
Symptom: reader không thể retrieve listed các tài nguyên tại của họ canonical các URL. có khả năng nguyên nhân: file là generated từ stale navigation hoặc relative paths. khắc phục: validate mỗi URL, replace retired locations, và ưu tiên ổn định canonical HTTPS các URL.
Không bots yêu cầu file
Symptom: nhật ký hiển thị không traffic để /llms.txt. có khả năng nguyên nhân: các hệ thống bạn care
về không hỗ trợ hoặc discover proposal. khắc phục: xác nhận log coverage và leave
file as thấp-cost thử nghiệm; không thêm crawler-cụ thể tricks hoặc infer trang web
vấn đề từ zero các yêu cầu.
Prompts cho curating llms.txt
Build a proposed llms.txt outline from this public URL inventory. Keep only canonical,
durable resources that help an agent understand the site or complete a task. Group them
under short Markdown headings, write one factual description per link, and flag URLs
that redirect, duplicate another page, require authentication, or may expose sensitive
information. Do not claim the file affects rankings or citations.
[paste inventory]Review this llms.txt file against the supplied crawl results. Report broken or
redirecting URLs, non-canonical links, descriptions unsupported by the destination,
missing high-value documentation, and sections that are too broad to be useful. Return
a corrected draft using only public URLs from the inputs.
[paste file and crawl results] Các framework cho deciding liệu để sử dụng llms.txt
cost–consumer–nội dung kiểm thử
- Cost: có thể team generate và maintain file với little ongoing hoạt động?
- Consumer: là ở đó identified agent hoặc workflow đó thực ra đọc nó?
- nội dung: Làm trang web có ổn định công khai tài liệu worth curating?
Implement Khi all three là credible. nếu consumer là hypothetical, treat file as thử nghiệm và cap maintenance budget.
giữ vai trò tách biệt
robots.txt: crawler access hướng dẫn.- XML sitemap: phát hiện URL Đối với tìm kiếm engines.
llms.txt: proposed curated reading hướng dẫn cho AI các hệ thống.- MCP: runtime giao thức qua mà agent có thể discover và invoke các tài nguyên hoặc tools.
Một file nên không là evaluated as nếu nó thực hiện một hệ thống job.
Tools cho llms.txt
- llms.txt Generator + Validator: Draft curated file và kiểm tra của nó structure và linked các tài nguyên trước khi xuất bản.
- AI-Crawler Access Checker: Audit thực tế crawler
access controls riêng;
llms.txtkhông phải substitute cho điều này kiểm tra. curl: Verify status, các chuyển hướng, nội dung loại, và thân phản hồi tại chính xác root path.- máy chủ hoặc CDN nhật ký: Determine liệu named người dùng agents yêu cầu file; absence là observation, không proof của xếp hạng vấn đề.
- ** link checker:** Revalidate mỗi curated đích on schedule so hướng dẫn không decay.
Validate llms.txt phát hành
Kiểm thử root phản hồi
Kiểm thử để chạy: yêu cầu /llms.txt trực tiếp không có cookies và follow không các chuyển hướng.
Dự kiến kết quả: thành công phản hồi contains dự kiến Markdown tại root
path. thất bại interpretation: Routing hoặc deployment không expose file.
Monitoring window: Immediate. Rollback trigger: path phục vụ HTML shell,
lỗi document, hoặc unrelated chuyển hướng.
Kiểm thử mỗi listed tài nguyên
Kiểm thử để chạy: Crawl các URL trong file và so sánh mỗi đích với của nó mô tả. Dự kiến kết quả: Công khai canonical các tài nguyên resolve successfully và match stated purpose. thất bại interpretation: hướng dẫn là stale hoặc misleading. Monitoring window: Immediate, sau đó on tài liệu cập nhật cadence. Rollback trigger: listed URL exposes riêng tư material hoặc consistently fails.
Kiểm thử stated thử nghiệm outcome honestly
Kiểm thử để chạy: Query máy chủ nhật ký cho /llms.txt by người dùng agent sau khi publication.
Dự kiến kết quả: Các yêu cầu, nếu bất kỳ, là recorded và có thể gán; zero là cũng
hợp lệ kết quả. thất bại interpretation: Logging có thể là incomplete hoặc consumers có thể
không hỗ trợ proposal. Monitoring window: Several thông thường crawl cycles.
Rollback trigger: None cho zero traffic alone; xóa chỉ Khi maintenance hoặc
exposure risk exceeds file giá trị.
Tự kiểm tra: llms.txt
các tài nguyên worth của bạn time
** gốc nguồn**
- llmstxt.org — spec.
- Jeremy Howard proposal post (Sept 3, 2024) — rationale, straight từ creator.
- AnswerDotAI/llms-txt — repo và tooling.
** dữ liệu (đọc những điều này trước khi bạn implement bất cứ điều gì)**
- Điều gì là llms.txt, và nên bạn Care về nó? — Ahrefs — Ryan Law overview; 28%-publish / 97%-zero-các yêu cầu cách diễn đạt.
- llms.txt adoption nghiên cứu — Ahrefs — Louise Linehan & Xibeijia Guan, 137 210 domains: ai thực ra requesting files (Claude Code dominates AI; retrieval bots barely register) và prompt-injection finding.
- Làm llms.txt quan trọng? chúng ta tracked 10 các trang — công cụ tìm kiếm Land — 180-day nghiên cứu; 8/10 các trang showed không measurable thay đổi.
- llms.txt thử nghiệm — OtterlyAI —
90 days, một trang web: 84 của 62 100+ AI bot visits hit
/llms.txt.
Đó debate
- Không, llms.txt không phải đó “new meta keywords” (bản dịch) «new meta từ khóa» — Carolyn Shelby, Search Engine Land — đó strongest defense of llms.txt.
- Đáp ứng llms.txt, một proposed tiêu chuẩn — Search Engine Land — đó pro/con summary và Brett Tabke critique.
My related writing on điều này trang web
- AI các crawler — controlling bots (mà là Điều gì llms.txt là không cho), và nơi access-control story thực ra lives.
từ khoảng ngành
- SE Xếp hạng: State của llms.txt — 300 000-domain XGBoost/SHAP nghiên cứu finding đó removing llms.txt từ citation model improved độ chính xác; cũng có adoption-rate breakdown by traffic tier.
- Wix AI Tìm kiếm Lab — llms.txt myths (Crystal Carter, June 2026) — nhiều hơn optimistic đọc; 1 400+ files reviewed, argues llms.txt files themselves surface trong AI kết quả ngay cả nếu adoption là patchy.
- Google nói LLMs.Txt Comparable để Từ khóa Meta Tag — công cụ tìm kiếm Journal — đầy đủ John Mueller quote và context từ April 2025.
- presenc.ai: State của llms.txt 2026 — provider-by-provider hỗ trợ matrix; treat Anthropic/Perplexity “confirmed” claims với care (chính sources không luôn cited).
- llms.txt là dead. nhiều hơn precisely: dud. — Kai Spriestersbach (Medium) — aggregates OtterlyAI 90-day dữ liệu, Gary Illyes xác nhận, và Google 24-hour incident vào skeptic summary.
Số liệu worth citing
- 97% of published llms.txt files đã nhận zero các yêu cầu trong một month (Ahrefs nghiên cứu, 137 210 domains, Có thể 2026). Chỉ ~3% saw bất kỳ traffic tại all. Nguồn
- Claude Code là đó dominant AI consumer of llms.txt — điều này “outfetched every AI retrieval bot, assistant, and training crawler except GPTBot” (bản dịch) «outfetched mỗi AI retrieval bot, assistant, và training crawler except GPTBot» (đó training bot). Đó retrieval bots đó xây dựng citation indexes đã là chỉ 1,1% of các yêu cầu. Nguồn
- 77% of bot các yêu cầu để llms.txt files nghĩ ra từ non-AI tools — SEO auditors, anonymous các crawler, tech-profiling bots — không AI tại all. Nguồn
- 84 of 62 100+ AI bot visits hit
/llms.txtover 90 days on một monitored site (~0,1%), performing “3x worse than average pages.” (bản dịch) «3x tệ hơn average các trang.» Nguồn - Removing llms.txt từ một citation model improved của nó độ chính xác — SE Xếp hạng, 300 000 domains, XGBoost + SHAP analysis. Không demonstrated citation benefit “at least not yet.” (bản dịch) «ít nhất chưa.» Nguồn
- Adoption ≠ usage. BuiltWith được tính 844 000+ các trang với llms.txt (Oct 2025), nhưng nhiều of đó là nền tảng auto-deployment (Mintlify rolled điều này out để all hosted tài liệu các trang tại khi), không informed riêng lẻ decisions.
- Chỉ 408 các yêu cầu targeted
/llms.txttrên 515 million+ AI bot traffic events analyzed over 90 days (limy.ai analysis) — described as “statistically negligible.” (bản dịch) «không đáng kể về mặt thống kê.» - 10,13% of các trang có adopted llms.txt trên một 300 000-domain sample (SE Xếp hạng nghiên cứu, 2026) — khoảng 1 trong 10, với không có ý nghĩa khác biệt by traffic tier (thấp-traffic 9,88%, cao-traffic 8,27%). Nguồn
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.