llms.txt Generator + Validator

Free, no signup. Build a clean llms.txt that points LLMs at your key pages, or paste one you already have and check that it's well-formed.

llms.txt structural rubric
CheckStatus when healthy
H1 One site-name H1
Summary A short blockquote is recommended
Links Grouped, absolute HTTPS URLs
Size Keep the map concise (under ~20 KiB)

Source: llmstxt.org proposal; not a crawler-consumption guarantee

Heads up: llms.txt is a community proposal (llmstxt.org), not an official standard. No major AI crawler is documented to read it today. Treat it as low-effort, low-risk housekeeping — not a ranking lever. If you want to influence what AI systems can access, that's robots.txt and your on-page content, not this file. More on how AI crawlers actually work →
Want evidence instead of assumptions?

Upload your own access log to the Log File Analyzer. Its browser-local /llms.txt report distinguishes verified AI crawler requests, unverifiable crawler claims, spoofed claims, other traffic, and no observation in the supplied period. A request is retrieval evidence only—not proof of rankings, citations, or broad provider adoption.

The Optional section (a crawler may skip it) is emitted last automatically — just name a section “Optional”.

Output — llms.txt

Runs entirely in your browser — nothing you paste is uploaded or stored. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.

Where this is heading: serving Markdown to LLMs

Looking ahead — not something to do today. llms.txt and per-page Markdown are proposals no major AI crawler is documented to consume yet.

llms.txt is a directory — one file pointing at your key pages. The more interesting idea, once AI clients adopt it, is content negotiation: same URL, two representations. Serve HTML to browsers and a clean Markdown rendering to LLM clients, and decide which at the edge.

A Cloudflare Worker picks the representation from the request — either Accept: text/markdown or a known AI user-agent — and emits Vary: Accept (or Vary: User-Agent) so shared caches don't hand the wrong variant to the wrong client.

It's exactly the move behind this site's hybrid /robots.txt — same bytes, only the Content-Type label negotiated — just extended to actual content, and to a real body difference rather than only a header.

This tool does two jobs: it generates a well-formed llms.txt from a simple form (or a sitemap you upload), and it validates one you paste or fetch from a live domain. Both run entirely in your browser. For context on why this file is optional housekeeping rather than a ranking move, see how AI crawlers actually work and the llms.txt explainer.

What the results mean

The validator groups findings into three severities, shown as pills:

  • The file breaks the proposal's shape — no # title, a malformed list item that isn't a [title](url) link, a relative (non-absolute) URL, or an empty file.
  • Well-formed but not ideal — no > summary, an http:// link, a duplicate URL, a second # H1, or links that appear before any ## section.
  • Advisory only — a summary placed somewhere other than right after the title, an off-site link, a deep ### heading that isn't part of the structure, or a file large enough (past 20 KiB) to have stopped being a concise map.

A green result means the file is well-formed against the proposal. The result summary is deliberately blunt that this validates form only and cannot guarantee any AI system consumes the file.

How it works

The generate/parse/validate engine (src/lib/tools/llms-txt.ts) is pure TypeScript with no DOM and no network calls, so everything you do in the Generate and paste flows happens locally in your browser — your draft never leaves the page.

The parser reads line by line: it recognises the # title, a > blockquote summary, ## section headings, and Markdown link list items, while skipping the insides of fenced ``` code blocks so they aren't misread as directives. It then runs file-level checks — required title, recommended summary, duplicate URLs, absolute-URL and HTTPS rules, and size — and rolls everything into the error/warning/info tally.

The one server touch is Fetch by URL: browsers can't fetch another site's /llms.txt because of CORS, so a small SSRF-guarded, cached proxy (/api/llms-fetch) pulls only that exact path — nothing else on the domain — and hands the body back for validation.

Features

  • Form-based generator with add/remove sections and link rows, live output, and self-validation as you type.
  • Prefill from sitemap.xml — upload a sitemap and it buckets URLs by first path segment into ready-to-edit sections.
  • Automatic handling of the Optional section (always emitted last).
  • Copy, download, and a share link that packs your draft into the URL fragment (compressed, never sent to a server).
  • Validator with paste or live Fetch /llms.txt by domain, an error/warning/info tally, per-line findings, and a parsed-structure preview.
  • Runs client-side; the only server call is the CORS-bypassing llms.txt fetch.

Limitations

It validates the form of the file against the community proposal — it cannot tell you whether any AI crawler reads it, because none is documented to. It checks link syntax (absolute, HTTPS, well-formed Markdown) but does not crawl the links to confirm they resolve or return the content you claim. The sitemap prefill is a starting point, not a curation step — an llms.txt is only useful if you trim it to your genuinely important pages. And this is not a lever for AI visibility: what AI systems can access is governed by your robots.txt and your on-page content.

Frequently asked questions

Do AI crawlers actually read llms.txt?

Not in any documented way. llms.txt is a community proposal from llmstxt.org, not an official standard, and no major AI crawler (OpenAI, Google, Anthropic, Perplexity) has published support for reading it. Treat it as low-effort, low-risk housekeeping — a tidy map of your best pages — not a ranking lever. What AI systems can actually reach is governed by your robots.txt and your on-page content.

What goes in an llms.txt file?

One required "# " H1 title (your site name), a recommended "> " blockquote one-line summary right after it, then "## " section headings (Docs, Guides, Blog) each containing Markdown link list items in the form "- [title](url): optional notes". A special "## Optional" section holds links a crawler may skip. Every link URL should be an absolute https:// address the LLM can fetch directly.

Where does the llms.txt file go on my site?

At the root of your domain, served at https://site.example/llms.txt — the same location convention as robots.txt. This validator can fetch that exact path for any domain (via a small server-side proxy, because browsers cannot fetch cross-origin) so you can check a live file without pasting it.

What does the validator check?

The form of the file against the proposal, not whether anything reads it. It flags a missing H1 title (error), missing summary (warning), malformed or relative link URLs, http:// links, duplicate URLs, links that appear before any section, deep "### " headings that are not part of the structure, off-site links, and files large enough to have stopped being a concise map. It cannot and does not guarantee any AI system consumes the file.

How big should llms.txt be?

Small. It is a directory of links, not a copy of your content. The validator raises an informational note once a file passes about 20 KiB, because at that size it has usually stopped being a concise map and started duplicating pages. If you want to serve full clean content to LLMs, that is per-page Markdown and content negotiation, not a bigger llms.txt.

Next stepAI Search Volume Estimator — score it against a calibrated model.

Feature requests for Llms Txt

Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.

Loading…

➕ Request a feature

New requests are reviewed before they appear here.

ツールについて

主要ページへLLMを案内する整形式のllms.txtを生成するか、既存ファイルを貼り付けまたは取得して検証します。H1、概要、絶対リンク、重複、外部サイトの参照を確認します。

llms.txtはコミュニティ提案であり公式標準ではないことを明示します。AIクローラーが読むと保証せず、すべてブラウザー内で処理します。

機能

  • H1と概要の検証
  • 整形式の絶対リンク、重複、外部リンクの検出
  • llms.txtの生成と貼り付け/取得検査
  • 提案仕様の限界を明示するレポート

仕組み

テキストを行ごとに解析し、H1、概要、セクション、絶対URLを正規化します。重複、相対リンク、外部リンク、形式エラーを分けて表示し、生成時は入力ページの一覧から整形式の出力を組み立てます。

制限事項

  • llms.txtは公式標準ではなく、AIクローラーが読むこと、検索結果、引用、トラフィックを保証しません。
  • 取得結果は入力URLの応答に依存し、認証、ネットワーク、robots、キャッシュの状態を推測しません。

よくある質問

llms.txtを置けばAIクローラーが必ず読みますか?

いいえ。コミュニティ提案であり、公式標準でもクローラーの実装保証でもありません。

外部リンクを含められますか?

含められますが、外部サイトの参照として警告します。対象と運用意図を確認してから公開してください。

このツールはファイルをアップロードしますか?

いいえ。生成と検証はブラウザー内で行います。取得を選んだ場合は指定URLへのリクエストだけを実行します。