Permission
Is the crawler allowed?
Free, no signup. “Am I letting ChatGPT, Claude, Perplexity, and Google’s AI read this page — and did I mean to?” Enter a URL and get the answer for every major AI crawler, with the exact robots.txt rule that decided it.
Is the crawler allowed?
Are useful facts present in raw HTML?
Does JavaScript successfully expose them?
Can an agent identify and use semantic controls?
Do page, schema, feed and checkout facts agree?
HTML remains primary. llms.txt, UCP, and other protocol files are optional distribution layers and cannot compensate for inaccessible or unextractable HTML.
Expected columns: crawler, timestamp, url, status, verification. A user-agent string alone is never treated as verified identity.
No log evidence uploaded.
How it works: matching runs in your browser with a matcher ported from Google’s open-source
robots.txt parser. The server only fetches the site’s robots.txt and checks for an llms.txt (browsers can’t — CORS).
Checks run from our server; we fetch the URL you enter and don't keep the results. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
Say a site's robots.txt contains:
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /private/ …and you check https://example.com/blog/post. The tool returns a per-crawler verdict like this
(run through the same matcher this tool ships with, not hand-typed):
Example data — captured by running this page's own matcher against the robots.txt above
| Crawler | Operator | Access | Winning rule |
|---|---|---|---|
OAI-SearchBot verifiableoai-searchbot | OpenAI | Allowed | no matching rule → default allow |
ClaudeBot published rangesclaudebot | Anthropic | Allowed | no matching rule → default allow |
GPTBot verifiablegptbot | OpenAI | Blocked | Disallow: / (line 2) |
User-agent: GPTBot group carries
Disallow: /, so this site opts out of OpenAI's model training. See
what Allowed/Blocked mean ↓.User-agent: * group only
disallows /private/, and /blog/post doesn't match it, so both fall through
to the default allow rather than matching any rule at all.https://example.com/some/page. The scheme is optional; a bare host is assumed
https://. The exact path matters, because a Disallow rule can block one
folder while the rest of the site stays open.robots.txt,
checks for an llms.txt, and evaluates every crawler in the roster against your URL.+ saves the current site or page. Use ☆ beside any saved site, page, or list to favorite it. Recent check history appears below.
Target filled from your local choices.
Saved targets, named lists, and recent check summaries remain only in this browser.
nosnippet, max-snippet, and data-nosnippet can limit preview/use but also affect normal Search snippets. Google-Extended only controls Gemini/Vertex training and grounding, not Search or Overviews.
nosnippet can block content from generative-model context without preventing title-based discovery, while Bing documents AI-answer/training effects for noarchive and noindex. Use the page audit or extension to inspect those meta and HTTP-header directives.
These fetch pages to answer questions and cite sources. Blocking a cooperating crawler can prevent that crawler from accessing the page, but it does not guarantee absence from every answer or citation.
These collect pages for model training. A provider's documented robots control governs future crawling; it does not remove content already collected or guarantee deletion from an existing training dataset.
Pick what to block; copy the generated rules into your robots.txt.
Blocking here means “documented posture,” not enforcement — a cooperating crawler will honor it; a hostile one may not.
Crawler roster last reviewed . Data is hand-maintained and editable (src/data/ai-crawlers.json).
Disallow rule in the crawler's group (or the
wildcard * group) covers this path. The Winning rule column shows the exact
directive and line number.robots.txt touches this crawler for this path, so it defaults to allowed.robots.txt,
so an "Allowed"/"Blocked" verdict here describes documented posture, not enforcement.The crawlers are split into Live retrieval & answer citation (bots that fetch pages now to answer and cite) and Model training (bots that collect pages to train models) so you can set each posture deliberately.
Matching runs in your browser with a matcher ported from Google's open-source
robots.txt parser — the same engine behind the site's robots.txt tester and Render Gap
tool — so verdicts follow Google's real precedence rules (most-specific path wins, not first-match).
The only server work is a small endpoint that fetches the target's robots.txt and checks
for an llms.txt, because a browser can't request another domain's files directly (CORS).
If robots.txt is missing or returns a 4xx, every crawler defaults to allowed
(Google's behavior). If it returns a 5xx or is unreachable, Google treats the whole site as temporarily
disallowed. The tool shows amber Not evaluated rows and withholds access verdicts until
the file can be fetched successfully again.
robots.txt, and tags each as verifiable or unverifiable by IP (cross-referenced with the Googlebot Verifier roster).llms.txt and reports it honestly (no crawler is documented to read it, so it changes no verdict).robots.txt block. It does not add a Googlebot disallow because that would prevent Google from crawling affected URLs. robots.txt is advisory, not a lock — it tells cooperating crawlers what
you prefer and enforces nothing. A crawler tagged "⚠ ignores robots" may fetch a "Blocked" page anyway;
to actually stop it you must block at your CDN/WAF and verify by IP. The roster is hand-maintained, so a
brand-new crawler may not appear until it's added. And a "verifiable" tag means the operator publishes a
way to check identity — it does not guarantee any given request is genuine.
Not on its own. GPTBot is OpenAI’s training crawler — blocking it opts your content out of future model training, but it does not stop the live-retrieval bots that fetch pages to answer questions and cite sources. Those are OAI-SearchBot (ChatGPT search results and citations) and ChatGPT-User (a page fetched when a user follows a link). To stay out of ChatGPT answers you would have to block those too, which the checker lists separately under "Live retrieval & answer citation."
There is no separate AI Overviews crawler, and Google does not offer a full AI-Overviews-only opt-out that keeps the page otherwise unchanged in Search. Blocking Googlebot prevents crawling but does not by itself guarantee that a known URL disappears from Search. Snippet controls also affect ordinary Search snippets. Google-Extended separately controls Gemini/Vertex training and grounding, not Search or AI Overviews.
robots.txt is advisory, not a lock. It tells cooperating crawlers what you prefer; it does not enforce anything. Well-behaved operators honor it, but some crawlers are reported to ignore it — the checker flags those with a "⚠ ignores robots" tag. To actually stop a crawler you have to block it at your CDN or WAF, and verify the request is genuine by IP rather than trusting the user-agent string.
The "verifiable" / "unverifiable" tag reflects whether the operator publishes a way to check a crawler claim — current IP ranges or reverse-DNS plus forward confirmation (FCrDNS), cross-referenced from the Bot Verifier roster. A user-agent alone is always spoofable. Anthropic now publishes an official crawler IP list at claude.com/crawling/bots.json, so Claude traffic can be checked for a published-range match; that is not the same as reverse-DNS confirmation.
No. llms.txt is a proposed community convention, and no major AI crawler is documented to read it, so its presence or absence does not affect whether any crawler is allowed or blocked. The checker reports whether an llms.txt exists at the domain purely for information — every access verdict is decided by robots.txt alone.
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
New requests are reviewed before they appear here.
Grátis e sem cadastro. “Estou permitindo que o ChatGPT, Claude, Perplexity e a IA do Google leiam esta página — e essa era a minha intenção?” Informe uma URL e obtenha a resposta para cada grande rastreador de IA, com a regra exata do robots.txt que determinou o resultado.
Veja se o ChatGPT, Claude, Perplexity e a IA do Google podem ler uma página — e se você pretendia permitir isso — com a regra exata do robots.txt que decidiu o resultado.
Como funciona: a correspondência é executada no navegador com um comparador adaptado do analisador de robots.txt de código aberto do Google. O servidor apenas busca o robots.txt do site e verifica se existe um llms.txt, pois o navegador não pode solicitar diretamente arquivos de outro domínio por causa do CORS.
Não por si só. O GPTBot é o rastreador de treinamento da OpenAI; bloqueá-lo impede que o conteúdo seja coletado para futuros treinamentos, mas não bloqueia os bots de recuperação ativa que buscam páginas para responder perguntas e citar fontes. Esses bots são o OAI-SearchBot, usado nos resultados e citações da busca do ChatGPT, e o ChatGPT-User, usado quando um usuário abre um link. Para não aparecer nas respostas do ChatGPT, você também teria de bloquear esses bots, listados separadamente em “Recuperação ativa e citação de respostas”.
Não existe um rastreador separado para as Visões Gerais Criadas por IA, e o Google não oferece uma desativação completa apenas para esse recurso que mantenha a página inalterada na Busca. Bloquear o Googlebot impede o rastreamento, mas não garante por si só que uma URL conhecida desapareça da Busca. Os controles de snippets também afetam os snippets comuns da Busca. O Google-Extended controla separadamente o treinamento e a fundamentação do Gemini/Vertex, não a Busca nem as Visões Gerais Criadas por IA.
robots.txt é uma orientação, não um bloqueio. Ele informa aos rastreadores cooperativos o que você prefere, mas não impõe nada. Alguns rastreadores abaixo são relatados como ignorando o arquivo, e a ferramenta os sinaliza com ⚠. Para realmente impedir um rastreador, bloqueie-o na CDN ou no WAF e verifique pelo IP se a solicitação é genuína.
A tag “verificável” ou “não verificável” indica se o operador publica uma forma de confirmar a identidade do rastreador: intervalos de IP atuais ou DNS reverso com confirmação direta (FCrDNS), com referência cruzada ao cadastro do Verificador de bots. Um user-agent isolado sempre pode ser falsificado. A Anthropic publica uma lista oficial de IPs de rastreadores em claude.com/crawling/bots.json; assim, o tráfego do Claude pode ser comparado aos intervalos publicados, o que não equivale à confirmação por DNS reverso.
Não. llms.txt é uma convenção proposta pela comunidade, e nenhum grande rastreador de IA está documentado como leitor desse arquivo. Sua presença ou ausência não muda se um rastreador é permitido ou bloqueado. A ferramenta relata a existência de llms.txt no domínio apenas como informação; todos os vereditos de acesso são determinados exclusivamente pelo robots.txt.