Explicit AI crawler blocking was rare across 47 readable robots.txt files

A named-agent robots.txt study covering 11 AI-related crawler tokens across a fixed cohort of 49 SEO software sites.

First published: Sep 13, 2026 · Beginner
2 evidence signals on this page
  • Original research or data studyThis article is explicitly classified as a study.
  • Linked source dataDownload the dated JSON snapshot

Across 47 successfully retrieved robots.txt files, only ByteSpider had an explicit Disallow: / observation, on one site. Most tracked agent tokens were not named at all. Named-agent absence is not an allow verdict: wildcard groups, path-specific rules, longest-match precedence, authentication, and crawler behavior were outside this classification. Two origins were unknown because their robots.txt endpoints returned 403 and 404.

I classified explicit named-agent rules for 11 AI-related crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. tokens across a fixed cohort of 49 SEO software sites on September 13, 2026.

The finding

Forty-seven robots files were readable. Only one tracked token, ByteSpider, had an explicit root disallow in the cohort, and it appeared on one site. No explicit Disallow: / was observed for GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Google-Extended, PerplexityBot, Meta-ExternalAgent, or CCBot.

Data study

Explicit named-agent root disallows

One of 47 readable robots.txt files explicitly root-blocked ByteSpider. No root disallow was observed for the other tracked tokens.

View chart data
MeasureReadable files
ByteSpider1/47
GPTBot0/47
ClaudeBot0/47
Google-Extended0/47
PerplexityBot0/47
CCBot0/47

Two of 49 origins were unknown because robots.txt returned a non-200 response. Unknown observations are excluded from the 47-file denominator.

Methodology: The collector classified matching named-agent groups as root disallow, root allow, other named rule, or not mentioned. It stored response hashes and rule-count evidence without retaining discovered page URLs. Robots Exclusion Protocol (RFC 9309).

The important caveat

This is named-agent directive presence, not effective crawler access. A token that is not named may still encounter wildcard rules, path-specific rules, authentication, WAF controls, or behavior outside robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere.. The result should not be read as “46 sites allow ByteSpider” or “the cohort allows AI crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor..”

Download the dated JSON snapshot.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin an expert quote first.

Languages