Googlebot
Google · 315 published ranges
FCrDNS supported (.googlebot.com, .google.com).
One practical reference for the bots in your logs: search crawlers, AI training crawlers, AI-search fetchers, their published ranges, and the difference between a user-agent claim and a verified request.
These bots discover, render, and index pages for search surfaces. Check a suspicious request by IP before treating it as a search-engine crawl.
Google · 315 published ranges
FCrDNS supported (.googlebot.com, .google.com).
Google · 270 published ranges
FCrDNS supported (.google.com).
Microsoft · 28 published ranges
FCrDNS supported (.search.msn.com).
Yandex · 0 published ranges
FCrDNS supported (.yandex.ru, .yandex.net, .yandex.com).
Baidu · 0 published ranges
No FCrDNS scheme listed; use the published range list when available.
Naver · 0 published ranges
No FCrDNS scheme listed; use the published range list when available.
Apple · 12 published ranges
FCrDNS supported (.applebot.apple.com).
DuckDuckGo · 481 published ranges
No FCrDNS scheme listed; use the published range list when available.
AI systems use separate crawler tokens for model training, search/citations, and user-requested page fetches. Blocking one does not automatically affect the others.
Showing all 19 AI crawlers.
| Bot | Purpose | robots.txt posture | Verification | What it does |
|---|---|---|---|---|
| GPTBot OpenAI | Training | Says it respects robots.txt | 21 ranges | OpenAI's training crawler. Blocking it opts your content out of future model training data. Docs ↗ |
| OAI-SearchBot OpenAI | Search / citations | Says it respects robots.txt | 35 ranges | Indexes pages for ChatGPT search results and citations. Blocking it can remove you from ChatGPT search answers. Docs ↗ |
| ChatGPT-User OpenAI | User-triggered assistant | Unknown | 286 ranges | Fetches a page live for a user-triggered request. OpenAI documents that robots.txt rules may not apply to user-initiated actions, so treat robots policy as advisory for this agent. Docs ↗ |
| ClaudeBot Anthropic | Training | Says it respects robots.txt | 20 ranges | Anthropic's training crawler for Claude. Anthropic publishes an official crawler IP list at claude.com/crawling/bots.json; match the current list when verifying traffic because a ClaudeBot UA alone is spoofable. Docs ↗ |
| Claude-User Anthropic | User-triggered assistant | Says it respects robots.txt | 20 ranges | Fetches a page live when a Claude user's request needs it (e.g. a shared link). User-triggered. Anthropic publishes an official crawler IP list; a UA string alone is not identity proof. Docs ↗ |
| Claude-SearchBot Anthropic | Search / citations | Says it respects robots.txt | 20 ranges | Indexes pages to answer and cite in Claude's web search. Anthropic publishes an official crawler IP list; verify against the current list rather than trusting the UA alone. Docs ↗ |
| anthropic-ai Anthropic | Training | Says it respects robots.txt | 20 ranges | Older Anthropic crawler token still seen in robots.txt files; ClaudeBot is the current one. Keep a rule for both. Docs ↗ |
| PerplexityBot Perplexity | Search / citations | Says it respects robots.txt | 8 ranges | Indexes pages for Perplexity answers and citations. Docs ↗ |
| Perplexity-User Perplexity | User-triggered assistant | Reported exceptions — verify before relying on it | 4 ranges | User-triggered fetch. Perplexity documents that Perplexity-User requests are user-initiated and may not follow robots.txt; independent testing has reported it fetching disallowed pages. Docs ↗ |
| Google-Extended | Training | Says it respects robots.txt | 315 ranges + FCrDNS | A robots.txt token, not a separate crawler. Controls whether your content trains Gemini and grounds Gemini/Vertex answers. Blocking it does NOT remove you from Google Search or AI Overviews. Docs ↗ |
| Googlebot (AI Overviews) | Search / citations | Says it respects robots.txt | 315 ranges + FCrDNS | Google's AI Overviews use regular Search systems; there is no separate crawler. You cannot fully opt out only from AI Overviews while staying in Search, but nosnippet, max-snippet, and data-nosnippet can limit preview/use and also affect ordinary Search snippets. Docs ↗ |
| Applebot-Extended Apple | Training | Says it respects robots.txt | 12 ranges + FCrDNS | A robots.txt token controlling whether Applebot-collected pages train Apple's generative models. Blocking it does not affect Siri/Spotlight indexing (regular Applebot). Docs ↗ |
| CCBot Common Crawl | Training | Says it respects robots.txt | No published method | Common Crawl's crawler. Its public dataset is a major training source for many LLMs, so blocking CCBot indirectly reduces your presence in many models' training data. Docs ↗ |
| Bytespider ByteDance | Training | Reported exceptions — verify before relying on it | No published method | ByteDance's training crawler. Widely reported to ignore robots.txt and crawl aggressively; publishes no IP ranges, so it cannot be verified. Block at the WAF/CDN if you mean it. |
| Meta-ExternalAgent Meta | Training | Says it respects robots.txt | No published method | Meta's crawler for training Llama and Meta AI. Meta publishes no IP list; genuine traffic comes from AS32934. Docs ↗ |
| cohere-ai Cohere | Training | Says it respects robots.txt | No published method | Cohere's crawler token. |
| Amazonbot Amazon | Training | Says it respects robots.txt | FCrDNS only | Amazon's crawler; feeds Alexa answers and, per Amazon, model training. Amazon says it may use a robots.txt copy cached within the previous 30 days and may behave as if the file is absent when it cannot fetch it; this checker cannot prove Amazon's cached state. Supports reverse-DNS verification (.crawl.amazonbot.amazon). Docs ↗ |
| Diffbot Diffbot | Training | Says it respects robots.txt | No published method | Structured-data / knowledge-graph crawler whose data is licensed for AI training and enrichment. Docs ↗ |
| Omgilibot Webz.io | Training | Says it respects robots.txt | No published method | Webz.io crawler; its web dataset is sold for LLM training. |
AI training · OpenAI
21 published IP ranges.
Verify an IP →AI search · OpenAI
35 published IP ranges.
Verify an IP →AI training · Anthropic
20 published IP ranges.
Verify an IP →AI search · Perplexity
8 published IP ranges.
Verify an IP →AI training · Amazon
No IP-range data is available.
Verify an IP →AI training · Meta
Meta does not publish a machine-readable IP list for its crawlers. Genuine Meta crawler traffic originates from AS32934 (Facebook, Inc.); treat the ASN as the only available signal.
Verify an IP →AI training · ByteDance
ByteDance publishes no IP ranges and no reverse-DNS scheme for Bytespider. This bot cannot be verified — treat any Bytespider user agent with suspicion.
Verify an IP →A robots.txt file applies to one origin. That means https://www.example.com/robots.txt does not set rules for https://shop.example.com/; each subdomain needs its own file. The standard protocol defines User-agent, Allow, and Disallow, while crawler extensions need primary-source confirmation before you rely on them.
| System | Documented robots behavior | Use instead / caveat | Primary source |
|---|---|---|---|
Uses its documented REP interpretation; crawl-delay is not supported. | Use Search Console crawl controls and fix capacity issues; do not expect a Crawl-delay line to throttle Googlebot. | Google robots.txt specification ↗ | |
| Bingbot | Documents Crawl-delay values from 1–20 seconds. | Use Bing Webmaster Tools if you need help with an unexpected crawl pattern. | Bingbot guidance ↗ |
| Yandex | Documents a per-domain/subdomain crawl-rate setting in Yandex Webmaster. | Use the Webmaster setting rather than assuming a non-standard extension carries over from another crawler. | Yandex crawl-rate guidance ↗ |
Use the robots.txt Tester for Google-style Allow/Disallow matching across URLs and the Generator for an explicitly opt-in Bingbot crawl-delay line. Neither tool claims an undocumented interpretation for another system.
Paste an IP address to find the published CIDR range that contains it, or filter by crawler name, operator, or CIDR text. A match is evidence about the IP at this point in time, not a guarantee about every request it sends.
Enter a search to inspect the 2,986 published ranges.
| CIDR range | Crawler | Operator | Verification |
|---|
| Crawler | Operator | Published ranges | Reverse DNS | Data state |
|---|---|---|---|---|
| Googlebot | 315 | Yes (.googlebot.com, .google.com) | Current snapshot | |
| Google special crawlers | 270 | Yes (.google.com) | Current snapshot | |
| Google user-triggered fetchers | 1,506 | Yes (.gae.googleusercontent.com, .google.com, .googleusercontent.com) | Current snapshot | |
| Bingbot | Microsoft | 28 | Yes (.search.msn.com) | Current snapshot |
| YandexBot | Yandex | No published ranges | Yes (.yandex.ru, .yandex.net, .yandex.com) | Not published |
| Baiduspider | Baidu | No published ranges | No | Not published |
| Yeti | Naver | No published ranges | No | Not published |
| GPTBot | OpenAI | 21 | No | Current snapshot |
| OAI-SearchBot | OpenAI | 35 | No | Current snapshot |
| ChatGPT-User | OpenAI | 286 | No | Current snapshot |
| ClaudeBot | Anthropic | 20 | No | Current snapshot |
| PerplexityBot | Perplexity | 8 | No | Current snapshot |
| Perplexity-User | Perplexity | 4 | No | Current snapshot |
| Applebot | Apple | 12 | Yes (.applebot.apple.com) | Current snapshot |
| Amazonbot | Amazon | No published ranges | Yes (.crawl.amazonbot.amazon) | Not published |
| DuckDuckBot | DuckDuckGo | 481 | No | Current snapshot |
| Meta-ExternalAgent | Meta | No published ranges | No | Not published |
| Bytespider | ByteDance | No published ranges | No | Not published |
Start with the source IP. Copy it from an access log. Do not use a user-agent string as proof: any client can claim to be Googlebot, GPTBot, or another crawler.
Match the operator’s published CIDRs. Use the IP-ranges tab above or the Googlebot Verifier. A range match is strong evidence where an operator publishes and maintains one.
Run forward-confirmed reverse DNS when it is documented. The PTR hostname must end at the operator’s official domain boundary, and a forward lookup of that hostname must return the same IP.
Keep “unverifiable” separate from “fake.” Some operators publish neither ranges nor FCrDNS. That means identity cannot be confirmed—not that every request with that user agent is necessarily spoofed.
Verify a request
Check a single source IP against the snapshot and live DNS.
Understand training, search, and user-triggered fetchers before changing access rules.
These examples match a user-agent header, which can be spoofed. Prefer robots.txt for compliant bots and WAF/rate-limit controls for abusive traffic. Blocking a search or retrieval bot can remove your pages from that product’s results or citations.
# Review the effect before publishing — this blocks these user agents site-wide.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: anthropic-ai
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: cohere-ai
Disallow: /
User-agent: Amazonbot
Disallow: /
User-agent: Diffbot
Disallow: /
User-agent: Omgilibot
Disallow: / # Review before deploying. Matches the User-Agent header, not a verified IP.
if ($http_user_agent ~* "GPTBot|ClaudeBot|anthropic-ai|Google-Extended|Applebot-Extended|CCBot|Bytespider|Meta-ExternalAgent|cohere-ai|Amazonbot|Diffbot|Omgilibot") {
return 403;
} # Review before deploying. Matches the User-Agent header, not a verified IP.
SetEnvIfNoCase User-Agent "GPTBot|ClaudeBot|anthropic-ai|Google-Extended|Applebot-Extended|CCBot|Bytespider|Meta-ExternalAgent|cohere-ai|Amazonbot|Diffbot|Omgilibot" blocked_ai_crawler
<RequireAll>
Require all granted
Require not env blocked_ai_crawler
</RequireAll> Start with the source IP and verify it against published ranges and documented DNS checks.
Open Googlebot Verifier →Read the trade-offs between model training, persistent search/citation bots, and one-off user-triggered fetches.
Read AI crawler guidance →Use log-file analysis to distinguish a verified crawler from wasted crawl activity and unexpected bot categories.
Read log-file analysis →No. User-agent text is easy to spoof. Start with the source IP, then use the crawler operator’s published range list and documented forward-confirmed reverse-DNS process where available.
Not necessarily. Operators often use separate tokens for training, search/citations, and a user-requested fetch. Review each token before changing a rule; Google’s regular Googlebot remains the important exception for Search and AI Overviews.
Some operators do not publish ranges or a reverse-DNS validation method. The honest verdict is unverifiable: the request cannot be confirmed from public evidence, but that is not proof of spoofing.
Published-list history
Weekly snapshots keep a short, citable change history. A change means an operator changed its own published list; it does not by itself mean a crawler changed behavior.