Guide : Googlebot
Ce que Googlebot en réalité is — Smartphone vs Desktop, evergreen Chromium rendering, user-agent strings, IP-range verification, byte limites, and the crawl-vs-rank distinction.
Langues
1 indice probant sur cette page
- Outil en ligne associérobots.txt Tester
Googlebot is Google's web robot d’exploration — the software que récupère pages so Google peut index and rank les. It comes in two variants que share un robots.txt token: Googlebot Smartphone (principal, sous indexation mobile-first) and Googlebot Desktop. It runs an evergreen Chromium and renders JavaScript in a separate, plus tard queue (explorer ≠ render). Exploration isn't a ranking factor, and blocking Googlebot in robots.txt isn't the même as deindexing — a blocked URL peut encore be indexé URL-only. The user-agent is trivially spoofed, so vérifier with reverse + forward DNS or Google's publié IP ranges. 'Googlebot' is really Recherche Google's slice of a beaucoup plus grand exploration platform.
Evidence for this claim Googlebot is Google's crawler, with smartphone and desktop crawler types that share the same product token. Scope: Current Googlebot crawler and user-agent documentation. Confidence: high · Verified: Google Search Central: Googlebot Evidence for this claim A claimed Google crawler can be verified using reverse and forward DNS or Google's published IP ranges. Scope: Google's current crawler-verification methods. Confidence: high · Verified: Google Search Central: Verify GooglebotTL;DR — Googlebot is Google’s web robot d’exploration — the program que visits votre pages, downloads les, and hands les off so Google peut index and rank les. Là are two of les (a smartphone un and a desktop un), and the smartphone un fait la plupart of the fonctionner. Being crawled is a requirement to montrer up in search — but getting crawled plus doesn’t faire vous rank plus élevé.
Ce que Googlebot is
Google doesn’t browse the web the façon vous do. It sends out an automated program — a robot d’exploration, or bot — que visits pages, downloads what’s on les, and follows liens to trouver plus pages. Que program is Googlebot. (Bing’s equivalent is Bingbot.)
I décrire it simply in my Ahrefs guide to Googlebot: “Googlebot is the web robot d’exploration utilisé by Google to gather the information nécessaire and construire a searchable index of the web.” Everything Google montre vous in résultats de recherche commencé with Googlebot fetching the page.
Là are en réalité two Googlebots
Googlebot comes in two flavors:
- Googlebot Smartphone — pretends to be someone on a phone. Ce is the principal un. Google mostly semble at the mobile version of votre site (that’s appelé indexation mobile-first), so la plupart crawls come from the smartphone bot.
- Googlebot Desktop — pretends to be someone on a desktop computer. It fait a plus petit share of the exploration.
Here’s the catch: in votre robots.txt fichier (the fichier que indique bots où ils
peut go), les deux share the même nom — Googlebot. So vous pouvez’t tell robots.txt
“let the desktop one in but not the mobile one.” It’s all-or-nothing.
Fait Googlebot run JavaScript?
Yes. Googlebot uses a recent version of Chrome sous the hood, so it peut lire modern, JavaScript-heavy sites. As I put it in my Googlebot guide: “Googlebot is evergreen, meaning it sees websites as utilisateurs voudrait in the latest Chrome navigateur.” But running que JavaScript se produit a little plus tard, in a separate step — pas the instant votre page is premier récupéré.
The two choses personnes obtenir incorrect
- Exploration n’est pas ranking. Getting crawled plus souvent won’t déplacer vous up the results. Exploration is simplement how Google trouve and downloads votre page — it’s a gate vous have to obtenir via, pas a scoreboard.
- Blocking Googlebot in
robots.txtne fait pas delete vous from Google. It seulement arrête Google from reading lune page. Si autre sites lien to it, l’URL peut encore montrer up in results (simplement sans a utile description). To en réalité garder a page out, vous let Google explorer it and ajouter anoindextag.
Vouloir the technical version — the exact user-agent strings, how to vérifier a réel Googlebot, byte limites, and the rendering queue? Switch to the Avancé tab.
Evidence for this claim Googlebot is Google's crawler, with smartphone and desktop crawler types that share the same product token. Scope: Current Googlebot crawler and user-agent documentation. Confidence: high · Verified: Google Search Central: Googlebot Evidence for this claim A claimed Google crawler can be verified using reverse and forward DNS or Google's published IP ranges. Scope: Google's current crawler-verification methods. Confidence: high · Verified: Google Search Central: Verify GooglebotTL;DR — Googlebot is Recherche Google’s robot d’exploration, split into Smartphone (principal, mobile-first) and Desktop, qui share un
Googlebotrobots.txt token — vous pouvez’t target les separately. It runs an evergreen Chromium and renders JavaScript in a separate, plus tard queue (explorer ≠ render). Exploration is requis to rank but is pas a ranking signal, and a robots-blocked URL peut encore be indexé URL-only. Vérifier it by reverse + forward DNS to a Google domain or contre Google’s publié IP ranges — the user-agent is trivially spoofed. And “Googlebot” is really simplement the Search-facing slice of a beaucoup bigger exploration platform.
Ce que Googlebot en réalité is
Google is precise à propos de the nom: “Googlebot is the generic nom pour two types of web robots d’exploration utilisé by Recherche Google.” Ceux two types are Googlebot Smartphone (“a mobile crawler that simulates a user on a mobile device”) and Googlebot Desktop (“a desktop crawler that simulates a user on desktop”).
It n’est pas un little program running on un machine. “Googlebot runs on thousands of machines,” as I describe it in my Googlebot guide, “ils determine how fast and ce que to explorer on websites,” distributed à travers datacenters worldwide but egressing primarily from US IP addresses. Discovery se produit mostly via liens — Google trouve nouveau URLs “primarily from links embedded in previously crawled pages” — plus sitemaps. (Pour the complet discovery and scheduling picture, that’s the exploration hub’s job.)
Smartphone vs Desktop — and pourquoi it’s “smartphone-first”
Sous indexation mobile-first, the smartphone robot d’exploration is the principal un. Google: “Pour la plupart sites Recherche Google primarily indexes the mobile version of le contenu. As tel the majority of Googlebot explorer requêtes va be made en utilisant the mobile robot d’exploration, and a minority en utilisant the desktop robot d’exploration.” Indexation mobile-first has been complet pour tout sites since October 2023, so the practical rule is: si content isn’t visible to the smartphone agent, it isn’t indexé. Match votre content, données structurées, metadata, and robots tags à travers mobile and desktop.
The robots.txt gotcha: “Les deux robot d’exploration types obey the même product token (utilisateur agent
token) in robots.txt, and so vous pouveznot selectively target soit Googlebot
Smartphone or Googlebot Desktop en utilisant robots.txt.” The seulement façon to differentiate is
to lire the HTTP user-agent requête header in votre propre server-side logic. (Pour the
mechanics of que header format, voir indexation mobile-first and user-agent.)
The user-agent strings
The robots.txt product token pour les deux is simplement Googlebot. The complet UA strings
differ:
Googlebot Desktop:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36Googlebot Smartphone:
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)W.X.Y.Z is a placeholder pour the current Chrome version, qui moves as Googlebot’s
evergreen Chromium is mis à jour. Don’t trust the string on its propre, though — it’s
trivially spoofed (voir verification ci-dessous).
Evergreen Chromium and the rendering queue
Ce is the distinction que trips personnes up la plupart: exploration and rendering are separate steps. Googlebot runs “an evergreen version of Chromium,” and it became evergreen back in May 2019 (jumping from the old Chrome 41 to current stable), qui is pourquoi it now handles ES6+, IntersectionObserver, Web Components, and modern CSS. But it doesn’t execute votre JavaScript the moment it récupère the HTML.
Google: “Googlebot queues tout pages with a 200 Code d’état HTTP pour rendering,
unless a robots meta tag or header indique Google pas to index lune page. Lune page may
stay on ce queue pour a few seconds, but it peut prendre plus long que que. Une fois
Google’s resources autoriser, a headless Chromium renders lune page and executes the
JavaScript.” The rendering service (WRS) behaves comme a modern navigateur but with
quirks worth knowing: it’s effectively stateless — local/session storage and
cookies are cleared à travers page loads — it doesn’t récupérer images or videos (to enregistrer
bandwidth), caches aggressively (and may ignore votre mise en cache headers), and doesn’t
prise en charge WebSockets or WebRTC. Si votre content seulement apparaît après a click or a
JS-driven navigation que isn’t a réel <a href> lien, expect rendering trouble.
(Depth lives in the rendering sibling.)
Byte limites
Googlebot doesn’t download an unlimited amount per URL. As of Google’s March 2026 À l’intérieur Googlebot mettre à jour, it récupère roughly the premier 2 MB of quelconque individual URL (notamment the HTTP header) and up to 64 MB pour a PDF. The figure is a moving target — Google’s propre framing is que “ce limite n’est pas définir in stone and may modifier over temps as the web evolves and HTML pages grow in size,” and précédent docs listed 15 MB (qui turned out to be the broader infrastructure par défaut, pas Search’s number). The practical point holds regardless: anything past the cutoff simply isn’t récupéré — “to Googlebot, they simply don’t exist.” Garder critical content and markup ci-dessus the bloat.
Worked échec: the canonical exists, but Googlebot jamais receives it
Imagine a product template returning 2,4 MB of HTML. An application serializes a huge product-state object and recommendation payload near the top of the document; the canonical, product description, données structurées, and related-product liens do pas apparaître jusqu’à roughly byte 2 180 000. A navigateur downloads the whole réponse, so View Source semble correct. Googlebot Search arrête autour its documented 2 MB limite, so ceux plus tard signals ne faites pas exist in the récupéré resource.
Diagnose la réponse in byte order, pas seulement in the rendered DOM:
curl -sS -D response-headers.txt -o page.html https://example.com/product
wc -c response-headers.txt page.html
LC_ALL=C grep -abo 'rel="canonical"' page.html
LC_ALL=C grep -abo 'application/ld+json' page.htmlThe local byte counts are an approximation parce que delivery intermediaries and réponse handling peut differ, but ils réponse the utile premier question: are critical signals comfortably early, or are ils sitting near or au-delà the boundary? The fix is to supprimer or defer oversized inline données and emit essential metadata, principal content, and crawlable liens early—pas to déplacer the même bloat autour and hope the cutoff changements.
How Googlebot récupère — politely
- Fréquence d’exploration is algorithmic and self-throttling. “Pour la plupart sites, Googlebot shouldn’t accès votre site plus que une fois every few seconds on average.” It speeds up or backs off fondé on votre serveur’s health.
- Code d’états are the lever. Returning
429,500, or503indique Googlebot to slow bas — but que affecte the entier hostname, pas simplement the erroring URLs, and seulement fonctionne pour a day or two avant sustained errors commencer dropping pages from the index. John Mueller: “I’d seulement expect the fréquence d’exploration to react que quickly si ils were returning 429 / 500 / 503 / timeouts,” and “404s are généralement fine & une fois découvert, Googlebot va retry les anyway.” crawl-delayis ignored. Google ne fait pas traiter the non-standardcrawl-delayrobots.txt directive at tout. (Bing fait honor it — un of the réel Googlebot/Bingbot divergences.)- Exploration tracks demande d’exploration, pas a flat quota — capacity (ce que votre serveur peut prendre) plus demand (popularity and staleness). Pour la plupart sites ce is a non-issue; it seulement bites at réel scale. The complet treatment is in budget d’exploration.
Verifying it’s really Googlebot
The user-agent header is “often spoofed by other crawlers” — so it alone proves nothing. Google’s robots d’exploration identifier themselves three façons: the user-agent header, the source IP, and the reverse-DNS hostname of que IP. Two réel verification méthodes:
- Manual (one-off). Reverse-DNS the source IP; confirmer it resolves to a
hostname ending in
googlebot.com,google.com, orgoogleusercontent.com(the mask semble commecrawl-***-***-***-***.googlebot.com); alors forward-DNS que hostname and confirmer it renvoie the original IP. - Automatic (at scale). Match the IP contre Google’s publié CIDR ranges.
Google has split ces from the old unique
googlebot.jsoninto several JSON fichiers by robot d’exploration category — the Googlebot un ishttps://www.gstatic.com/ipranges/common-crawlers.json(the legacygooglebot.jsonURL encore redirections to the même données).
Les deux are in the Scripts tab, pour macOS/Linux and Windows. Pourquoi bother? Plenty of trafic lies à propos de being Googlebot, so logs que “show Googlebot” peut be largely impostors — vérifier avant vous trust.
Googlebot is un bot in a fleet
“Googlebot” is genuinely a bit of a misnomer. Gary Illyes, March 2026: “I mean, appel it Googlebot, that’s a misnomer,” and “Googlebot n’est pas our robot d’exploration infrastructure.” The infrastructure underneath, in his words, is “software as a service, si vous comme. SaaS” — a shared platform nombreux Google products draw from. Ce que vous voir in votre logs is the Search slice of it: “Quand vous voir Googlebot in votre serveur logs, vous are simplement looking at Recherche Google.” He aussi notes là are “dozens, if not hundreds of different crawlers,” la plupart aussi petit to bother documenting.
The named ones you’ll en réalité meet alongside Googlebot inclure Googlebot-Image,
Googlebot-Video, and Googlebot-News (qui share Googlebot’s strings/tokens),
Storebot-Google, and Google-InspectionTool (powers the Inspection d’URL and Rich
Results outils). Two behave unusually: AdsBot ignores the global * robots.txt
rule (a Disallow: / sous User-agent: * encore won’t arrêter it), and
Google-Safety ignores robots.txt entirely. The AI-related robots d’exploration —
Google-Extended (contrôle Gemini training; pas a ranking signal) and
GoogleOther (R&D crawls, offloaded from Googlebot) — exist aussi, but the
AI robots d’exploration sibling covers ceux in depth, so I’ll point là plutôt que
duplicate.
How to contrôler Googlebot
Three contrôle, three différent effects:
robots.txtarrête exploration, pas indexation. Utiliser it to garder bots out of low-value URL spaces — jamais as a deindexing outil.noindexarrête indexation — but Googlebot doit be allowed to explorer lune page to voir the tag in the premier placer.- Password protection blocks accès entirely.
Qui brings us to the unique la plupart misunderstood Googlebot fact: “There’s a
difference entre exploration and indexation; blocking Googlebot from exploration une page
doesn’t prevent l’URL of lune page from appearing in résultats de recherche.” A
robots-blocked URL peut encore obtenir indexé URL-only si something liens to it. To
en réalité supprimer une page, autoriser exploration and ajouter noindex. I’ve written ce up
in detail in
Indexé, though blocked by robots.txt.
Pour the wider pipeline Googlebot lives à l’intérieur — URL discovery, the explorer scheduler, rendering, and the crawl-vs-index-vs-rank distinctions — voir the exploration hub and How Search Fonctionne.
AI summary
A condensed prendre on the Avancé version:
- Googlebot = Recherche Google’s web robot d’exploration, in two variants: Smartphone
(principal, sous indexation mobile-first) and Desktop. Ils share un
Googlebotrobots.txt token, so vous pouvez’t target les separately — seulement via the HTTP user-agent header. - Mobile-first since Oct 2023: si content isn’t visible to the smartphone agent, it isn’t indexé.
- Evergreen Chromium, separate rendering queue: Googlebot runs current Chromium but executes JavaScript plus tard, in a stateless render step (storage/cookies cleared; aucun images/video récupéré; aggressive mise en cache). Explorer ≠ render.
- Byte limite ~2 MB per URL (PDFs 64 MB) as of 2026 — content au-delà the cutoff isn’t récupéré. The exact figure is “not set in stone.”
- Polite, algorithmic fetching:
429/5xxdire “slow down” (hostname-wide, short-term);crawl-delayis ignored (Bing honors it). - Exploration ≠ ranking, and robots-block ≠ deindex — a blocked URL peut encore
be indexé URL-only. Supprimer pages with
noindex(exploration allowed), pas robots.txt. - Vérifier by reverse + forward DNS to
*.googlebot.com/*.google.com/*.googleusercontent.com, or contre Google’s publié IP ranges (common-crawlers.json) — the user-agent is trivially spoofed. - “Googlebot” is a fleet, pas un program — the Search-facing slice of a plus grand
exploration platform, alongside Googlebot-Image/Video/News, Google-InspectionTool,
AdsBot (ignores
*), Google-Safety (ignores robots.txt), and the AI robots d’exploration (Google-Extended, GoogleOther — voir the AI robots d’exploration sibling).
Documentation officielle
Primary-source documentation from the moteur de recherches.
- Googlebot — the canonical doc: the two robot d’exploration types, UA strings, byte limites, and explorer/index contrôle. Commencer ici.
- Overview of Google robots d’exploration and fetchers — every Google user-agent, the three robot d’exploration categories, and the publié IP ranges.
- Google’s courant robots d’exploration — the complet UA-string and robots-token référence pour Googlebot and its siblings.
- Vérifier Google robots d’exploration and fetchers — reverse/forward DNS and the JSON IP-range fichiers.
- Indexation mobile-first meilleur practices — pourquoi Smartphone is principal and ce que to garder consistent.
- JavaScript SEO basics — the evergreen Chromium renderer and the rendering queue.
- How Code d’état HTTPs affecter Google’s robots d’exploration — how
2xx/3xx/4xx/429/5xxmodifier explorer behavior. - À l’intérieur Googlebot (March 2026) — the current byte limites and “Googlebot is just Google Search” framing.
Bing / Microsoft (pour contrast)
- Bingbot explorer contrôler — Bing’s manual explorer scheduling and its
crawl-delayprise en charge, qui Googlebot doesn’t have.
Quotes from the source
On-the-record statements from Google. Chaque lien is a deep lien que jumps to the quoted passage on the source page.
Google — ce que Googlebot is
- “Googlebot is the generic name for two types of web crawlers used by Google Search.” — Recherche Google Central docs. Jump to quote
- “For most sites Google Search primarily indexes the mobile version of the content. As such the majority of Googlebot crawl requests will be made using the mobile crawler, and a minority using the desktop crawler.” Jump to quote
- “Google uses the mobile version of a site’s content, crawled with the smartphone agent, for indexing and ranking.” Jump to quote
Google — rendering and byte limites
- “While Google Search runs JavaScript with an evergreen version of Chromium, there are a few things that you can optimize.” Jump to quote
- “Googlebot queues all pages with a 200 HTTP status code for rendering, unless a robots meta tag or header tells Google not to index the page.” Jump to quote
- “Googlebot crawls the first 2MB of a supported file type, and the first 64MB of a PDF file.” Jump to quote
Google — explorer vs index
- “There’s a difference between crawling and indexing; blocking Googlebot from crawling a page doesn’t prevent the URL of the page from appearing in search results.” — Recherche Google Central docs. Jump to quote
Gary Illyes, Google (reported via Moteur de recherche Journal and Google’s March 2026 blog post)
- “I mean, calling it Googlebot, that’s a misnomer.” … “Googlebot is not our crawler infrastructure.” … “it’s software as a service, if you like. SaaS.” Lire the coverage
- “When you see Googlebot in your server logs, you are just looking at Google Search.” Lire the coverage
John Mueller, Google (via Moteur de recherche Journal’s Reddit coverage)
- “I’d only expect the crawl rate to react that quickly if they were returning 429 / 500 / 503 / timeouts.” … “404s are generally fine & once discovered, Googlebot will retry them anyway.” Lire the coverage
Googlebot readiness checklist
A rapide réussir to confirmer Googlebot peut trouver, récupérer, render, and correctement handle votre pages:
- Important content is reachable via réel
<a href>liens (pas click-only or JS-driven navigation the renderer can’t follow). - Mobile and desktop versions match — même content, données structurées, metadata, and robots tags (mobile-first signifie the smartphone agent’s view is ce que counts).
- Critical content and markup sit ci-dessus the ~2 MB récupérer cutoff; PDFs sous 64 MB.
-
robots.txtdoesn’t block anything vous vouloir indexé — and you’re pas en utilisant it to essayer to deindex (that’snoindex’s job, with exploration allowed). - Server renvoie fast, stable réponses — minimal
5xx/429/timeouts, since ceux throttle exploration hostname-wide. - You’re pas relying on
crawl-delay(Google ignores it). - Server logs are vérifié with verified Googlebot (reverse + forward DNS or IP ranges) — pas raw user-agent matches, qui impostors spoof.
- Vous comprendre a robots-blocked URL peut encore be indexé URL-only si lié to.
Googlebot — cheat sheet
The two robot d’exploration types
| Googlebot Smartphone | Googlebot Desktop | |
|---|---|---|
| Simulates | A mobile-device utilisateur | A desktop utilisateur |
| Share of crawls | Majority (mobile-first) | Minority |
| robots.txt token | Googlebot | Googlebot (même — can’t separate) |
| Tell apart via | HTTP user-agent header | HTTP user-agent header |
robots.txt product token (les deux): Googlebot
Autre Google robots d’exploration you’ll voir in logs
| Robot d’exploration | robots token | Remarque |
|---|---|---|
| Googlebot | Googlebot | Principal Search robot d’exploration |
| Googlebot-Image / -Video / -News | Googlebot-Image etc. (or Googlebot) | Utiliser Googlebot’s UA strings |
| Google-InspectionTool | Google-InspectionTool (or Googlebot) | Inspection d’URL / Résultats enrichis tests |
| Storebot-Google | Storebot-Google | Shopping |
| AdsBot | AdsBot-Google | Ignores the global * robots.txt rule |
| Google-Safety | — | Ignores robots.txt entirely |
| GoogleOther / Google-Extended | GoogleOther / Google-Extended | R&D / AI training — voir AI robots d’exploration |
Fast facts
- Récupérer limite: ~2 MB per URL (PDFs 64 MB), as of 2026 — au-delà it isn’t récupéré. (“Not set in stone.”)
- Renders with evergreen Chromium, in a separate, plus tard queue (explorer ≠ render).
crawl-delay: ignored by Google (Bing honors it).429/5xx: temporary “slow down” — affecte the whole hostname.- Vérifier with reverse + forward DNS or publié IP ranges — jamais the UA string alone.
- Exploration is pas a ranking factor; robots-block is pas deindexing.
Vérifier a bot is really Googlebot
The user-agent header is trivially spoofed, so confirmer with a reverse-DNS lookup (doit fin in a Google domain) followed by a forward-DNS lookup (doit resolve back to the même IP).
macOS / Linux
# 1) Reverse-DNS the IP from your logs — it should end in
# googlebot.com, google.com, or googleusercontent.com
host 66.249.66.1
# → 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com
# 2) Forward-DNS that hostname back — it must resolve to the same IP
host crawl-66-249-66-1.googlebot.com
# → crawl-66-249-66-1.googlebot.com has address 66.249.66.1Windows
nslookup 66.249.66.1
nslookup crawl-66-249-66-1.googlebot.comSi the reverse lookup doesn’t fin in a Google domain, or the forward lookup doesn’t match the original IP, it isn’t Googlebot.
Vérifier at scale contre Google’s IP ranges
Pour large-scale checks, match source IPs contre Google’s publié CIDR ranges
au lieu de doing per-request DNS. Google split the old unique googlebot.json into
several fichiers by category — Googlebot lives in common-crawlers.json:
# Fetch the current Googlebot (common crawlers) ranges
curl -s https://www.gstatic.com/ipranges/common-crawlers.json
# The legacy URL still works and redirects to the same data:
# https://developers.google.com/static/search/apis/ipranges/googlebot.jsonCharger the CIDR prefixes from que JSON and tester chaque logged IP pour membership avant trusting quelconque “Googlebot” hit.
Ressources utiles
My connexe writing
- Ce que Is Googlebot & How Fait It Fonctionner? — my complet Ahrefs guide on Googlebot, with the explorer/contrôler/vérifier detail.
- Meet the Nouveau Web Robots d’exploration: AI Bots Are Closing in on Moteur de recherche Bots — how Googlebot stacks up contre the AI robots d’exploration by share and speed.
- JavaScript SEO Problèmes & Meilleur Practices — the rendering side of ce que Googlebot fait.
- Indexé, though blocked by robots.txt — pourquoi a robots-blocked URL peut encore be indexé.
- The Beginner’s Guide to SEO technique — où Googlebot fits in the bigger picture.
My speaking
- How Search Fonctionne (SlideShare) — my walkthrough of exploration, rendering, indexation, and ranking. (Standing disclaimer s’applique: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From others
- Google’s Exploration December series — the meilleur concentrated définir of official explorer explainers.
- r/TechSEO — the community pour explorer/index debugging.
- Googlebot: Ce que c’est, Comment cela fonctionne & how to optimize (Moteur de recherche Land) — thorough practical guide with SEL’s editorial depth; bon pour a second opinion on the basics.
- Google explique how exploration fonctionne in 2026 (Barry Schwartz, Moteur de recherche Land) — solid write-up of Gary Illyes’s March 2026 À l’intérieur Googlebot post, notamment the 2 MB byte-limit clarification.
- Google Dit Ils Deploy Hundreds of Undocumented Robots d’exploration (Moteur de recherche Journal) — covers Gary Illyes’s “calling it Googlebot is a misnomer” remarks in depth.
- Is Google’s Two Waves of Indexation Over? (Onely) — Martin Splitt’s on-record conversation à propos de the crawl-then-render queue and how the gap is shrinking; encore the clearest industry treatment of WRS timing.
- Google Reveals JavaScript Rendering Secrets at Chrome Dev Summit 2018 (Lumar) — detailed notes on WRS stateless behavior (cookies/localStorage cleared, aucun image récupère, WebSockets unsupported) que are encore relevant today.
- From Googlebot to GPTBot: Who’s exploration votre site? (Cloudflare blog) — Cloudflare Radar données on robot d’exploration share and speed showing Googlebot leading “good bot” trafic.
- Was Que Really a Google Bot Exploration My Site? (Imperva Research) — data-backed regarder at Googlebot impersonation; context pour pourquoi DNS verification matters.
Stats worth citing
- Googlebot is the fastest robot d’exploration on the web. From my lire of Cloudflare Radar données: “Googlebot is the fastest robot d’exploration on the web according to Cloudflare Radar, with Ahrefsbot being the 2nd fastest.” Source
- AI bots are closing in on search bots. From my Cloudflare Radar analysis, search-engine robots d’exploration (Googlebot among les) encore explorer the la plupart, but AI bots are a clair #2 and on pace to overtake les dans a couple of années. Source
- Byte limite: ~2 MB per URL (PDFs 64 MB) — Google’s documented Googlebot récupérer limite as of 2026, with the explicit caveat que the figure isn’t fixed. Source
Googlebot mistakes to éviter
En utilisant robots.txt to essayer to deindex une page. robots.txt arrête exploration,
pas indexation. A robots-blocked URL peut encore montrer up in results (indexé
URL-only) si something liens to it. Do à la place: let Google explorer lune page
and ajouter a noindex tag — that’s the seulement reliable façon to supprimer it.
Trying to target the desktop or smartphone robot d’exploration separately in
robots.txt. Les deux share the unique Googlebot product token, so a rule
written pour un s’applique to les deux. Do à la place: si vous besoin différent
behavior per device, lire the HTTP user-agent requête header in votre propre
server-side logic — robots.txt can’t faire que distinction.
Letting mobile and desktop versions of une page drift apart. Sous indexation mobile-first, the smartphone agent’s view is ce que obtient indexé and ranked. Content, données structurées, metadata, or robots tags que seulement exist on desktop are effectively invisible to Google. Do à la place: garder les deux versions in sync, and vérifier ce que the smartphone agent en réalité sees.
Relying on crawl-delay to slow Googlebot bas. Google doesn’t traiter
the non-standard crawl-delay directive at tout — it’s a Bing-only lever.
Do à la place: retourner 429/500/503 si vous genuinely besoin Google to back
off, understanding que throttles the whole hostname and seulement holds pour a day
or two avant sustained errors commencer dropping pages from the index.
Trusting the Googlebot user-agent string in votre logs at face valeur.
The UA header is trivially spoofed, and plenty of trafic que claims to be
Googlebot isn’t. Do à la place: vérifier with reverse + forward DNS or contre
Google’s publié IP ranges avant treating “Googlebot” trafic in votre
analytics as réel.
Assuming JavaScript-only navigation (click handlers, aucun réel <a href>)
va obtenir découvert and rendered comme a normal lien. Googlebot discovers
URLs primarily via liens, and rendering se produit plus tard in a separate,
stateless queue que doesn’t fire arbitrary UI interactions. Do à la place:
expose every important chemin as a réel <a href> lien que fonctionne sans
JavaScript.
Treating explorer frequency as a ranking lever. Getting crawled plus souvent isn’t a scoreboard — it’s simplement the gate vous have to réussir via to be eligible to rank. Do à la place: spend effort on content and technical health, pas on trying to attract plus explorer hits pour leur propre sake.
Courant Googlebot problèmes
Une page is blocked in robots.txt but encore montre up in search
- Causer: robots.txt seulement arrête exploration, pas indexation. Si autre pages lien to the blocked URL, Google peut encore index it URL-only (usually with aucun snippet/description).
- Fix: si vous en réalité vouloir it out of results, autoriser exploration in
robots.txtand ajouter anoindextag à la place — Googlebot has to be able to lire lune page to voir the tag. Confirmer the fix with the robots-txt-tester outil and re-check l’URL’s status une fois it’s recrawled.
JavaScript-rendered content isn’t showing up dans l’index
- Causer: exploration and rendering are separate steps. The récupéré HTML obtient
queued pour rendering — usually seconds, parfois beaucoup plus long — avant a
headless Chromium executes the JavaScript. Content que dépend on
click-only or non-
<a href>navigation may jamais render at tout, since the Web Rendering Service doesn’t fire arbitrary UI interactions. - Fix: vérifier si le contenu exists in the raw HTML réponse versus
seulement après JS execution, and assurez-vous important paths are réel
<a href>liens. The render-gap outil is construit pour exactly ce vérifier — it montre the gap entre what’s récupéré and ce que en réalité renders.
Server logs montrer a lot of “Googlebot” trafic que seems off
- Causer: the
Googlebotuser-agent string is trivially spoofed. A meaningful share of trafic claiming to be Googlebot in raw logs is impostors. - Fix: vérifier with reverse-DNS (devrait resolve to a hostname ending in
googlebot.com,google.com, orgoogleusercontent.com) followed by forward-DNS back to the même IP, or match contre Google’s publié IP ranges (common-crawlers.json) pour bulk log checks — voir the Scripts tab pour les deux méthodes, or run a log pull via the log-file-analyzer outil, qui flags spoofed Googlebot hits automatically.
Fréquence d’exploration dropped suddenly
- Causer: Googlebot’s fréquence d’exploration is algorithmic and self-throttling —
sustained
429,500,503, or timeout réponses tell it to back off, and que throttle s’applique hostname-wide, pas simplement to the erroring URLs. - Fix: vérifier recent server error rates and réponse codes with the
http-status-checker outil or votre serveur logs. Si errors persist plus
que a day or two, expect pages to commencer dropping from the index, pas simplement
slower exploration. Fixing the underlying
5xx/429source is ce que restores the fréquence d’exploration — there’s aucun separate “unthrottle” lever.
Content que seulement exists on desktop isn’t ranking
- Causer: sous indexation mobile-first, the smartphone agent’s view is ce que obtient indexé and ranked. Si content, données structurées, or metadata seulement exists on the desktop version, it’s effectively invisible to Google.
- Fix: comparer ce que the mobile and desktop versions serve — utiliser the mobile-friendly-tester outil to voir ce que the smartphone agent renders, and bring the two versions back in sync.
Outils pour working with Googlebot
- robots-txt-tester — vérifier si a spécifique
URL is allowed or blocked pour
Googlebotavant vous assume votrerobots.txtis doing ce que vous think. - log-file-analyzer — parse votre serveur logs to voir réel Googlebot explorer activity and flag trafic que claims to be Googlebot but fails IP/DNS verification.
- render-gap — comparer ce que Googlebot récupère versus ce que en réalité renders après JavaScript executes, directement utile pour the crawl-vs-render distinction ce article covers.
- mobile-friendly-tester — vérifier ce que the smartphone robot d’exploration (the principal un sous indexation mobile-first) sees on a page.
- http-status-checker — confirmer the status
codes votre serveur is returning, since sustained
429/5xxréponses are ce que en réalité throttle Googlebot’s fréquence d’exploration.
Third-party: Recherche Google Console’s Inspection d’URL outil montre how Googlebot dernier crawled and rendered a spécifique URL, notamment qui robot d’exploration (mobile or desktop) récupéré it. Screaming Frog peut explorer a site rendering as Googlebot to comparer raw HTML contre rendered output at scale.
Quiz
Five questions to vérifier ce que stuck à propos de Googlebot.
Journal des modifications
Mis à jour le 28 juil. 2026.
Résumé éditorial et détails enregistrés des changements.Détails des changements
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
Comparaison complète indisponible — aucun instantané antérieur n’a été archivé pour cette révision.
Mis à jour le 18 juil. 2026.
Résumé éditorial et détails enregistrés des changements.Détails des changements
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
Comparaison complète indisponible — aucun instantané antérieur n’a été archivé pour cette révision.