Hướng dẫn về Googlebot

Điều gì Googlebot thực ra là — Smartphone so với Desktop, evergreen Chromium kết xuất, người dùng-agent strings, IP-range verification, byte limits, và đó crawl-so với-xếp hạng phân biệt.

Xuất bản lần đầu: 24 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Googlebot là Google web crawler — đó software đó fetches các trang so Google có thể chỉ mục và xếp hạng them. Điều này xuất hiện trong hai variants đó share một robots.txt token: Googlebot Smartphone (chính, dưới mobile-đầu tiên lập chỉ mục) và Googlebot Desktop. Điều này chạy an evergreen Chromium và renders JavaScript trong một tách biệt, sau đó queue (crawl ≠ render). Crawling không một xếp hạng factor, và blocking Googlebot trong robots.txt không đó giống nhau as deindexing — một blocked URL có thể vẫn là được lập chỉ mục URL-chỉ. Người dùng-agent là trivially spoofed, so verify với reverse + forward DNS hoặc Google published IP ranges. 'Googlebot' là thực sự Google Search's slice of một nhiều lớn hơn crawling nền tảng.

Tóm tắt — Googlebot là Google Search’s crawler, split vào Smartphone (chính, mobile-đầu tiên) và Desktop, mà share một Googlebot robots.txt token — Bạn có thể’t đích them riêng. nó chạy evergreen Chromium và renders JavaScript trong tách biệt, sau đó queue (crawl ≠ render). Crawling là bắt buộc để xếp hạng nhưng là không tín hiệu xếp hạng, và robots-blocked URL có thể vẫn là được lập chỉ mục URL-chỉ. Verify nó by reverse + forward DNS để Google domain hoặc so với Google published IP ranges — người dùng-agent là trivially spoofed. và “Googlebot” là thực sự chỉ Tìm kiếm-facing slice của nhiều bigger crawling nền tảng.

Evidence for this claim Googlebot is Google's crawler, with smartphone and desktop crawler types that share the same product token. Scope: Current Googlebot crawler and user-agent documentation. Confidence: high · Verified: Google Search Central: Googlebot Evidence for this claim A claimed Google crawler can be verified using reverse and forward DNS or Google's published IP ranges. Scope: Google's current crawler-verification methods. Confidence: high · Verified: Google Search Central: Verify Googlebot

Điều gì Googlebot thực ra là

Google là precise về đó name: “Googlebot is the generic name for two types of web crawlers used by Google Search.” (bản dịch) «Googlebot là đó generic name cho hai types of web các crawler dùng by Google Search.» Những hai types là Googlebot Smartphone (“a mobile crawler that simulates a user on a mobile device” (bản dịch) «một trình thu thập di động mô phỏng người dùng trên thiết bị di động») và Googlebot Desktop (“a desktop crawler that simulates a user on desktop” (bản dịch) «một trình thu thập máy tính mô phỏng người dùng trên máy tính»).

Điều này không phải một little program đang chạy on một machine. “Googlebot runs on thousands of machines,” (bản dịch) «Googlebot chạy on thousands of machines,» as I mô tả điều này trong my Googlebot hướng dẫn, “they determine how fast and what to crawl on websites,” (bản dịch) «they determine cách fast và điều cần crawl on websites,» distributed trên datacenters worldwide nhưng egressing primarily từ US IP addresses. Phát hiện happens mostly qua links — Google tìm thấy new URLs “primarily from links embedded in previously crawled pages” (bản dịch) «primarily từ links embedded trong previously được crawl các trang» — plus sitemaps. (Cho đó đầy đủ phát hiện và scheduling picture, đó là đó crawling hub job.)

Smartphone so với Desktop — và vì sao đây là “smartphone-first” (bản dịch) «smartphone-đầu tiên»

Dưới mobile-đầu tiên lập chỉ mục, đó smartphone crawler là đó chính một. Google: “For most sites Google Search primarily indexes the mobile version of the content. As such the majority of Googlebot crawl requests will be made using the mobile crawler, and a minority using the desktop crawler.” (bản dịch) «Đối với hầu hết trang web Google Search primarily indexes đó mobile version of đó nội dung. As such đó majority of Googlebot các yêu cầu crawl sẽ là đã làm dùng đó mobile crawler, và một minority dùng đó desktop crawler.» Mobile-đầu tiên lập chỉ mục đã được hoàn tất cho all các trang since October 2023, so đó practical rule là: nếu nội dung không visible để đó smartphone agent, điều này không được lập chỉ mục. Match nội dung của bạn, dữ liệu có cấu trúc, metadata, và robots tags trên mobile và desktop.

Đó robots.txt gotcha: “Both crawler types obey the same product token (user agent token) in robots.txt, and so you cannot selectively target either Googlebot Smartphone or Googlebot Desktop using robots.txt.” (bản dịch) «Cả hai crawler types obey đó giống nhau sản phẩm token (người dùng agent token) trong robots.txt, và so bạn không thể selectively đích either Googlebot Smartphone hoặc Googlebot Desktop dùng robots.txt.» Đó chỉ way để differentiate là để đọc đó HTTP user-agent header yêu cầu trong của bạn own máy chủ-side logic. (Cho đó mechanics of đó header format, see mobile-đầu tiên lập chỉ mụcngười dùng-agent.)

người dùng-agent strings

robots.txt sản phẩm token cho cả hai là chỉ Googlebot. đầy đủ UA strings differ:

Googlebot Desktop:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36

Googlebot Smartphone:

Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

W.X.Y.Z là placeholder cho hiện tại Chrome version, mà moves as Googlebot’s evergreen Chromium là đã cập nhật. không trust string on của nó own, though — nó trivially spoofed (see verification dưới).

Evergreen Chromium và kết xuất queue

Này là đó phân biệt đó trips mọi người lên hầu hết: crawling và kết xuất là tách biệt steps. Googlebot chạy “an evergreen version of Chromium,” (bản dịch) «an evergreen version of Chromium,» và điều này trở thành evergreen back trong Có thể 2019 (jumping từ đó old Chrome 41 để hiện tại ổn định), mà là vì sao điều này hiện tại xử lý ES6+, IntersectionObserver, Web Components, và modern CSS. Nhưng điều này không execute của bạn JavaScript đó moment điều này fetches đó HTML.

Google: “Googlebot queues all pages with a 200 HTTP status code for rendering, unless a robots meta tag or header tells Google not to index the page. The page may stay on this queue for a few seconds, but it can take longer than that. Once Google’s resources allow, a headless Chromium renders the page and executes the JavaScript.” (bản dịch) «Googlebot queues all các trang với một 200 HTTP mã trạng thái cho kết xuất, trừ khi một robots meta tag hoặc header tells Google không để chỉ mục đó trang. Đó trang có thể stay on này queue cho vài seconds, nhưng điều này có thể take lâu hơn đó. Khi Google các tài nguyên cho phép, một headless Chromium renders đó trang và executes đó JavaScript.» Đó kết xuất service (WRS) behaves như một modern trình duyệt nhưng với quirks worth knowing: đây là effectively stateless — local/session storage và cookies là cleared trên trang loads — điều này không fetch images hoặc videos (để save bandwidth), caches aggressively (và có thể bỏ qua của bạn bộ nhớ đệm các header), và không hỗ trợ WebSockets hoặc WebRTC. Nếu nội dung của bạn chỉ xuất hiện sau một nhấp hoặc một JS-driven navigation đó không một real <a href> link, expect kết xuất trouble. (Depth lives trong đó kết xuất sibling.)

Byte limits

Googlebot không download an unlimited amount theo URL. As of Google March 2026 Bên trong Googlebot cập nhật, điều này fetches khoảng đó đầu tiên 2 MB of bất kỳ individual URL (including đó HTTP header) và lên để 64 MB cho một PDF. Đó hình là một moving đích — Google own cách diễn đạt là đó “this limit is not set in stone and may change over time as the web evolves and HTML pages grow in size,” (bản dịch) «này limit không phải set trong stone và có thể thay đổi theo thời gian as đó web evolves và HTML các trang grow trong size,» và trước đó tài liệu listed 15 MB (mà turned out để là đó rộng hơn infrastructure default, không Tìm kiếm number). Đó practical point holds regardless: bất cứ điều gì past đó cutoff đơn giản không fetched — “to Googlebot, they simply don’t exist.” (bản dịch) «để Googlebot, they đơn giản không exist.» Giữ cốt yếu nội dung và markup trên đó bloat.

Worked thất bại: canonical tồn tại, nhưng Googlebot không bao giờ nhận nó

Imagine sản phẩm template returning 2,4 MB của HTML. application serializes huge sản phẩm-state object và khuyến nghị payload near top của document; canonical, sản phẩm mô tả, dữ liệu có cấu trúc, và related-sản phẩm links làm không xuất hiện cho đến khi khoảng byte 2 180 000. trình duyệt downloads toàn bộ phản hồi, so View Nguồn looks đúng. Googlebot Tìm kiếm dừng khoảng của nó được ghi lại 2 MB limit, so những điều đó sau đó các tín hiệu không exist trong fetched tài nguyên.

Diagnose phản hồi trong byte order, không chỉ trong được kết xuất DOM:

curl -sS -D response-headers.txt -o page.html https://example.com/product
wc -c response-headers.txt page.html
LC_ALL=C grep -abo 'rel="canonical"' page.html
LC_ALL=C grep -abo 'application/ld+json' page.html

local byte được tính là approximation vì phân phối intermediaries và phản hồi xử lý có thể differ, nhưng họ câu trả lời hữu ích đầu tiên câu hỏi: là cốt yếu các tín hiệu comfortably sớm, hoặc là họ sitting near hoặc beyond boundary? khắc phục là để xóa hoặc defer oversized inline dữ liệu và emit essential metadata, chính nội dung, và crawlable links sớm—không để move giống nhau bloat khoảng và hope cutoff thay đổi.

Cách Googlebot fetches — politely

  • Tốc độ crawl là algorithmic và self-throttling. “For most sites, Googlebot shouldn’t access your site more than once every few seconds on average.” (bản dịch) «Đối với hầu hết trang web, Googlebot không nên access trang web của bạn hơn khi mỗi một vài seconds on average.» Điều này speeds lên hoặc backs off dựa trên máy chủ của bạn health.
  • Các mã trạng thái là đó lever. Returning 429, 500, hoặc 503 tells Googlebot để chậm xuống — nhưng đó ảnh hưởng đó entire hostname, không chỉ đó erroring URLs, và chỉ hoạt động cho một day hoặc hai trước sustained các lỗi bắt đầu dropping các trang từ đó chỉ mục. John Mueller: “I’d only expect the crawl rate to react that quickly if they were returning 429 / 500 / 503 / timeouts,” (bản dịch) «I’d chỉ expect đó tốc độ crawl để react đó quickly nếu they đã là returning 429 / 500 / 503 / timeouts,»“404s are generally fine & once discovered, Googlebot will retry them anyway.” (bản dịch) «404s là generally fine & khi discovered, Googlebot sẽ retry them anyway.»
  • crawl-delay là đã bỏ qua. Google không xử lý đó non-tiêu chuẩn crawl-delay robots.txt directive tại all. (Bing làm honor điều này — một of đó real Googlebot/Bingbot divergences.)
  • Crawling tracks crawl demand, không một flat quota — capacity (điều gì máy chủ của bạn có thể take) plus demand (popularity và staleness). Đối với hầu hết trang web này là một non-vấn đề; điều này chỉ bites tại real quy mô. Đó đầy đủ treatment là trong ngân sách crawl.

Verifying nó thực sự Googlebot

Người dùng-agent header là “often spoofed by other crawlers” (bản dịch) «thường spoofed by other các crawler» — so điều này alone proves không có gì. Google các crawler identify themselves three ways: người dùng-agent header, đó nguồn IP, và đó reverse-DNS hostname of đó IP. Hai real verification các phương thức:

  1. Manual (một-off). Reverse-DNS nguồn IP; xác nhận nó resolves để hostname ending trong googlebot.com, google.com, hoặc googleusercontent.com ( mask looks như crawl-***-***-***-***.googlebot.com); sau đó forward-DNS đó hostname và xác nhận nó trả về gốc IP.
  2. Tự động (tại quy mô). Match IP so với Google published CIDR ranges. Google có split những điều này từ old single googlebot.json vào several JSON files by crawler category — Googlebot một là https://www.gstatic.com/ipranges/common-crawlers.json ( legacy googlebot.json URL vẫn các chuyển hướng để giống nhau dữ liệu).

Cả hai là trong đó Scripts tab, cho macOS/Linux và Windows. Vì sao bother? Plenty of traffic lies về đang Googlebot, so logs đó “show Googlebot” (bản dịch) «cho thấy Googlebot» có thể là largely impostors — verify trước khi bạn trust.

Googlebot là một bot trong fleet

“Googlebot” là genuinely một bit of một misnomer. Gary Illyes, March 2026: “I mean, calling it Googlebot, that’s a misnomer,” (bản dịch) «I có nghĩa là, calling điều này Googlebot, đó là một misnomer,»“Googlebot is not our crawler infrastructure.” (bản dịch) «Googlebot không phải của chúng ta crawler infrastructure.» Đó infrastructure underneath, trong his words, là “software as a service, if you like. SaaS” (bản dịch) «software as một service, nếu bạn như. SaaS» — một shared nền tảng nhiều Google các sản phẩm draw từ. Điều gì bạn see trong của bạn logs là đó Tìm kiếm slice of điều này: “When you see Googlebot in your server logs, you are just looking at Google Search.” (bản dịch) «Khi bạn see Googlebot trong máy chủ của bạn logs, bạn là chỉ looking tại Google Search.» He cũng notes có “dozens, if not hundreds of different crawlers,” (bản dịch) «dozens, nếu không hundreds of khác nhau các crawler,» hầu hết cũng nhỏ để bother documenting.

named ones bạn’ll thực ra đáp ứng alongside Googlebot bao gồm Googlebot-Image, Googlebot-Video, và Googlebot-News (mà share Googlebot’s strings/tokens), Storebot-Google, và Google-InspectionTool (powers URL Inspection và Rich Kết quả tools). Hai behave unusually: AdsBot bỏ qua global * robots.txt rule ( Disallow: / dưới User-agent: * vẫn sẽ không dừng nó), và Google-Safety bỏ qua robots.txt hoàn toàn. AI-related các crawler — Google-Extended (controls Gemini training; không tín hiệu xếp hạng) và GoogleOther (R&D crawl, offloaded từ Googlebot) — exist cũng, nhưng AI các crawler sibling covers những điều đó trong depth, so I’ll point ở đó thay vì duplicate.

Cách control Googlebot

Three controls, three khác effects:

  • robots.txt dừng crawling, không lập chỉ mục. sử dụng nó để giữ bots out của thấp-giá trị URL spaces — không bao giờ as deindexing tool.
  • noindex dừng lập chỉ mục — nhưng Googlebot phải là được phép để crawl trang để see tag trong đầu tiên place.
  • Password protection chặn access hoàn toàn.

Mà brings us để đó single hầu hết misunderstood Googlebot fact: “There’s a difference between crawling and indexing; blocking Googlebot from crawling a page doesn’t prevent the URL of the page from appearing in search results.” (bản dịch) «có một khác biệt giữa crawling và lập chỉ mục; blocking Googlebot từ crawling một trang không ngăn đó URL of đó trang từ appearing trong kết quả tìm kiếm.» MỘT robots-blocked URL có thể vẫn nhận được lập chỉ mục URL-chỉ nếu điều gì đó links để điều này. Để thực ra xóa một trang, cho phép crawling và thêm noindex. I’ve được viết này lên trong detail trong Được lập chỉ mục, though blocked by robots.txt.

cho wider pipeline Googlebot lives bên trong — phát hiện URL, crawl scheduler, kết xuất, và crawl-so với-chỉ mục-so với-xếp hạng distinctions — see crawling hub và Cách Tìm kiếm Hoạt động.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.