Người dùng Agent
Điều gì một người dùng agent là — đó HTTP header các crawler và các trình duyệt dùng để identify themselves, đó robots.txt token so với. đó đầy đủ string, và cách verify một bot là real.
Ngôn ngữ
2 tín hiệu bằng chứng trên trang này
- Dữ liệu nguồn được liên kếtphổ biến-các crawler.json
- Công cụ trực tuyến liên quanGooglebot Verifier
MỘT người dùng agent là đó HTTP header mỗi client — trình duyệt, crawler, hoặc bot — gửi để identify itself. Hai điều nhận confused: đó đầy đủ người dùng-agent *string* trong đó header yêu cầu, và đó ngắn người dùng-agent *token* (Googlebot, bingbot, Google-Extended) bạn đích trong robots.txt. Đó token là một substring of đó string (RFC 9309); some tokens, như Google-Extended, có không yêu cầu string tại all. Đó string là trivially spoofed — Google says của nó own là 'thường spoofed' — so không bao giờ trust điều này cho access control. Verify Googlebot/Bingbot by reverse DNS plus một forward lookup, hoặc so với published IP ranges. Và watch đó gotchas: AdsBot và Google-Safety bỏ qua `User-agent: *`, version numbers và wildcards trong đó token line là đã bỏ qua, matching là case-insensitive, và serving khác nhau nội dung để một bot UA hơn để người dùng là cloaking.
Evidence for this claim HTTP User-Agent is a request field containing product information supplied by the client; it is descriptive text and not proof of identity. Scope: HTTP semantics for User-Agent. Confidence: high · Verified: IETF RFC 9110: User-Agent Evidence for this claim robots.txt User-agent matching is defined by the Robots Exclusion Protocol and controls crawler access, not authentication or general HTTP content negotiation. Scope: RFC 9309 robots matching behavior. Confidence: high · Verified: IETF RFC 9309: Robots Exclusion ProtocolTL;DR — MỘT người dùng agent là một little line of text mỗi trình duyệt và mỗi bot gửi với mỗi yêu cầu để chẳng hạn “here’s who I am.” (bản dịch) «ở đây ai I am.» Google crawler says đây là Googlebot; Bing says bingbot. Trong
robots.txtbạn không ghi đó toàn bộ line — bạn dùng một ngắn name (một “token”) nhưGooglebot. Và ở đây đó catch: đó line là chỉ text, so anyone có thể fake điều này. Đó chỉ real way để know một bot là ai điều này claims để là để kiểm tra nơi của nó yêu cầu thực ra nghĩ ra từ.
Điều gì người dùng agent là
Mỗi khi trình duyệt của bạn loads một trang, điều này gửi along một ngắn text label đó says điều gì điều này là — điều gì đó như “I’m Chrome on a Mac.” (bản dịch) «I’m Chrome on một Mac.» Đó label là đó người dùng agent, và điều này travels trong an HTTP header on mỗi yêu cầu. Các máy chủ có thể đọc điều này và react để điều này.
quan trọng catch lên front: client fills trong đó label itself. Không có gì kiểm tra nó. nó claim, không credential — so người dùng agent đó nói “Googlebot” không phải giống nhau điều as yêu cầu đó thực ra verified as Googlebot.
Các crawler làm giống nhau điều. Khi Googlebot fetches của bạn trang, nó gửi người dùng agent
đó bao gồm Googlebot. Khi Bingbot fetches nó, người dùng agent bao gồm
bingbot. đó Cách bot announces itself trong của bạn máy chủ nhật ký.
string so với. ngắn name
có thực sự hai điều mọi người có nghĩa là by “người dùng agent,” và mixing them lên gây ra lot của confusion:
- ** người dùng-agent string** là đầy đủ line trong yêu cầu header. Googlebot’s là dài và looks lot như trình duyệt.
- ** người dùng-agent token** là ngắn name bạn sử dụng trong
robots.txtđể đích bot — nhưGooglebothoặcbingbot. token là chỉ piece của đầy đủ string, không toàn bộ điều.
So Khi bạn ghi rule trong robots.txt, bạn sử dụng ngắn token:
User-agent: Googlebot
Disallow: /private/bạn không paste giant trình duyệt-looking string trong ở đó.
Bạn có thể’t trust string
Này là đó một điều để remember. Người dùng-agent line là đơn giản text, so bất cứ điều gì có thể fake điều này. Bất kỳ script có thể claim để là Googlebot trong một single line of code — và plenty làm, để sneak past chặn. Google itself says đó Googlebot header là “often spoofed.” (bản dịch) «thường spoofed.»
đó có nghĩ là bạn nên không bao giờ quyết định ai nhận access để của bạn trang web based chỉ on người dùng agent. nếu bạn thực ra cần để xác nhận khách truy cập là thực Googlebot (chẳng hạn, bạn’re reading của bạn nhật ký), bạn verify by kiểm tra nơi yêu cầu nghĩ ra từ — không Điều gì nó nói nó là. Nâng cao tab walks qua chính xác Cách.
một vài gotchas
robots.txtngười dùng-agent names là case-insensitive —Googlebotvàgooglebotlà giống nhau.- Blocking mọi thứ với
User-agent: *làm không block tất cả Google bots — của nó quảng cáo các crawler và safety crawler bỏ qua wildcard. - Cho thấy một version của trang để crawler và khác một để thực mọi người là cloaking, và Google xử lý nó as spam.
Muốn đầy đủ picture — token các bảng, mỗi Google và Bing crawler, chính xác verification commands, và cloaking rules — chuyển để Nâng cao tab.
Evidence for this claim HTTP User-Agent is a request field containing product information supplied by the client; it is descriptive text and not proof of identity. Scope: HTTP semantics for User-Agent. Confidence: high · Verified: IETF RFC 9110: User-Agent Evidence for this claim robots.txt User-agent matching is defined by the Robots Exclusion Protocol and controls crawler access, not authentication or general HTTP content negotiation. Scope: RFC 9309 robots matching behavior. Confidence: high · Verified: IETF RFC 9309: Robots Exclusion ProtocolTL;DR — MỘT người dùng agent là đó HTTP header yêu cầu bất kỳ client gửi để identify itself; đây là tùy chọn, client-filled metadata, không authenticated identity. Của nó giá trị là người dùng-agent string. Tách biệt từ đó là người dùng-agent token (sản phẩm token) dùng trong
robots.txt— RFC 9309 says điều này NÊN là một substring of đó string, một mạnh convention với được ghi lại exceptions (Google-Extended có không yêu cầu string tại all). Matching là case-insensitive, version numbers/wildcards trong đó token line là đã bỏ qua, đó hầu hết-cụ thể group wins, và giống nhau-token groups hợp nhất nhưng không bao giờ hợp nhất với*. Đó string là trivially spoofed — Google calls của nó own “often spoofed” (bản dịch) «thường spoofed» — so verify by reverse + forward DNS (behind bất kỳ proxy/CDN, dùng đó client thật IP) so vớigooglebot.com/google.com/googleusercontent.comcho Google hoặcsearch.msn.comcho Bing, hoặc match published IP ranges — và ngay cả một verified yêu cầu chỉ proves một yêu cầu arrived, không đó trang đã là được lập chỉ mục, retrieved, hoặc dùng cho AI training. AdsBot và Google-Safety bỏ quaUser-agent: *. Chrome là cũng freezing detail out of trình duyệt UA strings (Người dùng-Agent reduction); Client Hints là đó structured nhưng opt-trong replacement, và neither substitutes cho crawler verification. Người dùng-agent adaptation có thể là legitimate, nhưng deceptively cho thấy các crawler materially khác nhau nội dung có thể là cloaking.
header, string, và token
Three điều, và giữ them straight là phần lớn của điều này topic.
- ** header.**
User-Agentlà HTTP yêu cầu header. mỗi client gửi nó: của bạn trình duyệt,curl, crawler, bot. Theo RFC 9110 ( cốt lõi HTTP semantics tiêu chuẩn), nó tùy chọn trường client fills trong — client-supplied descriptive metadata, không authenticated identity máy chủ có verified. - ** string.** header giá trị — freeform line describing software, version, kết xuất engine, và đôi khi OS.
- ** token.** ngắn identifier được sử dụng trong
robots.txtUser-agent:lines để đích crawler —Googlebot,bingbot,Google-Extended.
Đó mối quan hệ là đó part đó trips mọi người lên. RFC 9309 (đó formal Robots
Exclusion Giao thức tiêu chuẩn) says đó token “SHOULD be a substring of the
identification string that the crawler sends… in the case of HTTP, the product token
SHOULD be a substring in the User-Agent header.” (bản dịch) «NÊN là một substring of đó identification string đó crawler gửi… trong đó case of HTTP, đó sản phẩm token NÊN là một substring trong người dùng-Agent header.» đó là một SHOULD, không một MUST —
một mạnh convention đó tiêu chuẩn khuyến nghị, không một hard requirement mỗi crawler là
mechanically bound để. Google-Extended (dưới) là đó clearest ví dụ of một
được ghi lại exception để điều này. không đọc đó substring rule as universal chỉ vì
Google follows điều này cho hầu hết of của nó own tokens. Đó token là part of đó string khi
một provider làm supply một; bạn đích đó token trong robots.txt và đọc đó string
trong của bạn logs.
Google own cách diễn đạt of cách của nó bots identify themselves là hữu ích ở đây: “Google’s
crawlers identify themselves through three things: the HTTP user-agent request
header, the source IP address of the request, and the reverse DNS hostname of the
source IP.” (bản dịch) «Google các crawler identify themselves qua three điều: đó HTTP header yêu cầu, đó nguồn IP address of đó yêu cầu, và đó reverse DNS hostname of đó nguồn IP.» Note đó người dùng-agent là chỉ một of đó three — đó other hai là
cách bạn thực ra verify điều này.
Google-Extended: token với không string
Đó sạch nhất illustration of token ≠ string là Google-Extended. Điều này controls
liệu Google có thể dùng nội dung của bạn cho Gemini training và grounding — và điều này có
không dedicated HTTP yêu cầu người dùng-agent string tại all. Đó crawling itself là đã xong
với existing Googlebot strings; Google-Extended tồn tại chỉ as một robots.txt
control token. Bạn’ll không bao giờ see “Google-Extended” (bản dịch) «Google-Extended» trong một header yêu cầu trong của bạn logs.
practical consequence: blocking Google-Extended ảnh hưởng chỉ AI-training sử dụng
của bạn nội dung — nó làm không dừng Googlebot từ crawling và lập chỉ mục bạn cho
Tìm kiếm. họ’re tách biệt decisions controlled by tách biệt tokens. (cho rộng hơn
picture của bots reading của bạn trang web, see AI các crawler và crawler.)
Googlebot’s người dùng-agent strings
Googlebot là “evergreen” — nó chạy on gần đây version của Chrome, và Chrome
version trong của nó string cập nhật periodically (nó có since December 2019). Đó là lý làm
version xuất hiện as W.X.Y.Z placeholder:
Googlebot Smartphone (mobile):
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)Googlebot Desktop:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36Hai điều để internalize. đầu tiên, không hardcode version — W.X.Y.Z thay đổi,
và matching on nó sẽ break. Match ổn định token Googlebot thay vì. thứ hai,
Bạn có thể’t tách biệt mobile từ desktop trong robots.txt. Cả hai variants share một
Googlebot token, so robots.txt rule áp dụng để cả hai.
Google crawler tokens
Google chạy toàn bộ family của các crawler và fetchers, mỗi với của nó own token. ones bạn’ll đáp ứng phần lớn:
| Crawler | robots.txt token | Notes |
|---|---|---|
| Googlebot | Googlebot | Tìm kiếm, Images, Video, News, Discover — mobile + desktop share điều này token |
| Googlebot Image | Googlebot-Image | Google Images |
| Googlebot Video | Googlebot-Video | Video Tìm kiếm |
| Googlebot News | Googlebot-News | Dùng various Googlebot strings |
| Google StoreBot | Storebot-Google | Shopping |
| Google-InspectionTool | Google-InspectionTool | Powers Tìm kiếm kiểm thử tools |
| GoogleOther | GoogleOther | Internal research/fetching |
| Google-Extended | Google-Extended | robots.txt-chỉ — Gemini training, không yêu cầu string |
và ones đó break thông thường rules — special-case các crawler đó bỏ qua
User-agent: *:
- AdsBot (
AdsBot-Google) và AdsBot Mobile (AdsBot-Google-Mobile) — họ không obey wildcard. để block them bạn phải name them explicitly. - AdSense (
Mediapartners-Google) — giống nhau; bỏ qua global*. - Google-Safety — được sử dụng cho malware/abuse detection; nó bỏ qua robots.txt hoàn toàn.
Đó implication là đó một mọi người miss: User-agent: * không block AdsBot hoặc
Google-Safety. Nếu bạn “block all bots” (bản dịch) «block all bots» với một wildcard và assume AdsBot là đã biến mất,
điều này không. (Này là chính xác đó kind of surprise đó lands một trang trong được lập chỉ mục though
Bị chặn bởi robots.txt territory — see robots.txt cho đó đầy đủ control story.)
Bingbot’s người dùng-agent strings
Bing rebuilt Bingbot’s string trong 2022 để reflect đó nó renders với Microsoft Edge. hiện tại strings:
Bingbot Desktop:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36Bingbot Mobile:
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)robots.txt token là chỉ bingbot. điều để watch: post-2022, Bingbot’s
string looks gần như chính xác như thực Chrome/Edge trình duyệt — chỉ tell là
bingbot/2.0 fragment bên trong nó. nếu bạn có bất kỳ logic đó filters hoặc detects bots
by UA, đó thay đổi matters.
Cách robots.txt thực ra matches token
một vài rules govern mà group của rules crawler obeys (theo Google robots.txt spec và RFC 9309):
- Hầu hết-cụ thể match wins. Google “determines the correct group of rules by
finding… the group with the most specific user agent that matches the crawler’s
user agent.” (bản dịch) «determines đó correct group of rules by finding… đó group với đó hầu hết cụ thể người dùng agent đó matches đó crawler người dùng agent.» MỘT
Googlebotgroup beats một*group cho Googlebot. - Giống nhau-token groups hợp nhất — nhưng không bao giờ với
*. Multiple groups naming đó giống nhau agent là combined vào một. MỘT cụ thể-agent group và đó*group là không đã hợp nhất;*là chỉ đó fallback khi không có gì cụ thể matches. - Case-insensitive. Trường name và giá trị cả hai —
Googlebot,googlebot,GOOGLEBOTlà tương đương. - Version numbers và wildcards trong đó token line là đã bỏ qua. Theo Google,
“both
googlebot/1.2andgooglebot*are equivalent togooglebot.” (bản dịch) «cả hai và là tương đương để .» Bạn không thể ghiUser-agent: Googlebot*để match một family — đó*ở đó làm không có gì.
So User-agent: line takes token và matches nó as đơn giản (case-insensitive)
substring của crawler identity — không version pinning, không wildcards bên trong nó.
Vì sao Bạn có thể’t trust string — và Cách verify
Người dùng-agent string là freeform text. Bất cứ điều gì có thể set điều này. Một line of curl
sẽ claim để là Googlebot, và plenty of tools và malicious bots làm chính xác đó để
slip past chặn. Google says so trong của nó own Googlebot tài liệu: “the HTTP user-agent
request header used by Googlebot is often spoofed by other crawlers.” (bản dịch) «đó HTTP người dùng-agent header yêu cầu dùng by Googlebot là thường spoofed by other các crawler.» As I’ve put điều này
trong my Googlebot hướng dẫn, “Many SEO tools and some malicious bots will pretend to be
Googlebot. This may allow them to access websites that try to block them.” (bản dịch) «Nhiều SEO tools và some malicious bots sẽ pretend để là Googlebot. Này có thể cho phép them để access websites đó try để block them.»
So không bao giờ làm access hoặc nội dung decision on string alone. Verify thay vì.
Một prerequisite trước khi either phương thức: nhận thực nguồn IP. nếu của bạn trang web sits
behind reverse proxy, load balancer, hoặc CDN, address trong của bạn default access log
có thể là proxy IP, không crawler — bạn cần gốc client IP (thường
forwarded trong header như X-Forwarded-For, configured correctly tại của bạn proxy) hoặc
neither verification phương thức dưới có nghĩ là bất cứ điều gì.
Phương thức 1 — reverse + forward DNS (best cho spot kiểm tra). Google hai steps:
- “Run a reverse DNS lookup on the accessing IP address from your logs, using the
hostcommand. Verify that the domain name is eithergooglebot.com,google.com, orgoogleusercontent.com.” (bản dịch) «Chạy một reverse DNS lookup on đó accessing IP address từ của bạn logs, dùng đó command. Verify đó domain name là either , , hoặc .» - “Run a forward DNS lookup on the domain name retrieved in step 1… Verify that it’s the same as the original accessing IP address from your logs.” (bản dịch) «Chạy một forward DNS lookup on đó domain name retrieved trong step 1… Verify đó đây là đó giống nhau as đó original accessing IP address từ của bạn logs.»
cho Bingbot, giống nhau hai-step dance, nhưng hostname phải end trong search.msn.com
(không Bing-branded domain — phổ biến surprise). Commands là trong Scripts tab.
Phương thức 2 — published IP ranges (best tại quy mô). Google không publish một static allowlist cho hardcoding (“these IP address ranges can change” (bản dịch) «những IP address ranges có thể thay đổi»), nhưng điều này làm publish machine-readable CIDR JSON files bạn có thể match so với (phổ biến-các crawler.json và đó rộng hơn crawler files). Bing hiện tại publishes của nó ranges cũng. I được xây dựng một Googlebot IP verification tool cho chính xác này — paste trong IPs và điều này classifies them. Bing Quản trị viên web Tools có một được xây dựng-trong “Verify Bingbot” (bản dịch) «Verify Bingbot» tool as well.
DNS là tốt hơn cho một-off log kiểm tra; IP-range matching là tốt hơn cho verifying tại volume. sử dụng whichever fits — nhưng sử dụng một của them. và treat cả hai dự kiến hostnames và range files as hiện tại as của hôm nay, không vĩnh viễn — Google và Bing có changed những điều này paths trước khi ( IP-range JSON files moved và là renamed since điều này bài viết là đầu tiên được viết), so re-kiểm tra trực tiếp verification doc nếu lookup đó được sử dụng để hoạt động dừng matching.
UA match không phải proof của downstream outcomes
Ngay cả fully verified yêu cầu — thực Googlebot IP, forward-confirmed reverse DNS, mọi thứ kiểm tra out — chỉ proves một điều: đó yêu cầu reached của bạn máy chủ. nó tempting để round đó lên vào nhiều bigger claim, nhưng mỗi của những điều này là tách biệt fact requiring tách biệt evidence:
- yêu cầu đã nhận — yêu cầu với đó người dùng agent hit của bạn máy chủ. (Điều gì log verification thực ra proves.)
- Identity confirmed — yêu cầu thực sự nghĩ ra từ crawler nó claims để là. (Điều gì reverse DNS / IP-range matching adds on top.)
- nội dung fetched và được kết xuất — crawler successfully được kết xuất trang (không các lỗi, không blocked các tài nguyên). không guaranteed chỉ vì yêu cầu landed.
- được lập chỉ mục — URL đã làm nó vào tìm kiếm chỉ mục. thành công fetch không bảo đảm lập chỉ mục.
- được sử dụng cho retrieval, citation, hoặc training — cho AI các crawler especially (Google-Extended, GPTBot, và rest), crawl không phải proof của bạn nội dung là retrieved cho cụ thể câu trả lời, cited, hoặc được sử dụng trong model training. những điều đó là tách biệt, mostly unobservable steps downstream của crawl.
verified Googlebot hit trong của bạn nhật ký là thực tín hiệu — chỉ không stretch nó further hơn Điều gì nó thực ra hiển thị.
Web Bot Auth: nơi verification là heading
Trong 2026 Google began experimenting với Web Bot Auth — “an experimental
cryptographic protocol used to authenticate requests sent by bots.” (bản dịch) «an experimental cryptographic giao thức được dùng để authenticate các yêu cầu đã gửi by bots.» Đó ý tưởng là để
“move beyond easily spoofed headers to a verified identity and decouple agent
identity from IP addresses.” (bản dịch) «move beyond easily spoofed các header để một verified identity và decouple agent identity từ IP addresses.» Bots cryptographically sign của họ các yêu cầu; các trang verify
đó signature so với Google published công khai keys, và signed các yêu cầu carry một
Signature-Agent header. Google own caveat matters: “We don’t sign every request
of a particular agent. Be sure that you fall back to the established methods of bot
verification.” (bản dịch) «We không sign mỗi yêu cầu of một particular agent. Là sure đó bạn fall lại để đó established các phương thức of bot verification.» So đây là additive, không một replacement — reverse DNS và IP ranges vẫn
của bạn baseline hôm nay.
các trình duyệt là getting harder để parse từ UA string cũng
Mọi thứ trên là về các crawler, nhưng đó giống nhau “don’t over-trust the string” (bản dịch) «không over-trust đó string» lesson áp dụng để các trình duyệt, và đây là getting stronger. Chrome đã được rolling out Người dùng-Agent reduction: freezing hoặc coarsening parts of của nó UA string (đầy đủ trình duyệt version, OS version, device model) thay vì reporting them chính xác, so đó string không thể là được dùng để fingerprint một cụ thể người dùng. Google own cách diễn đạt: “The granularity and abundance of detail can lead to user identification. The default availability of this information can lead to covert tracking.” (bản dịch) «Đó granularity và abundance of detail có thể lead để người dùng identification. Đó default availability of này information có thể lead để covert tracking.» Practically, đó có nghĩa là UA-string phân tích cú pháp cho chính xác trình duyệt/OS/device version — analytics, device detection, bug triage — là increasingly unreliable và sẽ chỉ nhận hơn so.
Đó replacement Chrome khuyến nghị là Người dùng-Agent Client Hints (UA-CH): structured
dữ liệu đó trình duyệt gửi chỉ khi một máy chủ explicitly asks cho điều này. Thấp-entropy hints
(trình duyệt brand, major version, mobile flag) go out theo mặc định; cao-entropy hints
(chính xác version, nền tảng version, device model) require đó máy chủ để opt trong qua an
Accept-CH header phản hồi đầu tiên — an rõ ràng negotiation, không một broadcast. Hai
caveats trước khi bạn lean on điều này: đây là một Chrome/Chromium-family mechanism, không điều gì đó
mỗi trình duyệt gửi, và ngay cả nơi đây là supported, “the value may be blank, not
returned, or populated with a varying value.” (bản dịch) «đó giá trị có thể là blank, không đã trả về, hoặc populated với một varying giá trị.» Client Hints solve đó trình duyệt-string
vấn đề; they không phải một crawler-verification mechanism — Google và Bing vẫn
verify của họ own các crawler qua DNS và IP ranges, không Client Hints.
Người dùng-agent targeting và cloaking
Đó tempting move — “detect Googlebot by its UA and serve it something special” (bản dịch) «detect Googlebot by của nó UA và serve điều này điều gì đó special» — là cả hai technically fragile và một policy violation.
Fragile, vì Google không crawl với một UA. bạn’d có để correctly xử lý Googlebot (mobile và desktop), Google-InspectionTool, AdsBot, GoogleOther, và nhiều hơn, từ rotating IPs — practically không thể để whitelist cleanly.
MỘT policy violation, vì serving khác nhau nội dung để một crawler hơn để người dùng là cloaking: “presenting different content to users and search engines with the intent to manipulate search rankings and mislead users.” (bản dịch) «presenting khác nhau nội dung để người dùng và các công cụ tìm kiếm với đó intent để manipulate tìm kiếm thứ hạng và mislead người dùng.» Đó hình phạt ranges từ algorithmic demotion để đầy đủ deindexing. Note đó line: legitimate adaptation (responsive layouts, nội dung negotiation) là fine — đây là swapping đó nội dung itself giữa bots và người dùng đó crosses vào cloaking.
cho nơi người dùng agent sits trong bigger pipeline, see crawling ( hub) và crawler. cho controlling Điều gì những điều đó bots là được phép để fetch, see robots.txt.
AI summary
condensed take on Nâng cao version:
- Three điều, kept straight: đó
User-Agentyêu cầu header — tùy chọn, client-filled metadata theo RFC 9110, không authenticated identity; của nó giá trị, đó người dùng-agent string; và đó token dùng trongrobots.txt. Theo RFC 9309 đó token NÊN là một substring of đó string — một mạnh convention, không một universal rule. - Some tokens có không string.
Google-Extendedlà đó được ghi lại exception: điều này tồn tại chỉ as một robots.txt control (Gemini training); crawling dùng thông thường Googlebot strings. Blocking điều này không ảnh hưởng Tìm kiếm lập chỉ mục. - Googlebot/Bingbot strings là evergreen — đó Chrome version cho thấy as
W.X.Y.Zvà thay đổi; match đó ổn định token (Googlebot,bingbot), không bao giờ đó version. Mobile và desktop Googlebot share một token. Post-2022 Bingbot looks như một real trình duyệt except cho đóbingbot/2.0fragment. - robots.txt matching: hầu hết-cụ thể group wins; giống nhau-token groups hợp nhất nhưng không bao giờ
với
*; matching là case-insensitive; version numbers và wildcards trong đóUser-agent:line là đã bỏ qua (Googlebot*=Googlebot). - AdsBot và Google-Safety bỏ qua
User-agent: *— block them by name hoặc không tại all. - Đó string là trivially spoofed (Google calls của nó own “often spoofed” (bản dịch) «thường spoofed»). Nhận đó
client thật IP đầu tiên (proxies/CDNs có thể mask điều này trong của bạn logs), thì verify by
reverse + forward DNS (
googlebot.com/google.com/googleusercontent.com; Bing →search.msn.com) hoặc published IP ranges — files Google có renamed/moved trước, so re-kiểm tra đó trực tiếp doc nếu một lookup dừng matching. Không bao giờ trust đó string cho access control. - MỘT verified yêu cầu vẫn không proof of mọi thứ downstream. Yêu cầu đã nhận, identity confirmed, nội dung được kết xuất, trang được lập chỉ mục, và nội dung retrieved/cited/trained-on là tách biệt claims needing tách biệt evidence — một log hit proves đó đầu tiên, không có gì khác tự động.
- Web Bot Auth (2026, experimental) signs các yêu cầu cryptographically — additive, với DNS/IP vẫn đó fallback.
- Trình duyệt UA strings là getting harder để parse cũng: Chrome Người dùng-Agent reduction freezes chính xác version/OS/device detail out of đó string; Client Hints là đó structured, opt-trong replacement — nhưng họ là Chrome-cụ thể và không substitute cho crawler verification.
- Serving khác nhau nội dung by UA là cloaking — một spam-policy violation, và fragile vì Google crawl với nhiều UAs.
Tài liệu chính thức
Chính-nguồn tài liệu từ các công cụ tìm kiếm và tiêu chuẩn.
- Overview of Google các crawler và fetchers (người dùng agents) — đó three identification các tín hiệu và đó đầy đủ crawler list.
- Google phổ biến các crawler — đó token + người dùng-agent-string bảng.
- Google Special-Case Các crawler — AdsBot, Mediapartners-Google, Google-Safety và đó
*exceptions. - Verify Các yêu cầu từ Google Các crawler và Fetchers — đó hai-step DNS phương thức và đó IP-range files.
- Điều gì Là Googlebot — đó Googlebot strings và đó “often spoofed” (bản dịch) «thường spoofed» note.
- Cách Google Interprets đó robots.txt Đặc tả — token matching, case-insensitivity, đã bỏ qua wildcards.
- Authenticating Các yêu cầu với Web Bot Auth (Experimental) — đó cryptographic signing giao thức.
- Updating người dùng agent of Googlebot (2019) — vì sao đó string cho thấy
W.X.Y.Z. - Spam Policies — Cloaking — đó definition và đó hình phạt.
Bing / Microsoft
- Announcing người dùng-agent thay đổi cho Bingbot (Apr 2022) — hiện tại desktop + mobile strings (Fabrice Canel).
- Cách Verify đó Bingbot là Bingbot (Aug 2012) —
*.search.msn.comreverse-DNS phương thức. - Bing Quản trị viên web Tools — Verify Bingbot — được xây dựng-trong verification tool.
** tiêu chuẩn**
- RFC 9110 — HTTP Semantics, §10.1.5 Người dùng-Agent — cốt lõi HTTP spec:
User-Agentlà tùy chọn, client-supplied trường, không authenticated identity. - RFC 9309: Robots Exclusion Giao thức — formal
SHOULD-cấp độ definition của sản phẩm token as substring của người dùng-agent string. - MDN — Người dùng-Agent header — HTTP syntax của header itself.
trình duyệt UA strings
- Chrome Quyền riêng tư Sandbox — Người dùng-Agent reduction — Điều gì nhận frozen/coarsened trong Chrome UA string, và Vì sao.
- Chrome cho Nhà phát triển — Người dùng-Agent Client Hints — thấp- so với. cao-entropy hints và
Accept-CHopt-trong.
Quotes từ nguồn
On—record statements từ Google, Bing, và RFC. mỗi link là deep link đó jumps để quoted passage on nguồn trang.
Google — Cách các crawler identify themselves
- “Google’s crawlers identify themselves through three things: the HTTP
user-agentrequest header, the source IP address of the request, and the reverse DNS hostname of the source IP.” (bản dịch) «Google các crawler identify themselves qua three điều: đó HTTP header yêu cầu, đó nguồn IP address of đó yêu cầu, và đó reverse DNS hostname of đó nguồn IP.» — Google Search Central tài liệu. Nhảy đến trích dẫn
Google — string là spoofed
- “The HTTP user-agent request header used by Googlebot is often spoofed by other crawlers.” (bản dịch) «Đó HTTP người dùng-agent header yêu cầu dùng by Googlebot là thường spoofed by other các crawler.» — Google Search Central tài liệu. Nhảy đến trích dẫn
Google — verifying by DNS
- “Run a reverse DNS lookup on the accessing IP address from your logs, using the
hostcommand. Verify that the domain name is eithergooglebot.com,google.com, orgoogleusercontent.com.” (bản dịch) «Chạy một reverse DNS lookup on đó accessing IP address từ của bạn logs, dùng đó command. Verify đó domain name là either , , hoặc .» Nhảy đến trích dẫn - “Google doesn’t post a public list of IP addresses for website owners to allowlist because these IP address ranges can change.” (bản dịch) «Google không post một công khai list of IP addresses cho website owners để allowlist vì những IP address ranges có thể thay đổi.» Nhảy đến trích dẫn
Google — robots.txt token matching
- “All non-matching text is ignored (for example, both
googlebot/1.2andgooglebot*are equivalent togooglebot).” (bản dịch) «All non-matching text là đã bỏ qua (ví dụ, cả hai và là tương đương để ).» — Google robots.txt spec. Nhảy đến trích dẫn
Google — Web Bot Auth
- “An experimental cryptographic protocol used to authenticate requests sent by bots.” (bản dịch) «An experimental cryptographic giao thức được dùng để authenticate các yêu cầu đã gửi by bots.» Nhảy đến trích dẫn
- “We don’t sign every request of a particular agent. Be sure that you fall back to the established methods of bot verification.” (bản dịch) «We không sign mỗi yêu cầu of một particular agent. Là sure đó bạn fall lại để đó established các phương thức of bot verification.» Nhảy đến trích dẫn
Google — cloaking
- “Cloaking refers to the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users.” (bản dịch) «Cloaking refers để đó practice of presenting khác nhau nội dung để người dùng và các công cụ tìm kiếm với đó intent để manipulate tìm kiếm thứ hạng và mislead người dùng.» — Google Search Essentials, Spam Policies. Nhảy đến trích dẫn
RFC 9309 — token là substring của string
- “The product token SHOULD be a substring of the identification string that the crawler sends to the service. For example, in the case of HTTP, the product token SHOULD be a substring in the User-Agent header.” (bản dịch) «Đó sản phẩm token NÊN là một substring of đó identification string đó crawler gửi để đó service. Ví dụ, trong đó case of HTTP, đó sản phẩm token NÊN là một substring trong người dùng-Agent header.» Nhảy đến trích dẫn
Patrick Stox — on spoofing
- “Many SEO tools and some malicious bots will pretend to be Googlebot. This may allow them to access websites that try to block them.” (bản dịch) «Nhiều SEO tools và some malicious bots sẽ pretend để là Googlebot. Này có thể cho phép them để access websites đó try để block them.» — từ my Googlebot hướng dẫn on Ahrefs. Đọc điều này
Crawler → token → string → verify
reference bảng. token là Điều gì bạn put trong robots.txt; verify
hostname là Điều gì genuine yêu cầu reverse-resolves để.
| Crawler | robots.txt token | UA string contains | Verify hostname (reverse DNS) |
|---|---|---|---|
| Googlebot (Tìm kiếm) | Googlebot | Googlebot/2.1 | googlebot.com / google.com / googleusercontent.com |
| Googlebot Image | Googlebot-Image | Googlebot-Image/1.0 | giống nhau as Googlebot |
| Googlebot Video | Googlebot-Video | Googlebot-Video/1.0 | giống nhau as Googlebot |
| Google StoreBot | Storebot-Google | Storebot-Google/1.0 | giống nhau as Googlebot |
| Google-InspectionTool | Google-InspectionTool | Google-InspectionTool/1.0 | giống nhau as Googlebot |
| GoogleOther | GoogleOther | GoogleOther | varies (see Google IP files) |
| Google-Extended | Google-Extended | none — robots.txt-chỉ token | n/ (không yêu cầu string) |
| AdsBot | AdsBot-Google | AdsBot-Google | bỏ qua User-agent: * |
| AdSense | Mediapartners-Google | Mediapartners-Google | bỏ qua User-agent: * |
| Google-Safety | (bỏ qua robots.txt) | Google-Safety | bỏ qua robots.txt hoàn toàn |
| Bingbot | bingbot | bingbot/2.0 | search.msn.com |
robots.txt matching rules tại glance
| Rule | Ý nghĩ |
|---|---|
| phần lớn-cụ thể group wins | Googlebot group beats * cho Googlebot |
| giống nhau-token groups hợp nhất | Multiple Googlebot groups combine vào một |
…nhưng không bao giờ hợp nhất với * | * là chỉ fallback Khi không có gì cụ thể matches |
| Case-insensitive | Googlebot = googlebot = GOOGLEBOT |
| Version/wildcards trong token đã bỏ qua | Googlebot/1.2 và Googlebot* cả hai = Googlebot |
Fast facts
- Token = substring của UA string (RFC 9309). không toàn bộ string.
- Mobile và desktop Googlebot share một token — Bạn có thể’t split them trong robots.txt.
- Chrome version trong string là
W.X.Y.Z— nó thay đổi; không bao giờ hardcode nó. - UA string là trivially spoofed — verify by DNS hoặc IP, không bao giờ trust string.
Verify bot by reverse + forward DNS
người dùng-agent string có thể là faked trong một line của curl. xác nhận bot là genuine by
kiểm tra IP nó thực ra nghĩ ra từ. pattern là giống nhau cho Google và Bing —
chỉ dự kiến hostname differs.
Nhận right IP đầu tiên: nếu các yêu cầu truyền qua reverse proxy, load balancer, hoặc CDN trước khi hitting của bạn máy chủ, của bạn default nhật ký có thể hiển thị proxy address, không crawler. sử dụng đúng client IP (từ correctly configured forwarding header) trước khi đang chạy either kiểm tra dưới.
macOS / Linux
# --- Googlebot ---
# 1) Reverse DNS the IP from your logs — must end in googlebot.com, google.com, or googleusercontent.com
host 66.249.66.1
# → 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com
# 2) Forward DNS that hostname back — it must resolve to the same IP
host crawl-66-249-66-1.googlebot.com
# → crawl-66-249-66-1.googlebot.com has address 66.249.66.1
# --- Bingbot ---
# Reverse DNS must end in search.msn.com, then forward-confirm back to the IP
host 157.55.39.1
host <the-hostname-it-returned>Windows
:: Googlebot
nslookup 66.249.66.1
nslookup crawl-66-249-66-1.googlebot.com
:: Bingbot
nslookup 157.55.39.1
nslookup <the-hostname-it-returned>nếu reverse lookup không end trong dự kiến domain — googlebot.com /
google.com / googleusercontent.com cho Google, search.msn.com cho Bing — hoặc
forward lookup không match gốc IP, nó không phải thực bot, không quan trọng Điều gì
người dùng-agent string nói.
Match so với published IP ranges (tại quy mô)
cho verifying lots của hits, skip theo-yêu cầu DNS và match IP so với engine published CIDR ranges thay vì. Google publishes machine-readable JSON:
https://developers.google.com/static/crawling/ipranges/common-crawlers.json
https://developers.google.com/static/crawling/ipranges/special-crawlers.json
https://developers.google.com/static/crawling/ipranges/user-triggered-fetchers.jsonPull đó file, xây dựng đó CIDR set, và kiểm thử mỗi logged IP cho membership. (I được xây dựng một Googlebot IP verification tool đó làm này cho bạn.) Bing publishes của nó ranges cũng, và Bing Quản trị viên web Tools có một được xây dựng-trong “Verify Bingbot” (bản dịch) «Verify Bingbot» kiểm tra.
Người dùng-agent sanity checklist
trước khi bạn ghi UA rules hoặc act on bot trong của bạn nhật ký:
- bạn là targeting đó token trong
robots.txt(e.g.Googlebot), không pasting đó đầy đủ UA string. - Không version numbers hoặc
*bên trong đóUser-agent:line — họ là đã bỏ qua (Googlebot*làm không có gì). - Bạn haven’t assumed
User-agent: *chặn AdsBot hoặc Google-Safety — điều này không; name them explicitly nếu bạn cần để. - Nếu bạn blocked
Google-Extended, bạn understand điều này chỉ ảnh hưởng AI-training dùng — Googlebot vẫn crawl và indexes cho Tìm kiếm. - bạn là không keying access control hoặc nội dung on đó thô UA string — đây là spoofable.
- Bất kỳ “is this really Googlebot?” (bản dịch) «là này thực sự Googlebot?» kiểm tra goes qua reverse + forward DNS
(
googlebot.com/google.com/googleusercontent.com) hoặc đó published IP ranges — Bing quasearch.msn.com. - Không version-pinned UA matching anywhere — đó Chrome
W.X.Y.Zportion thay đổi. - bạn là không serving khác nhau nội dung để một bot UA hơn để người dùng (đó là cloaking).
Tools cho hoạt động với người dùng agents
Three của my own free tools cover three tasks mọi người thực ra come để điều này topic cho: confirming claimed bot là thực, seeing mà người dùng agents là thực ra hitting của bạn trang web, và kiểm tra Điều gì AI các crawler là được phép để làm.
Googlebot Verifier — tool cho chính xác vấn đề điều này bài viết giữ coming lại để: người dùng-agent string là chỉ text, so Bạn có thể’t trust nó on của nó own. Paste trong IP address từ của bạn nhật ký, pick mà crawler nó claims để là (Googlebot, Bingbot, và others), và nó chạy reverse + forward DNS kiểm tra và published-IP-range match cho bạn, sau đó trả về tiered verdict — reverse-DNS confirmed, trong published ranges, spoofed, unverifiable, hoặc không known crawler. Đã nhận toàn bộ day của hits để kiểm tra thay vì một IP? Paste lên để 500 IPs hoặc thô log lines vào bulk box.
Log File Trình phân tích — cho đó câu hỏi “which user agents are actually crawling my site?” (bản dịch) «mà người dùng agents là thực ra crawling my site?» Drop trong một máy chủ access log (nginx, Apache, IIS/W3C, hoặc JSON) và điều này parses điều này hoàn toàn trong trình duyệt của bạn, breaking crawl activity xuống by bot và by section, flagging status-code waste, splitting AI các crawler từ tìm kiếm các crawler, và — đó part đó matters hầu hết cho này topic — đang chạy một spoofer báo cáo đó names các yêu cầu claiming để là một known crawler người dùng-agent token không có đó IP để lại điều này lên.
AI-Crawler Access Checker — cho newer
family của tokens đó không behave như Googlebot hoặc bingbot. Enter URL và
nó kiểm tra của bạn robots.txt so với mỗi major AI-crawler người dùng-agent token
(GPTBot, ClaudeBot, PerplexityBot, Google-Extended, và nhiều hơn), hiển thị chính xác
rule đó wins cho mỗi, và flags liệu llms.txt tồn tại. hữu ích cho
confirming token như Google-Extended là đang làm Điều gì bạn think nó đang làm —
since, as covered trên, nó có không yêu cầu string của nó own để spot trong của bạn
nhật ký.
Prompts cho người dùng-agent tasks
Hai prompts được xây dựng khoảng đó cụ thể traps này topic sets — spoofing và
robots.txt token syntax — không generic “audit my SEO” (bản dịch) «kiểm tra SEO của tôi» filler. Paste của bạn own
dữ liệu vào đó placeholders.
Prompt 1 — triage batch của người dùng-agent strings từ của bạn nhật ký cho signs của spoofing
Paste cột của thô người dùng-agent strings pulled từ của bạn access log (không IPs — điều này prompt có thể’t verify identity, chỉ spot inconsistencies trong string itself):
Here is a list of raw User-Agent strings from my server access log, one per
line. For each one:
1. Say which crawler token it claims to be (e.g. Googlebot, bingbot,
GPTBot), or "no recognizable token" if none.
2. Flag anything internally inconsistent for that claimed crawler — e.g. a
claimed Googlebot string missing "compatible; Googlebot" or the
"+http://www.google.com/bot.html" URL, a claimed bingbot string missing
"bingbot/2.0", or a Chrome version that looks hand-typed rather than a
real evergreen build.
3. Remind me that this is a text-pattern check only — it cannot confirm
identity. Real verification requires reverse+forward DNS or matching
against the crawler's published IP ranges.
[paste user-agent strings here]Prompt 2 — kiểm tra robots.txt cho token-matching mistakes
Paste của bạn đầy đủ robots.txt file:
Review this robots.txt file for user-agent token mistakes:
1. Flag any User-agent line that includes a version number or a wildcard
inside the token (e.g. "Googlebot/1.2" or "Googlebot*") — these are
ignored, not matched as a family.
2. Check whether User-agent: * is being relied on to block AdsBot-Google,
AdsBot-Google-Mobile, Mediapartners-Google, or Google-Safety — these
ignore the wildcard and need their own named group if I want them
blocked.
3. Note any duplicate groups for the same token that could be merged, and
confirm token matching here is case-insensitive so I don't need
near-duplicate groups for casing variants.
4. List which named groups exist and which of Google's/Bing's common
crawler tokens (Googlebot, Googlebot-Image, Google-Extended, bingbot)
have no explicit group at all, so I know they're falling through to *.
[paste robots.txt here] Tự kiểm tra: người dùng agent
Five nhanh các câu hỏi on header, string, token, và Cách verify bot là thực. Pick câu trả lời cho mỗi, sau đó kiểm tra.
các tài nguyên worth của bạn time
My related writing
- Điều gì là Googlebot & Cách Làm nó Hoạt động? — đầy đủ Googlebot UA strings, verification các phương thức, và my IP verification tool.
- được lập chỉ mục, though Bị chặn bởi robots.txt — nơi UA-based blocking và robots.txt blocking collide.
- Robots.txt và SEO: Mọi thứ bạn cần Know — Cách người dùng-agent groups và rules thực ra hoạt động.
- Đáp ứng New Web Các crawler: AI Bots là Closing trong on công cụ tìm kiếm Bots — thay đổi cast của người dùng agents trong của bạn nhật ký.
Chính thức / các tiêu chuẩn
- RFC 9110 — HTTP Semantics, §10.1.5 Người dùng-Agent — base definition: tùy chọn, client-supplied metadata.
- RFC 9309: Robots Exclusion Giao thức — formal
SHOULD-cấp độ sản phẩm-token-as-substring definition. - Google Overview của các crawler và fetchers và Verify Google các crawler.
- Chrome Quyền riêng tư Sandbox — Người dùng-Agent reduction và Người dùng-Agent Client Hints — Vì sao trình duyệt UA strings là getting harder để parse, và Điều gì replaces them.
từ others
- John Mueller — Bots đó impersonate Googlebot — on spoofing và Vì sao reverse DNS là câu trả lời.
- MDN — Người dùng-Agent header — HTTP spec view của header.
- r/TechSEO — community cho crawl/log gỡ lỗi.
- Web Bot Auth: Google new experimental phương thức để validate authentic bots (công cụ tìm kiếm Land, Barry Schwartz, có thể 2026) — best news-desk summary của Cách cryptographic bot signing hoạt động và Ý nghĩ trên thực tế.
- Google-Agent người dùng agent identifies AI agent traffic trong máy chủ nhật ký (công cụ tìm kiếm Land) — covers new người dùng-triggered fetcher đó bỏ qua robots.txt và dùng Web Bot Auth.
- Google là Kiểm thử New Bot Authorization tiêu chuẩn (công cụ tìm kiếm Journal) — rộng hơn ngành context on IETF tiêu chuẩn và mà companies (Amazon, Cloudflare, Akamai, OpenAI) là backing nó.
- Announcing tương lai người dùng-agents cho Bingbot (Bing Quản trị viên web Blog, Fabrice Canel, Dec 2019) — gốc announcement của Bingbot’s shift để Edge-based kết xuất trước khi 2022 rollout.
- Microsoft list của Bingbot IP addresses đã phát hành (công cụ tìm kiếm Land) — coverage của Bing decision để publish IP ranges cho tại-quy mô bot verification.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 18 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.