Người dùng Agent

Điều gì một người dùng agent là — đó HTTP header các crawler và các trình duyệt dùng để identify themselves, đó robots.txt token so với. đó đầy đủ string, và cách verify một bot là real.

Xuất bản lần đầu: 24 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
2 tín hiệu bằng chứng trên trang này

MỘT người dùng agent là đó HTTP header mỗi client — trình duyệt, crawler, hoặc bot — gửi để identify itself. Hai điều nhận confused: đó đầy đủ người dùng-agent *string* trong đó header yêu cầu, và đó ngắn người dùng-agent *token* (Googlebot, bingbot, Google-Extended) bạn đích trong robots.txt. Đó token là một substring of đó string (RFC 9309); some tokens, như Google-Extended, có không yêu cầu string tại all. Đó string là trivially spoofed — Google says của nó own là 'thường spoofed' — so không bao giờ trust điều này cho access control. Verify Googlebot/Bingbot by reverse DNS plus một forward lookup, hoặc so với published IP ranges. Và watch đó gotchas: AdsBot và Google-Safety bỏ qua `User-agent: *`, version numbers và wildcards trong đó token line là đã bỏ qua, matching là case-insensitive, và serving khác nhau nội dung để một bot UA hơn để người dùng là cloaking.

TL;DR — MỘT người dùng agent là đó HTTP header yêu cầu bất kỳ client gửi để identify itself; đây là tùy chọn, client-filled metadata, không authenticated identity. Của nó giá trị là người dùng-agent string. Tách biệt từ đó là người dùng-agent token (sản phẩm token) dùng trong robots.txt — RFC 9309 says điều này NÊN là một substring of đó string, một mạnh convention với được ghi lại exceptions (Google-Extended có không yêu cầu string tại all). Matching là case-insensitive, version numbers/wildcards trong đó token line là đã bỏ qua, đó hầu hết-cụ thể group wins, và giống nhau-token groups hợp nhất nhưng không bao giờ hợp nhất với *. Đó string là trivially spoofed — Google calls của nó own “often spoofed” (bản dịch) «thường spoofed» — so verify by reverse + forward DNS (behind bất kỳ proxy/CDN, dùng đó client thật IP) so với googlebot.com/google.com/googleusercontent.com cho Google hoặc search.msn.com cho Bing, hoặc match published IP ranges — và ngay cả một verified yêu cầu chỉ proves một yêu cầu arrived, không đó trang đã là được lập chỉ mục, retrieved, hoặc dùng cho AI training. AdsBot và Google-Safety bỏ qua User-agent: *. Chrome là cũng freezing detail out of trình duyệt UA strings (Người dùng-Agent reduction); Client Hints là đó structured nhưng opt-trong replacement, và neither substitutes cho crawler verification. Người dùng-agent adaptation có thể là legitimate, nhưng deceptively cho thấy các crawler materially khác nhau nội dung có thể là cloaking.

Evidence for this claim HTTP User-Agent is a request field containing product information supplied by the client; it is descriptive text and not proof of identity. Scope: HTTP semantics for User-Agent. Confidence: high · Verified: IETF RFC 9110: User-Agent Evidence for this claim robots.txt User-agent matching is defined by the Robots Exclusion Protocol and controls crawler access, not authentication or general HTTP content negotiation. Scope: RFC 9309 robots matching behavior. Confidence: high · Verified: IETF RFC 9309: Robots Exclusion Protocol

header, string, và token

Three điều, và giữ them straight là phần lớn của điều này topic.

  • ** header.** User-Agent là HTTP yêu cầu header. mỗi client gửi nó: của bạn trình duyệt, curl, crawler, bot. Theo RFC 9110 ( cốt lõi HTTP semantics tiêu chuẩn), nó tùy chọn trường client fills trong — client-supplied descriptive metadata, không authenticated identity máy chủ có verified.
  • ** string.** header giá trị — freeform line describing software, version, kết xuất engine, và đôi khi OS.
  • ** token.** ngắn identifier được sử dụng trong robots.txt User-agent: lines để đích crawler — Googlebot, bingbot, Google-Extended.

Đó mối quan hệ là đó part đó trips mọi người lên. RFC 9309 (đó formal Robots Exclusion Giao thức tiêu chuẩn) says đó token “SHOULD be a substring of the identification string that the crawler sends… in the case of HTTP, the product token SHOULD be a substring in the User-Agent header.” (bản dịch) «NÊN là một substring of đó identification string đó crawler gửi… trong đó case of HTTP, đó sản phẩm token NÊN là một substring trong người dùng-Agent header.» đó là một SHOULD, không một MUST — một mạnh convention đó tiêu chuẩn khuyến nghị, không một hard requirement mỗi crawler là mechanically bound để. Google-Extended (dưới) là đó clearest ví dụ of một được ghi lại exception để điều này. không đọc đó substring rule as universal chỉ vì Google follows điều này cho hầu hết of của nó own tokens. Đó token là part of đó string khi một provider làm supply một; bạn đích đó token trong robots.txt và đọc đó string trong của bạn logs.

Evidence for this claim A robots.txt user-agent line selects a crawler product token, not an arbitrary full HTTP User-Agent string; RFC 9309 says the token should be a substring of the identification string, but this SHOULD-level convention has documented product-specific exceptions and is not authentication. Scope: robots.txt parsing and matching Confidence: high · Verified: Robots Exclusion Protocol

Google own cách diễn đạt of cách của nó bots identify themselves là hữu ích ở đây: “Google’s crawlers identify themselves through three things: the HTTP user-agent request header, the source IP address of the request, and the reverse DNS hostname of the source IP.” (bản dịch) «Google các crawler identify themselves qua three điều: đó HTTP header yêu cầu, đó nguồn IP address of đó yêu cầu, và đó reverse DNS hostname of đó nguồn IP.» Note đó người dùng-agent là chỉ một of đó three — đó other hai là cách bạn thực ra verify điều này.

Google-Extended: token với không string

Đó sạch nhất illustration of token ≠ string là Google-Extended. Điều này controls liệu Google có thể dùng nội dung của bạn cho Gemini training và grounding — và điều này có không dedicated HTTP yêu cầu người dùng-agent string tại all. Đó crawling itself là đã xong với existing Googlebot strings; Google-Extended tồn tại chỉ as một robots.txt control token. Bạn’ll không bao giờ see “Google-Extended” (bản dịch) «Google-Extended» trong một header yêu cầu trong của bạn logs.

practical consequence: blocking Google-Extended ảnh hưởng chỉ AI-training sử dụng của bạn nội dung — nó làm không dừng Googlebot từ crawling và lập chỉ mục bạn cho Tìm kiếm. họ’re tách biệt decisions controlled by tách biệt tokens. (cho rộng hơn picture của bots reading của bạn trang web, see AI các crawler và crawler.)

Googlebot’s người dùng-agent strings

Googlebot là “evergreen” — nó chạy on gần đây version của Chrome, và Chrome version trong của nó string cập nhật periodically (nó có since December 2019). Đó là lý làm version xuất hiện as W.X.Y.Z placeholder:

Googlebot Smartphone (mobile):

Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Googlebot Desktop:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36

Hai điều để internalize. đầu tiên, không hardcode versionW.X.Y.Z thay đổi, và matching on nó sẽ break. Match ổn định token Googlebot thay vì. thứ hai, Bạn có thể’t tách biệt mobile từ desktop trong robots.txt. Cả hai variants share một Googlebot token, so robots.txt rule áp dụng để cả hai.

Google crawler tokens

Google chạy toàn bộ family của các crawler và fetchers, mỗi với của nó own token. ones bạn’ll đáp ứng phần lớn:

Crawlerrobots.txt tokenNotes
GooglebotGooglebotTìm kiếm, Images, Video, News, Discover — mobile + desktop share điều này token
Googlebot ImageGooglebot-ImageGoogle Images
Googlebot VideoGooglebot-VideoVideo Tìm kiếm
Googlebot NewsGooglebot-NewsDùng various Googlebot strings
Google StoreBotStorebot-GoogleShopping
Google-InspectionToolGoogle-InspectionToolPowers Tìm kiếm kiểm thử tools
GoogleOtherGoogleOtherInternal research/fetching
Google-ExtendedGoogle-Extendedrobots.txt-chỉ — Gemini training, không yêu cầu string

và ones đó break thông thường rules — special-case các crawler đó bỏ qua User-agent: *:

  • AdsBot (AdsBot-Google) và AdsBot Mobile (AdsBot-Google-Mobile) — họ không obey wildcard. để block them bạn phải name them explicitly.
  • AdSense (Mediapartners-Google) — giống nhau; bỏ qua global *.
  • Google-Safety — được sử dụng cho malware/abuse detection; nó bỏ qua robots.txt hoàn toàn.

Đó implication là đó một mọi người miss: User-agent: * không block AdsBot hoặc Google-Safety. Nếu bạn “block all bots” (bản dịch) «block all bots» với một wildcard và assume AdsBot là đã biến mất, điều này không. (Này là chính xác đó kind of surprise đó lands một trang trong được lập chỉ mục though Bị chặn bởi robots.txt territory — see robots.txt cho đó đầy đủ control story.)

Bingbot’s người dùng-agent strings

Bing rebuilt Bingbot’s string trong 2022 để reflect đó nó renders với Microsoft Edge. hiện tại strings:

Bingbot Desktop:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36

Bingbot Mobile:

Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)

robots.txt token là chỉ bingbot. điều để watch: post-2022, Bingbot’s string looks gần như chính xác như thực Chrome/Edge trình duyệt — chỉ tell là bingbot/2.0 fragment bên trong nó. nếu bạn có bất kỳ logic đó filters hoặc detects bots by UA, đó thay đổi matters.

Cách robots.txt thực ra matches token

một vài rules govern mà group của rules crawler obeys (theo Google robots.txt spec và RFC 9309):

  • Hầu hết-cụ thể match wins. Google “determines the correct group of rules by finding… the group with the most specific user agent that matches the crawler’s user agent.” (bản dịch) «determines đó correct group of rules by finding… đó group với đó hầu hết cụ thể người dùng agent đó matches đó crawler người dùng agent.» MỘT Googlebot group beats một * group cho Googlebot.
  • Giống nhau-token groups hợp nhất — nhưng không bao giờ với *. Multiple groups naming đó giống nhau agent là combined vào một. MỘT cụ thể-agent group và đó * group là không đã hợp nhất; * là chỉ đó fallback khi không có gì cụ thể matches.
  • Case-insensitive. Trường name và giá trị cả hai — Googlebot, googlebot, GOOGLEBOT là tương đương.
  • Version numbers và wildcards trong đó token line là đã bỏ qua. Theo Google, “both googlebot/1.2 and googlebot* are equivalent to googlebot.” (bản dịch) «cả hai và là tương đương để .» Bạn không thể ghi User-agent: Googlebot* để match một family — đó * ở đó làm không có gì.

So User-agent: line takes token và matches nó as đơn giản (case-insensitive) substring của crawler identity — không version pinning, không wildcards bên trong nó.

Vì sao Bạn có thể’t trust string — và Cách verify

Người dùng-agent string là freeform text. Bất cứ điều gì có thể set điều này. Một line of curl sẽ claim để là Googlebot, và plenty of tools và malicious bots làm chính xác đó để slip past chặn. Google says so trong của nó own Googlebot tài liệu: “the HTTP user-agent request header used by Googlebot is often spoofed by other crawlers.” (bản dịch) «đó HTTP người dùng-agent header yêu cầu dùng by Googlebot là thường spoofed by other các crawler.» As I’ve put điều này trong my Googlebot hướng dẫn, “Many SEO tools and some malicious bots will pretend to be Googlebot. This may allow them to access websites that try to block them.” (bản dịch) «Nhiều SEO tools và some malicious bots sẽ pretend để là Googlebot. Này có thể cho phép them để access websites đó try để block them.»

So không bao giờ làm access hoặc nội dung decision on string alone. Verify thay vì.

Một prerequisite trước khi either phương thức: nhận thực nguồn IP. nếu của bạn trang web sits behind reverse proxy, load balancer, hoặc CDN, address trong của bạn default access log có thể là proxy IP, không crawler — bạn cần gốc client IP (thường forwarded trong header như X-Forwarded-For, configured correctly tại của bạn proxy) hoặc neither verification phương thức dưới có nghĩ là bất cứ điều gì.

Phương thức 1 — reverse + forward DNS (best cho spot kiểm tra). Google hai steps:

  1. “Run a reverse DNS lookup on the accessing IP address from your logs, using the host command. Verify that the domain name is either googlebot.com, google.com, or googleusercontent.com.” (bản dịch) «Chạy một reverse DNS lookup on đó accessing IP address từ của bạn logs, dùng đó command. Verify đó domain name là either , , hoặc .»
  2. “Run a forward DNS lookup on the domain name retrieved in step 1… Verify that it’s the same as the original accessing IP address from your logs.” (bản dịch) «Chạy một forward DNS lookup on đó domain name retrieved trong step 1… Verify đó đây là đó giống nhau as đó original accessing IP address từ của bạn logs.»

cho Bingbot, giống nhau hai-step dance, nhưng hostname phải end trong search.msn.com (không Bing-branded domain — phổ biến surprise). Commands là trong Scripts tab.

Phương thức 2 — published IP ranges (best tại quy mô). Google không publish một static allowlist cho hardcoding (“these IP address ranges can change” (bản dịch) «những IP address ranges có thể thay đổi»), nhưng điều này làm publish machine-readable CIDR JSON files bạn có thể match so với (phổ biến-các crawler.json và đó rộng hơn crawler files). Bing hiện tại publishes của nó ranges cũng. I được xây dựng một Googlebot IP verification tool cho chính xác này — paste trong IPs và điều này classifies them. Bing Quản trị viên web Tools có một được xây dựng-trong “Verify Bingbot” (bản dịch) «Verify Bingbot» tool as well.

DNS là tốt hơn cho một-off log kiểm tra; IP-range matching là tốt hơn cho verifying tại volume. sử dụng whichever fits — nhưng sử dụng một của them. và treat cả hai dự kiến hostnames và range files as hiện tại as của hôm nay, không vĩnh viễn — Google và Bing có changed những điều này paths trước khi ( IP-range JSON files moved và là renamed since điều này bài viết là đầu tiên được viết), so re-kiểm tra trực tiếp verification doc nếu lookup đó được sử dụng để hoạt động dừng matching.

UA match không phải proof của downstream outcomes

Ngay cả fully verified yêu cầu — thực Googlebot IP, forward-confirmed reverse DNS, mọi thứ kiểm tra out — chỉ proves một điều: đó yêu cầu reached của bạn máy chủ. nó tempting để round đó lên vào nhiều bigger claim, nhưng mỗi của những điều này là tách biệt fact requiring tách biệt evidence:

  • yêu cầu đã nhận — yêu cầu với đó người dùng agent hit của bạn máy chủ. (Điều gì log verification thực ra proves.)
  • Identity confirmed — yêu cầu thực sự nghĩ ra từ crawler nó claims để là. (Điều gì reverse DNS / IP-range matching adds on top.)
  • nội dung fetched và được kết xuất — crawler successfully được kết xuất trang (không các lỗi, không blocked các tài nguyên). không guaranteed chỉ vì yêu cầu landed.
  • được lập chỉ mục — URL đã làm nó vào tìm kiếm chỉ mục. thành công fetch không bảo đảm lập chỉ mục.
  • được sử dụng cho retrieval, citation, hoặc training — cho AI các crawler especially (Google-Extended, GPTBot, và rest), crawl không phải proof của bạn nội dung là retrieved cho cụ thể câu trả lời, cited, hoặc được sử dụng trong model training. những điều đó là tách biệt, mostly unobservable steps downstream của crawl.

verified Googlebot hit trong của bạn nhật ký là thực tín hiệu — chỉ không stretch nó further hơn Điều gì nó thực ra hiển thị.

Web Bot Auth: nơi verification là heading

Trong 2026 Google began experimenting với Web Bot Auth“an experimental cryptographic protocol used to authenticate requests sent by bots.” (bản dịch) «an experimental cryptographic giao thức được dùng để authenticate các yêu cầu đã gửi by bots.» Đó ý tưởng là để “move beyond easily spoofed headers to a verified identity and decouple agent identity from IP addresses.” (bản dịch) «move beyond easily spoofed các header để một verified identity và decouple agent identity từ IP addresses.» Bots cryptographically sign của họ các yêu cầu; các trang verify đó signature so với Google published công khai keys, và signed các yêu cầu carry một Signature-Agent header. Google own caveat matters: “We don’t sign every request of a particular agent. Be sure that you fall back to the established methods of bot verification.” (bản dịch) «We không sign mỗi yêu cầu of một particular agent. Là sure đó bạn fall lại để đó established các phương thức of bot verification.» So đây là additive, không một replacement — reverse DNS và IP ranges vẫn của bạn baseline hôm nay.

các trình duyệt là getting harder để parse từ UA string cũng

Mọi thứ trên là về các crawler, nhưng đó giống nhau “don’t over-trust the string” (bản dịch) «không over-trust đó string» lesson áp dụng để các trình duyệt, và đây là getting stronger. Chrome đã được rolling out Người dùng-Agent reduction: freezing hoặc coarsening parts of của nó UA string (đầy đủ trình duyệt version, OS version, device model) thay vì reporting them chính xác, so đó string không thể là được dùng để fingerprint một cụ thể người dùng. Google own cách diễn đạt: “The granularity and abundance of detail can lead to user identification. The default availability of this information can lead to covert tracking.” (bản dịch) «Đó granularity và abundance of detail có thể lead để người dùng identification. Đó default availability of này information có thể lead để covert tracking.» Practically, đó có nghĩa là UA-string phân tích cú pháp cho chính xác trình duyệt/OS/device version — analytics, device detection, bug triage — là increasingly unreliable và sẽ chỉ nhận hơn so.

Đó replacement Chrome khuyến nghị là Người dùng-Agent Client Hints (UA-CH): structured dữ liệu đó trình duyệt gửi chỉ khi một máy chủ explicitly asks cho điều này. Thấp-entropy hints (trình duyệt brand, major version, mobile flag) go out theo mặc định; cao-entropy hints (chính xác version, nền tảng version, device model) require đó máy chủ để opt trong qua an Accept-CH header phản hồi đầu tiên — an rõ ràng negotiation, không một broadcast. Hai caveats trước khi bạn lean on điều này: đây là một Chrome/Chromium-family mechanism, không điều gì đó mỗi trình duyệt gửi, và ngay cả nơi đây là supported, “the value may be blank, not returned, or populated with a varying value.” (bản dịch) «đó giá trị có thể là blank, không đã trả về, hoặc populated với một varying giá trị.» Client Hints solve đó trình duyệt-string vấn đề; they không phải một crawler-verification mechanism — Google và Bing vẫn verify của họ own các crawler qua DNS và IP ranges, không Client Hints.

Người dùng-agent targeting và cloaking

Đó tempting move — “detect Googlebot by its UA and serve it something special” (bản dịch) «detect Googlebot by của nó UA và serve điều này điều gì đó special» — là cả hai technically fragile và một policy violation.

Fragile, vì Google không crawl với một UA. bạn’d có để correctly xử lý Googlebot (mobile và desktop), Google-InspectionTool, AdsBot, GoogleOther, và nhiều hơn, từ rotating IPs — practically không thể để whitelist cleanly.

MỘT policy violation, vì serving khác nhau nội dung để một crawler hơn để người dùng là cloaking: “presenting different content to users and search engines with the intent to manipulate search rankings and mislead users.” (bản dịch) «presenting khác nhau nội dung để người dùng và các công cụ tìm kiếm với đó intent để manipulate tìm kiếm thứ hạng và mislead người dùng.» Đó hình phạt ranges từ algorithmic demotion để đầy đủ deindexing. Note đó line: legitimate adaptation (responsive layouts, nội dung negotiation) là fine — đây là swapping đó nội dung itself giữa bots và người dùng đó crosses vào cloaking.

cho nơi người dùng agent sits trong bigger pipeline, see crawling ( hub) và crawler. cho controlling Điều gì những điều đó bots là được phép để fetch, see robots.txt.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.