Hướng dẫn về Robots.txt

Điều gì robots.txt thực ra làm — điều này controls crawling, không lập chỉ mục — plus đó chính xác syntax, cách Google xử lý điều này dưới đó hood, và đó mistakes đó break các trang.

Xuất bản lần đầu: 23 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Robots.txt là một đơn giản-text file tại đó root of mỗi host đó tells các crawler mà URLs they có thể và có thể không yêu cầu. Đó một điều để nhận right: điều này controls crawling, không lập chỉ mục. MỘT disallowed URL có thể vẫn là được lập chỉ mục không có một snippet nếu đây là linked từ elsewhere — để giữ một trang out of đó chỉ mục bạn dùng noindex, và đó trang không được là blocked trong robots.txt hoặc Google không bao giờ sees đó noindex. Google hỗ trợ chỉ người dùng-agent, cho phép, disallow, và sitemap (noindex, nofollow, và crawl-delay đã là dropped Sept 1, 2019). Điều này lives tại /robots.txt, là scoped để một host+giao thức+port, caps tại 500 KiB, caches ~24h, và một 4xx có nghĩa là không restrictions trong khi một 5xx có thể stall crawling site-wide. không block render-cốt yếu CSS/JS, và không rely on điều này để hide bất cứ điều gì — đó file là công khai.

Tóm tắt — Robots.txt là đơn giản-text file tại root của mỗi host (/robots.txt, lowercase) đó implements Robots Exclusion Giao thức (RFC 9309). nó controls crawling, không lập chỉ mục — disallowed URL có thể vẫn là được lập chỉ mục không có snippet nếu linked elsewhere; để deindex sử dụng noindex on trang Đó là không blocked. Google hỗ trợ chỉ user-agent, allow, disallow, và sitemap; noindex/nofollow/crawl-delay là dropped Sept 1, 2019. Phạm vi là một host+giao thức+port. Matching dùng phần lớn-cụ thể (longest) rule, least-restrictive on ties; *$ là wildcards; paths là case-sensitive. Google caps file tại 500 KiB, caches ~24h, xử lý 4xx (except 429) as không-restrictions, và on 5xx stalls crawling cho ~12h sau đó falls lại để cuối cùng good copy cho ~30 days. không block render-cốt yếu CSS/JS, và không treat nó as access control — file là công khai.

Điều gì nó là và nơi nó lives

Robots.txt implements đó Robots Exclusion Giao thức, đã tạo by Martijn Koster trong 1994 và finally standardized trong 2022 as RFC 9309 — co-authored by Google Gary Illyes, Henner Zeller, Lizzi Sassman, và Koster himself. Đó tiêu chuẩn own wording: “This document specifies and extends the ‘Robots Exclusion Protocol’ method originally defined by Martijn Koster in 1994 for service owners to control how content served by their services may be accessed, if at all, by automatic clients known as crawlers.” (bản dịch) «Này document specifies và extends đó ‘Robots Exclusion Giao thức’ phương thức originally được định nghĩa by Martijn Koster trong 1994 cho service owners để control cách nội dung phân phối by của họ services có thể là accessed, nếu tại all, by tự động clients known as các crawler.»

một vài facts đó catch mọi người out:

  • Điều này phải được tại đó root, lowercase. RFC 9309 là rõ ràng: “The rules MUST be accessible in a file named ‘/robots.txt’ (all lowercase) in the top-level path of the service.” (bản dịch) «Đó rules Phải được accessible trong một file named ‘/robots.txt’ (all lowercase) trong đó top-cấp độ path of đó service.» Google adds đó URL itself là case-sensitive, như bất kỳ URL.
  • Phạm vi là một host + giao thức + port. Google: “The rules listed in the robots.txt file apply only to the host, protocol, and port number where the robots.txt file is hosted.” (bản dịch) «Đó rules listed trong đó robots.txt file apply chỉ để đó host, giao thức, và port number nơi đó robots.txt file là hosted.» So https://example.com, https://www.example.com, https://blog.example.com, và http://example.com mỗi cần của họ own file. Subdomains và các giao thức không share một.
  • Supported các giao thức cho Google là HTTP, HTTPS, và FTP.

misconception đó defines điều này topic: crawling so với lập chỉ mục

nếu bạn take một điều từ điều này trang, take điều này: robots.txt controls crawling, không lập chỉ mục. Blocking URL không phải giống nhau as removing nó từ Google. Evidence for this claim A robots.txt rule controls crawling rather than guaranteeing removal from Google Search; a URL can still appear when Google cannot crawl it. Scope: Google Search crawler behavior. Other crawlers can interpret robots.txt differently. Confidence: high · Verified: Google: Introduction to robots.txt

Google own intro doc says điều này plainly: robots.txt “is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.” (bản dịch) «không phải một mechanism cho giữ một web trang out of Google. Để giữ một web trang out of Google, block lập chỉ mục với hoặc password-bảo vệ đó trang.» Và on điều gì thực ra happens để một blocked URL: “While Google won’t crawl or index the content blocked by a robots.txt file, we might still find and index a disallowed URL if it is linked from other places on the web.” (bản dịch) «Trong khi Google sẽ không crawl hoặc chỉ mục đó nội dung blocked by một robots.txt file, we có thể vẫn tìm và chỉ mục một disallowed URL nếu điều này là linked từ other places on đó web.» Đó kết quả là đó familiar snippet-ít hơn listing: “its URL can still appear in search results, but the search result won’t have a description.” (bản dịch) «của nó URL có thể vẫn xuất hiện trong kết quả tìm kiếm, nhưng đó tìm kiếm kết quả sẽ không có một mô tả.»

Evidence for this claim Robots.txt controls crawler access, not index eligibility; Google may still index a disallowed URL discovered through links, typically without a content snippet. Scope: web crawling Confidence: high · Verified: Robots.txt Introduction and Guide

Đó spec restates đó giống nhau nuance cho đó disallow rule itself: “Google can’t index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” (bản dịch) «Google không thể chỉ mục đó nội dung of các trang mà là disallowed cho crawling, nhưng điều này có thể vẫn chỉ mục đó URL và cho thấy điều này trong kết quả tìm kiếm không có một snippet.»

Vì sao bạn phải không block trang bạn muốn để noindex

Này là đó trap đó âm thầm breaks deindexing efforts. MỘT noindex chỉ hoạt động nếu Google có thể crawl đó trang để đọc điều này. Google block-lập chỉ mục doc spells out đó dependency: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.” (bản dịch) «Cho đó rule để là effective, đó trang hoặc tài nguyên không được là blocked by một robots.txt file, và điều này có để là nếu không accessible để đó crawler. Nếu đó trang là blocked by một robots.txt file hoặc đó crawler không thể access đó trang, đó crawler sẽ không bao giờ see đó rule, và đó trang có thể vẫn xuất hiện trong kết quả tìm kiếm, ví dụ nếu other các trang link để điều này.» Evidence for this claim Google must be able to crawl a URL to see a noindex rule; blocking the URL in robots.txt can prevent the rule from being observed. Scope: Google Search indexing controls for HTML meta robots and X-Robots-Tag rules. Confidence: high · Verified: Google: Block indexing with noindex

So nếu của bạn goal là để nhận trang out của chỉ mục, John Mueller hướng dẫn là sạch nhất way để remember nó: Khi bạn muốn để unindex các trang, bạn nên không block Google với robots.txt, nhưng rather sử dụng noindex.

lived proof: I blocked hai của chúng ta own cao-xếp hạng các trang

I không có để argue này từ theory. Trong my thử nghiệm blocking hai cao-xếp hạng Ahrefs các trang, I có chủ ý blocked them trong robots.txt và tracked điều gì happened. Đó các trang stayed được lập chỉ mục và kept xếp hạng — they đã không vanish. Điều gì we lost đã là đó freshness Google nhận từ re-crawling: “We lost a position here or there and all of the featured snippets for the pages.” (bản dịch) «We lost một position ở đây hoặc ở đó và all of đó featured snippets cho đó các trang.» Traffic dropped, nhưng ít hơn I dự kiến: “Both pages lost some traffic. But it didn’t result in much change to our traffic estimate like I was expecting.” (bản dịch) «Cả hai các trang lost some traffic. Nhưng điều này đã không kết quả trong nhiều thay đổi để của chúng ta traffic estimate như I đã là expecting.»

My takeaway từ đó dữ liệu: “Accidentally blocking pages (that Google already ranks) from being crawled using robots.txt probably isn’t going to have much impact on your rankings, and they will likely still show in the search results.” (bản dịch) «Accidentally blocking các trang (đó Google đã ranks) từ đang được crawl dùng robots.txt probably không going để có nhiều impact on của bạn thứ hạng, và they sẽ có khả năng vẫn cho thấy trong đó kết quả tìm kiếm.» Và đó blunt version: “Don’t block pages you want indexed. It hurts. Not as bad as you might think it does—but it still hurts.” (bản dịch) «không block các trang bạn muốn được lập chỉ mục. Điều này hurts. Không as bad as bạn có thể think điều này làm—nhưng điều này vẫn hurts.»

Đó flip side là reassurance: khi Search Console flags “Indexed, though blocked by robots.txt” (bản dịch) «Được lập chỉ mục, though Bị chặn bởi robots.txt» cho một utility URL — cart, filter, parameter junk — đây là thường một non-vấn đề. As Mueller put điều này về thêm-để-cart URLs, blocking them là fine, và ngay cả khi they nhận “được lập chỉ mục,” đây là khó có khả năng they’ll là shown trong tìm kiếm trừ khi ai đó chạy một very cụ thể query cho những URLs, mà real người dùng không làm. Distinguish đó scary-sounding warning từ an thực tế vấn đề: điều này chỉ matters nếu đó blocked URL là một trang bạn thực ra wanted được crawl và được lập chỉ mục.

syntax ( reference)

robots.txt là đặt của groups. mỗi group bắt đầu với một hoặc nhiều hơn User-agent lines naming mà crawler(s) nó áp dụng để, followed by rules cho them.

User-agent: *
Disallow: /cart/
Disallow: /search
Allow: /search/help

User-agent: Googlebot
Disallow: /no-google/

Sitemap: https://example.com/sitemap.xml

Người dùng-agent và groups. MỘT crawler obeys chính xác một group — đó một với đó hầu hết cụ thể người dùng-agent đó matches điều này — và bỏ qua đó rest. Google: “Google’s crawlers determine the correct group of rules by finding in the robots.txt file the group with the most specific user agent that matches the crawler’s user agent. Other groups are ignored.” (bản dịch) «Google các crawler determine đó correct group of rules by finding trong đó robots.txt file đó group với đó hầu hết cụ thể người dùng agent đó matches đó crawler người dùng agent. Other groups là đã bỏ qua.» Và: “Only one group is valid for a particular crawler.” (bản dịch) «Chỉ một group là hợp lệ cho một particular crawler.» (Bing behaves cùng cách — hơn on đó dưới.)

Đó cũng có nghĩa là một cụ thể group không nhận topped lên với đó wildcard group rules — đây là dùng on của nó own, không đã hợp nhất với User-agent: *. Google spec là rõ ràng đó “user agent specific groups and global groups (*) are not combined.” (bản dịch) «người dùng agent cụ thể groups và global groups ( ) không phải combined.» So nếu bạn ghi một User-agent: googlebot-news group, điều này có để là self-contained: bất cứ điều gì bạn vẫn muốn điều này để obey từ đó * group có để là repeated bên trong điều này, hoặc Googlebot-News đơn giản sẽ không see những rules tại all.

Evidence for this claim For Google's crawlers, the most specific matching user-agent group applies; rules from that specific group are not combined with the global asterisk group, although multiple matching specific groups are merged internally. Scope: robots.txt parsing and fetching Confidence: high · Verified: How Google Interprets the robots.txt Specification

Disallow và Cho phép. Disallow lists paths một crawler không được yêu cầu; Allow carves exceptions lại out. Đó disallow rule “specifies paths that must not be accessed by the crawlers identified by the user-agent line the disallow rule is grouped with.” (bản dịch) «specifies paths đó không được là accessed by đó các crawler identified by người dùng-agent line đó disallow rule là grouped với.» Đó allow rule “specifies paths that may be accessed by the designated crawlers. When no path is specified, the rule is ignored.” (bản dịch) «specifies paths đó có thể là accessed by đó designated các crawler. Khi không path là specified, đó rule là đã bỏ qua.»

Đó matching rule (hầu hết các hướng dẫn nhận này sai). Khi hai rules conflict, đó hầu hết cụ thể một wins, và “most specific” (bản dịch) «hầu hết cụ thể» có nghĩa là longest path: “When matching robots.txt rules to URLs, crawlers use the most specific rule based on the length of the rule path. In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.” (bản dịch) «Khi matching robots.txt rules để URLs, các crawler dùng đó hầu hết cụ thể rule dựa trên đó length of đó rule path. Trong case of conflicting rules, including những với wildcards, Google dùng đó least restrictive rule.» So on một genuine tie, đó least restrictive rule wins — Allow beats Disallow. RFC 9309 frames điều này as đó “Longest Match” (bản dịch) «Longest Match»: “The following example shows that in the case of two rules, the longest one is used for matching.” (bản dịch) «Đó sau ví dụ cho thấy đó trong đó case of hai rules, đó longest một được dùng cho matching.» Evidence for this claim Google resolves matching robots.txt rules by path specificity and uses the least restrictive rule when equally specific rules conflict. Scope: Google crawler interpretation of robots.txt rules; other crawlers may implement different extensions. Confidence: high · Verified: Google: Robots.txt interpretation

Worked ví dụ:

User-agent: *
Allow: /folder/page
Disallow: /folder/

URL /folder/page matches cả hai rules. Allow: /folder/page (12 chars) là lâu hơn Disallow: /folder/ (8 chars), so lâu hơn, nhiều hơn cụ thể Cho phép wins và trang là crawlable.

Wildcards *$. Google: * designates 0 or more instances of any valid character. $ designates the end of the URL.” (bản dịch) «designates 0 hoặc hơn instances of bất kỳ hợp lệ character. designates đó end of đó URL.» So Disallow: /*.pdf$ chặn mỗi URL ending trong .pdf, và Disallow: /*? chặn mỗi URL containing một query string. Matching là prefix-based: Disallow: /fish matches /fish, /fish.html, và /fish/salmon.html, nhưng không /Fish (case-sensitive) hoặc /catfish (đây là một prefix, không một substring).

Case sensitivity (đó subtle một). Trường và người dùng-agent names là case-insensitive; path các giá trị là case-sensitive. Google: “Both the user-agent field name and its value are case-insensitive,” (bản dịch) «Cả hai đó trường name và của nó giá trị là case-insensitive,» nhưng “The field name (disallow) is case-insensitive, but its value is case-sensitive,” (bản dịch) «Đó trường name ( ) là case-insensitive, nhưng của nó giá trị là case-sensitive,»“The path value must start with / to designate the root and the value is case-sensitive.” (bản dịch) «Đó path giá trị phải bắt đầu với để designate đó root và đó giá trị là case-sensitive.» So Disallow: /Folder/ không block /folder/.

Sitemap. Sitemap: directive takes đầy đủ absolute URL và là independent của groups — nó có thể sit anywhere trong file.

Comments. Bất cứ điều gì sau # là đã bỏ qua: “To include comments, precede your comment with the # character.” (bản dịch) «Để bao gồm comments, precede của bạn comment với đó character.»

noindex, nofollow, và crawl-delay không phải robots.txt directives

Này là một persistent myth. As of September 1, 2019, Google retired hỗ trợ cho unsupported, undocumented rules — including noindex, nofollow, và crawl-delay. Google announcement focused on rules unsupported by đó internet draft, such as crawl-delay, nofollow, và noindex, noting they đã là không bao giờ được ghi lại by Google, và đã nói Google đã là retiring all code đó xử lý unsupported và unpublished rules (such as noindex) on đó date. Đó supported trường list là ngắn, và đó spec calls out đó exclusion trực tiếp: Google hỗ trợ user-agent, allow, disallow, và sitemap, và “other fields such as crawl-delay aren’t supported.” (bản dịch) «other các trường such as không supported.»

nếu bạn relied on noindex trong robots.txt, alternatives là noindex meta tag hoặc X-Robots-Tag header, 404/410 các mã trạng thái, password protection, Disallow, hoặc Search Console removal tool.

Cách Google xử lý robots.txt dưới hood

  • Size limit: 500 KiB. “Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.” (bản dịch) «Google enforces một robots.txt file size limit of 500 kibibytes (KiB). Nội dung mà là sau đó maximum file size là đã bỏ qua.» RFC 9309 aligns: “The parsing limit MUST be at least 500 kibibytes [KiB].” (bản dịch) «Đó phân tích cú pháp limit Phải được ít nhất 500 kibibytes [KiB].»
  • Bộ nhớ đệm: ~24 hours. “Google generally caches the contents of robots.txt file for up to 24 hours, but may cache it longer in situations where refreshing the cached version isn’t possible.” (bản dịch) «Google generally caches đó nội dung of robots.txt file cho lên để 24 hours, nhưng có thể bộ nhớ đệm điều này lâu hơn trong situations nơi refreshing đó được lưu đệm version không có thể.» So một thay đổi không nhất thiết picked lên instantly. Evidence for this claim Google generally caches robots.txt for up to 24 hours and changes crawling behavior according to the HTTP status returned for the file. Scope: Google crawler handling of robots.txt fetches, including documented 4xx, 5xx, and redirect behavior. Confidence: high · Verified: Google: Robots.txt file handling
  • Các mã trạng thái quan trọng site-wide. Này là đó part hầu hết các hướng dẫn skip:
    • 4xx (except 429) → không restrictions. “Google’s crawlers treat all 4xx errors, except 429, as if a valid robots.txt file didn’t exist. This means that Google assumes that there are no crawl restrictions.” (bản dịch) «Google các crawler treat all 4xx các lỗi, except 429, as nếu một hợp lệ robots.txt file đã không exist. Này có nghĩa là đó Google assumes đó có không crawl restrictions.» MỘT 404 on /robots.txt có nghĩa là “crawl everything.” (bản dịch) «crawl mọi thứ.» (không dùng 401/403 để throttle crawling.)
    • 5xx / unreachable → dangerous. “For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can’t fetch a new version, for the next 30 days Google will use the last good version, while still trying to fetch a new version.” (bản dịch) «Cho đó đầu tiên 12 hours, Google dừng crawling đó site nhưng giữ trying để fetch đó robots.txt file. Nếu Google không thể fetch một new version, cho đó tiếp theo 30 days Google sẽ dùng đó cuối cùng good version, trong khi vẫn trying để fetch một new version.» So một máy chủ lỗi on /robots.txt có thể effectively disallow của bạn toàn bộ site cho đó đầu tiên ~12 hours, thì chạy on đó cuối cùng được lưu đệm copy cho ~30 days. MỘT persistently erroring robots.txt là một site-wide crawl risk. Và nếu đây là vẫn hỏng sau những 30 days: “If the errors are still not fixed after 30 days: If the site is generally available to Google, Google will behave as if there is no robots.txt file (but still keep checking for a new version).” (bản dịch) «Nếu đó các lỗi là vẫn không fixed sau 30 days: Nếu đó site là generally khả dụng để Google, Google sẽ behave as nếu có không robots.txt file (nhưng vẫn giữ kiểm tra cho một new version).» Nói cách khác, một robots.txt đó không bao giờ recovers không stay disallowed forever — Google eventually falls lại để crawling với không restrictions, đó giống nhau as một 404.
    • 3xx → Google follows ít nhất five chuyển hướng hops, thì xử lý điều này as một 404.

robots.txt trong Bing, Yandex, và beyond

grouping và syntax là essentially shared, nhưng hai divergences quan trọng:

  • crawl-delay. Google bỏ qua điều này, Bing vẫn honors điều này, và Yandex dropped điều này trong 2018 — Yandex own tài liệu trạng thái đó “From February 22, 2018, Yandex doesn’t take into account the Crawl-delay directive,” (bản dịch) «Từ February 22, 2018, Yandex không take vào account đó Crawl-delay directive,» pointing bạn để đó site tốc độ crawl setting trong Yandex Quản trị viên web thay vì. Bing là rõ ràng đó “The robots.txt file is the only valid place to set a crawl-delay directive for MSNBot,” (bản dịch) «Đó robots.txt file là đó chỉ hợp lệ place để set một crawl-delay directive cho MSNBot,» và đó directive “accepts only positive, whole numbers as values… the higher the value, the more throttled down the crawl rate will be.” (bản dịch) «accepts chỉ positive, toàn bộ numbers as các giá trị… đó cao hơn đó giá trị, đó hơn throttled xuống đó tốc độ crawl sẽ là.» Note Bing xử lý đó giá trị as một relative throttle, không theo nghĩa đen N seconds.
  • Đó bingbot-section gotcha. Chỉ như Google “only one group per crawler” (bản dịch) «chỉ một group theo crawler» rule, nếu bạn tạo một User-agent: bingbot section, Bing áp dụng chỉ đó section và bỏ qua đó User-agent: * defaults (crawl-delay excepted). So một bingbot-cụ thể group phải repeat mỗi directive bạn vẫn muốn enforced.
  • Amazon bộ nhớ đệm và failure behavior. Amazon says của nó các crawler có thể dùng một robots.txt copy được lưu đệm trong đó trước đó 30 days. Nếu they không thể fetch đó file, they behave as though điều này không exist. MỘT checker có thể báo cáo đó copy điều này fetched, nhưng điều này không thể prove mà được lưu đệm version Amazon dùng—hoặc đó Amazon observed đó giống nhau failure as đó checker. Evidence for this claim Amazon says its crawlers may use a robots.txt copy cached within the previous 30 days and behave as though the file does not exist when they cannot fetch it. Scope: Amazon crawler behavior only; a checker result cannot establish which cached copy Amazon used or whether Amazon observed the same fetch failure. Confidence: high · Verified: Amazon: Amazonbot

Managing AI các crawler với robots.txt

Robots.txt là hiện tại main lever cho managing AI các crawler, và họ obey giống nhau group/người dùng-agent syntax. catch: những điều này là tách biệt tokens, so blocking một không block others.

  • OpenAI chạy several distinct bots, và đó controls cho mỗi là independent — allowing một không cho phép đó others, và blocking một không block đó others. GPTBot crawl nội dung cho training OpenAI’s models; OAI-SearchBot surfaces các trang trong ChatGPT’s tìm kiếm features; OAI-AdsBot kiểm tra đó safety of các trang được gửi as quảng cáo (của nó dữ liệu không dùng cho training). Block training với User-agent: GPTBot / Disallow: / — đó alone sẽ không dừng đó tìm kiếm hoặc quảng cáo bots. ChatGPT-User là khác nhau again: điều này fires cho actions một person triggers bên trong ChatGPT hoặc một Custom GPT, không tự động crawling, và OpenAI says “robots.txt rules may not apply” (bản dịch) «robots.txt rules có thể không apply» để điều này — so không count on một Disallow để giữ điều này out. Nếu bạn làm thay đổi điều gì OAI-SearchBot có thể crawl, OpenAI notes điều này có thể take về 24 hours cho đó cập nhật để reach của họ tìm kiếm các hệ thống.
  • Google-Extended controls Gemini/Vertex training và là tách biệt từ Googlebot.
  • Others worth naming: CCBot (Phổ biến Crawl), ClaudeBot (Anthropic), PerplexityBot, và Bytespider.

hard caveat: compliance là voluntary. Robots.txt các yêu cầu; nó không enforce. Well-behaved các crawler obey nó; scrapers có thể và làm bỏ qua nó. nếu bạn truly cần để giữ điều gì đó away từ bot, đó authentication/blocking vấn đề, không robots.txt một.

phổ biến mistakes (và các cách sửa)

** file là 200, nhưng nó không phải thực ra usable robots file.** Status alone là không đủ. Capture phản hồi Content-Type và đầu tiên bytes: CDN/custom-lỗi template có thể trả về HTML tại /robots.txt với 200, mà phải là warning rather hơn “cho phép all” truyền. Google documents robots.txt as UTF-8 đơn giản text và có thể bỏ qua không hợp lệ characters. single UTF-8 BOM tại beginning là tolerated, nhưng thứ hai BOM, BOM trong middle, UTF-16 bytes, NULs, hoặc invisible/control characters có thể alter đầu tiên token hoặc invalidate line. hiển thị byte offset và affected line; không silently normalize file trước khi telling người dùng Điều gì crawler đã nhận. Apply Google effective 500 KiB phân tích cú pháp limit trước khi calculating cho phép/disallow kết quả, trong khi vẫn reporting discarded tail.

  • Blocking một trang bạn cũng muốn deindexed. Block + noindex có nghĩa là Google không bao giờ crawl điều này để see đó noindex. Dùng noindex không có đó block.
  • Dùng robots.txt để deindex. Sai tool hoàn toàn — đó là noindex job.
  • Blocking render-cốt yếu CSS/JS. Google cần những assets để see đó trang as một người dùng làm; Google own sample robots.txt explicitly re-cho phép .css/.js so Googlebot có thể crawl them.
  • Trying để hide sensitive dữ liệu. RFC 9309 là blunt: “The Robots Exclusion Protocol is not a substitute for valid content security measures. Listing paths in the robots.txt file exposes them publicly and thus makes the paths discoverable.” (bản dịch) «Đó Robots Exclusion Giao thức không phải một substitute cho hợp lệ nội dung security measures. Listing paths trong đó robots.txt file exposes them publicly và thus làm đó paths discoverable.» Disallowing /secret-admin/ theo nghĩa đen advertises điều này. Dùng auth.
  • MỘT stray Disallow: /. Này chặn đó entire site cho đó named crawler — đó classic staging leftover đó takes một site out of Google.
  • Ignoring đó phản hồi code on /robots.txt. MỘT 5xx có thể stall crawling site-wide; treat đó file availability as production-cốt yếu.

cho rộng hơn pipeline điều này sits bên trong — phát hiện, crawl scheduler, kết xuất, và Cách crawling differs từ lập chỉ mục — see crawling hub. sibling topics (ngân sách crawl, và Cách Google xử lý sitemaps) mỗi go deeper on một piece của điều này.

Who's been ignoring my robots.txt?

This is live data from this site, not an illustration. My robots.txt disallows /api/trap/, and the only link to it is invisible to humans — so a compliant crawler will never request it. Every user-agent below fetched it anyway. (Humans poking at it with curl show up too; the user-agent usually gives them away.)

Loading trap log…

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.