Hướng dẫn về Crawl Delay

Đó crawl-delay robots.txt directive — điều gì điều này làm, vì sao Google đã dừng honoring điều này trong 2019 và Yandex dropped điều này trong 2018, cách Bing vẫn interprets đó giá trị, mà other các crawler respect điều này, và điều cần dùng thay vì.

Xuất bản lần đầu: 27 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Crawl-delay là một non-tiêu chuẩn robots.txt directive đó asks bots to chờ giữa fetches to ease máy chủ load. Google có đã bỏ qua điều này since September 1, 2019; chậm Googlebot với tạm thời 429/503 các phản hồi hoặc by sửa máy chủ capacity. Bing hiện tại Quản trị viên web hướng dẫn documents các giá trị từ 1–20 seconds. Yandex dropped hỗ trợ on February 22, 2018 và hiện tại dùng một crawl-rate setting bên trong Yandex Quản trị viên web. Kiểm tra mỗi crawler own tài liệu trước relying on này non-tiêu chuẩn trường.

Tóm tắt — Crawl-delay là không chính thức robots.txt directive cho throttling bots; nó là không bao giờ trong formal tiêu chuẩn (RFC 9309). Google retired nó September 1, 2019 và explicitly không xử lý nó — chậm Googlebot với 429/503 (1-2 days max) hoặc by sửa máy chủ, không crawl-delay. Bing hiện tại hướng dẫn documents các giá trị từ 1–20 seconds. Yandex dropped hỗ trợ cho nó on February 22, 2018 — nó không honored ở đó either; Yandex hiện tại dùng crawl-rate setting bên trong Yandex Quản trị viên web. nhiều SEO các crawler (AhrefsBot, Semrush) và some AI bots (ClaudeBot) respect nó cũng — mà là một thực sử dụng case left: controlling non-Google, non-Bing, non-Yandex bots.

Evidence for this claim Google does not support or process the non-standard crawl-delay robots.txt field. Scope: Google crawlers; other crawlers may support the field. Confidence: high · Verified: Google: robots.txt specifications

Vì sao crawl-delay tồn tại

Crawl-delay là throttle. idea là straightforward: aggressive crawler fetching các trang back-để-back có thể put thực load on máy chủ, especially nhỏ hoặc chậm một. Crawl-delay là courtesy lever — way để ask bot để pause một vài seconds giữa các yêu cầu so của bạn máy chủ có thể giữ up. nó lives trong robots.txt bên trong User-agent group, giống nhau place của bạn disallowallow lines go. (cho đầy đủ file, see robots-txt.)

catch là đó nó là không bao giờ standardized. Robots Exclusion Giao thức là finally được chuẩn hóa as RFC 9309 trong 2022, và crawl-delay không phải trong nó. nó là luôn không chính thức extension đó khác các crawler chose để implement — hoặc không — và interpreted trong của họ own way. đó inconsistency là chính xác Vì sao Google walked away từ nó.

Google retired nó on September 1, 2019

On July 2, 2019, Gary Illyes published note on unsupported rules trong robots.txt on Google Search Central. Google là open-sourcing của nó robots.txt parser và, as part của đó, retiring all code đó handled rules đó là không bao giờ part của internet draft — cụ thể noindex, nofollow, và crawl-delay. retirement took effect September 1, 2019.

Google reasoning: những rules đã là undocumented, không chính thức, và interpreted inconsistently across các crawler, mà đã tạo ambiguity. Đó hiện tại robots.txt tài liệu là hiện tại rõ ràng — Google hỗ trợ four các trường (user-agent, allow, disallow, và sitemap), và đó tài liệu say “other fields such as crawl-delay aren’t supported.” (bản dịch) «other các trường such as crawl-delay không supported.» Google Myths và facts về crawling trang repeats điều này: “The non-standard ‘crawl-delay’ robots.txt rule is not processed by Google’s crawlers.” (bản dịch) «Đó non-tiêu chuẩn ‘crawl-delay’ robots.txt rule không phải processed by Google các crawler.» Evidence for this claim Google does not support or process the non-standard crawl-delay robots.txt field. Scope: Google crawlers; other crawlers may support the field. Confidence: high · Verified: Google: robots.txt specifications

Worth là clear về phạm vi: nó đã bỏ qua không quan trọng mà người dùng-agent group bạn put nó trong. User-agent: Googlebot block với Crawl-delay line là chỉ as đã bỏ qua as một under User-agent: *.

Điều gì để sử dụng thay vì cho Google

Google tốc độ crawl là fully tự động hiện tại, tuned để của bạn máy chủ health — và old manual crawl-rate slider trong Search Console là đã xóa on January 8, 2024 (announced prior November). So Khi bạn genuinely cần Googlebot để back off, bạn có three levers:

  1. Cách sửa đó máy chủ. Đó real cách sửa. Nếu máy chủ của bạn có thể take đó load, bạn không cần to throttle bất cứ điều gì.
  2. Trả về 429, 500, hoặc 503 cho emergencies. Google Reduce đó Googlebot tốc độ crawl tài liệu say to “return 500, 503, or 429 HTTP response status code instead of 200 to the crawl requests.” (bản dịch) «trả về 500, 503, hoặc 429 HTTP phản hồi mã trạng thái thay vì 200 to đó các yêu cầu crawl.» Googlebot đọc những as “chậm xuống” gần như immediately. Đó hard rule: “We don’t recommend that you do this for a long period of time (meaning, longer than 1-2 days).” (bản dịch) «We không khuyến nghị đó bạn làm này cho một dài period of time (meaning, lâu hơn hơn 1-2 days).» Sustained 5xx risks các trang getting dropped (và Google Quảng cáo pausing). Evidence for this claim Google recommends temporarily returning 500, 503, or 429 to reduce crawl rate and warns against doing so for longer than one or two days. Scope: Temporary Googlebot overload response, not routine crawl management. Confidence: high · Verified: Google: Reduce Googlebot crawl rate
  3. File một crawl-rate reduction yêu cầu qua Search Console cho một persistent vấn đề. đây là chậm, và điều này chỉ reduces, không bao giờ raises.

Một trap to tránh: không dùng 401, 403, hoặc 404 to throttle. Theo Google HTTP các mã trạng thái hướng dẫn, “the 4xx status codes, except 429, have no effect on crawl rate,” (bản dịch) «đó 4xx các mã trạng thái, except 429, có không effect on tốc độ crawl,» và bạn không nên “use 401 and 403 status codes for limiting the crawl rate.” (bản dịch) «dùng 401403 các mã trạng thái cho limiting đó tốc độ crawl.»

Đây là cũng Vì sao, trong my crawl-rate hướng dẫn, I treat crawl-delay as điều không để sử dụng cho Google — see đó bài viết cho đầy đủ speed-up / chậm-xuống playbook.

Bing documents 1–20 thứ hai range

hiện tại Bing Quản trị viên web hướng dẫn documents crawl-delay các giá trị từ 1–20 seconds. đó không làm trường portable tiêu chuẩn: Google bỏ qua nó, và mỗi khác crawler cần của nó own được ghi lại xác nhận.

Yandex dropped nó trong 2018

Yandex được dùng để honor crawl-delay as một literal minimum number of seconds giữa các yêu cầu — đó là đó version of đó fact bạn’ll vẫn see repeated across hầu hết SEO blogs. đây là stale. Yandex own hiện tại tài liệu là rõ ràng: “From February 22, 2018, Yandex doesn’t take into account the Crawl-delay directive.” (bản dịch) «Từ February 22, 2018, Yandex không take vào account đó Crawl-delay directive.» Evidence for this claim Yandex stopped honoring the Crawl-delay directive on February 22, 2018, and now recommends setting crawl rate inside Yandex Webmaster instead. Scope: Yandex's crawler only; does not apply to Google, Bing, or other crawlers. Confidence: high · Verified: Yandex Webmaster: The Crawl-delay directive To set cách fast Yandex bots crawl trang web của bạn hiện tại, bạn dùng đó tốc độ crawl setting bên trong Yandex Quản trị viên web trực tiếp — không robots.txt. So as of hôm nay, đó honest scorecard là: Bing honors crawl-delay trong của nó được ghi lại range, Google không bao giờ đã làm, và Yandex được dùng để nhưng đã dừng trong 2018.

khác các crawler: SEO tools và AI bots

Plenty của well-behaved non-tìm kiếm-engine các crawler honor crawl-delay as courtesy:

  • AhrefsBotSemrush bot cả hai respect nó.
  • AI các crawler là newer audience. Anthropic ClaudeBot là được ghi lại as supporting non-tiêu chuẩn crawl-delay directive. Others (GPTBot, Perplexity) nên là checked so với của họ own published tài liệu thay vì assumed.

pattern: nhỏ hơn, compliance-focused, hoặc courtesy-minded các crawler tend để honor crawl-delay; major các công cụ tìm kiếm có largely moved để internal các tín hiệu (Google) hoặc richer manual tools (Bing). và none của điều này binds thực tế bad actors — scrapers đó bỏ qua robots.txt hoàn toàn sẽ bỏ qua của bạn crawl-delay cũng.

sử dụng case đó left

So là crawl-delay dead? không quite. nó useless cho Google, redundant-tại-best cho Bing (sử dụng Crawl Control), nhưng nó genuinely simplest way để ask rộng middle — SEO tools, AI bots, và khác minor các crawler đó respect nó — để ease off. nếu cụ thể tool bot là overrunning của bạn máy chủ, targeted User-agent: group với crawl-delay là reasonable đầu tiên move:

User-agent: SomeBot
Crawl-delay: 10

chỉ remember nó yêu cầu, không bảo đảm, và nó scopes theo host và theo người dùng-agent group như mọi thứ khác trong robots.txt. cho rộng hơn efficiency picture — Điều gì bots thực ra cost bạn và ai cần để care — see crawl-budget.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.