Hướng dẫn về Crawl Demand

Đó "muốn" side of ngân sách crawl — điều gì làm Google muốn to crawl của bạn các trang (popularity, staleness, perceived inventory), cách demand đáp ứng host-load capacity, và vì sao bạn không thể force điều này.

Xuất bản lần đầu: 3 thg 7, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ

Crawl demand là đó 'muốn' side of ngân sách crawl — cách nhiều một công cụ tìm kiếm wants to crawl một site hoặc URL, as opposed to cách fast điều này có thể (đó là tốc độ crawl/capacity). Google names popularity (links/PageRank), staleness (cách thường một trang thay đổi), và perceived inventory (cách nhiều URLs Google thinks exist, junk được bao gồm — 'đó factor bạn có thể positively control đó hầu hết') as significant chung demand factors — không một closed three-item formula; Google cũng points to site size, cập nhật frequency, trang quality, và cách một site compares to similar ones. Site moves temporarily spike demand. Host load acts as một ceiling on realized demand, không một driver — my synthesis: demand sets đó priority order of URLs, capacity decides cách far xuống đó queue Googlebot nhận. Bạn không thể set demand trực tiếp — bạn move của nó inputs by earning links, giữ nội dung genuinely fresh, và cutting junk-URL inventory so quality các tín hiệu feed back vào đó scheduler. Google là thực ra trying to crawl ít hơn overall trong khi routing demand hơn precisely, so đó real goal đã là không bao giờ volume — đây là correct prioritization. Hầu hết các trang không bao giờ cần to manage này.

TL;DR — Crawl demand là đó muốn side of ngân sách crawl; tốc độ crawl/capacity là đó có thể side. Google names popularity (links / PageRank), staleness (cách thường một trang thay đổi), và perceived inventory (cách nhiều URLs Google thinks exist, junk được bao gồm — “the factor you can positively control the most” (bản dịch) «đó factor bạn có thể positively control đó hầu hết») as significant chung demand factors — không một closed formula; site size, cập nhật frequency, trang quality, và comparative relevance cũng factor trong. Site moves spike demand temporarily. My synthesis cho tying điều này together: demand sets đó priority order of URLs; host-load capacity decides cách far xuống đó queue Googlebot nhận — Google không document này as một literal algorithm, nhưng đây là đó model đó fits đó evidence. MỘT healthy máy chủ không manufacture demand, và cao demand có thể vẫn là capacity-throttled. Bạn không thể set demand trực tiếp — đó chỉ levers là của nó inputs, và đó scheduler turns demand up khi quality các tín hiệu từ lập chỉ mục improve. Meanwhile Google là actively trying to crawl ít hơn trong khi routing demand hơn precisely, so đó real goal đã là không bao giờ volume — đây là correct prioritization. Hầu hết các trang không bao giờ cần to manage any of này.

Evidence for this claim Google says Googlebot demand varies by site size, update frequency, page quality, and relevance compared with other sites; significant general demand factors are perceived inventory, popularity, and staleness. Scope: large or rapidly changing websites Confidence: high · Verified: Optimize your crawl budget

Crawl demand là “muốn,” tốc độ crawl là “có thể”

Google là rõ ràng đó ngân sách crawl có hai halves: đó amount of time và các tài nguyên Google devotes to crawling một site “is determined by two main elements: crawl capacity limit and crawl demand.” (bản dịch) «là determined by hai main elements: crawl capacity limit và crawl demand.» Đó way I frame điều này trong my Ahrefs crawl-budget hướng dẫn: ngân sách crawl là “made up crawl demand which is how many pages a search engine wants to crawl on your site and crawl rate which is how fast they can crawl.” (bản dịch) «đã làm up crawl demand mà là cách nhiều các trang một công cụ tìm kiếm wants to crawl on của bạn site và tốc độ crawl mà là cách fast they có thể crawl.» Demand là đó muốn; rate là đó có thể. Evidence for this claim Google describes crawl demand as one of the two main elements of crawl budget, alongside crawl capacity limit. Scope: Google Search crawling. Confidence: high · Verified: Google: Large site crawl budget guide

điều này trang là chỉ về muốn. có thể side — crawl capacity limit, GSC rate slider đó là đã xóa trong January 2024, Cách 5xx/429 các phản hồi chậm Googlebot, và Bing manual Crawl Control grid — all lives on tốc độ crawl trang. I’m không going để re-derive nó ở đây; Khi hai interact I’ll link across.

three demand inputs

Google hiện tại hướng dẫn calls perceived inventory, popularity, và staleness significant chung factors đó drive Cách nhiều nó wants để crawl — nó không present them as exhaustive, closed formula. giống nhau hướng dẫn cũng names trang web size, Cách thường nó đã cập nhật, trang quality, và Cách nó compares để similar các trang as factors Googlebot weighs. three dưới là ones Google giải thích trong phần lớn depth và ones Bạn có thể act on trực tiếp. Evidence for this claim Google's crawl-demand guidance discusses perceived inventory, popularity, and staleness. Scope: These inputs influence crawling but do not provide a user-controlled demand setting. Confidence: high · Verified: Google: Large site crawl budget guide

Demand decides which URLs sit at the front of the queue; a healthy server only determines how much of that demand can be realized. Nguồn: Google Search Central

Crawl demand orders URLs using popularity, genuine change, and the perceived value of the site's URL inventory. Crawl capacity, based on server response speed, stability, and errors, determines how far Googlebot can proceed through that ordered queue. Faster infrastructure raises the capacity ceiling but does not create demand for low-priority URLs.

© Patrick Stox LLC · CC BY 4.0 ·

Popularity

“URLs that are more popular on the Internet tend to be crawled more often to keep them fresher in our systems.” (bản dịch) «URLs đó là hơn popular on đó Internet tend to là được crawl hơn thường to giữ them fresher trong của chúng ta các hệ thống.» Hơn links và hơn PageRank pointing tại một URL là một demand tín hiệu — đây là vì sao của bạn homepage nhận được crawl constantly và một deep, unlinked trang barely tại all. As I put điều này trong my crawl-budget hướng dẫn, “Popular pages, or those with more links and PageRank, will generally receive priority over other pages.” (bản dịch) «Popular các trang, hoặc những với hơn links và PageRank, sẽ generally nhận priority over other các trang.» Liên kết nội bộ count ở đây cũng: một trang không có gì links to (an orphan) có gần như không demand hoạt động cho điều này.

Staleness

“Our systems want to recrawl documents frequently enough to pick up any changes.” (bản dịch) «Của chúng ta các hệ thống muốn to recrawl documents frequently đủ to pick up any thay đổi.» Google learns mỗi trang rhythm. MỘT trang đó thay đổi constantly earns frequent recrawls; một trang đó không bao giờ thay đổi nhận checked ít hơn và ít hơn. Trong my crawl-budget hướng dẫn I mô tả đó backoff Google áp dụng to một static trang: “if they crawl a page and see no changes after a day, they may wait three days before crawling again, ten days the next time, 30 days, 100 days, etc.” (bản dịch) «nếu they crawl một trang và see không thay đổi sau một day, they có thể chờ three days trước crawling again, ten days đó tiếp theo time, 30 days, 100 days, etc.» Đó theo-URL recrawl cadence là thực sự một crawl frequency câu hỏi — I cover đó mechanics ở đó — nhưng đó underlying force là demand, và cụ thể staleness.

Perceived inventory ( một bạn control phần lớn)

Này là đó headline lever. Google: “Without guidance from you, Google tries to crawl all or most of the URLs that it knows about on your site. If many of these URLs are duplicates, or you don’t want them crawled for some other reason (removed, unimportant, and so on), this wastes a lot of Google crawling time on your site. This is the factor that you can positively control the most.” (bản dịch) «Không có hướng dẫn từ bạn, Google tries to crawl all hoặc hầu hết of đó URLs đó điều này knows về trên trang web của bạn. Nếu nhiều of những URLs là duplicates, hoặc bạn không muốn them được crawl cho some other reason (đã xóa, unimportant, và so on), này wastes một lot of Google crawling time trên trang web của bạn. Này là đó factor đó bạn có thể positively control đó hầu hết.»

Evidence for this claim Google calls perceived inventory the crawl-demand factor site owners can positively control the most; duplicate, removed, and unimportant known URLs can waste crawling time. Scope: large or rapidly changing websites Confidence: high · Verified: Optimize your crawl budget

Đó subtle part là điều gì điều này làm to demand, không chỉ capacity. đây là easy to think of junk URLs as “wasting crawl budget” (bản dịch) «wasting ngân sách crawl» — spending fetches on copies thay vì new nội dung. Đúng. Nhưng có một demand-side effect cũng: một site whose knowable inventory là mostly thấp-giá trị duplicates và parameter sprawl looks, to Google, như một thấp hơn-giá trị site to crawl. Cutting perceived inventory không chỉ free capacity; theo thời gian điều này concentrates demand on đó URLs đó deserve điều này. Faceted navigation, session IDs, infinite calendar spaces, và other spider traps là đó classic inventory inflators — và đó classic demand suppressors.

trang web moves và khác demand spikes

Một demand driver không về any single URL: “Additionally, site-wide events like site moves may trigger an increase in crawl demand in order to reprocess the content under the new URLs.” (bản dịch) «Additionally, site-wide events như site moves có thể trigger an increase trong crawl demand trong order to reprocess đó nội dung under đó new URLs.» Nếu bạn làm một domain migration hoặc một big replatform và notice Googlebot hitting bạn far harder hơn thông thường cho vài weeks, đó là dự kiến — Google có to re-fetch và reprocess mọi thứ under đó new addresses. đây là một tạm thời spike, không một new baseline, và đây là một demand event đối thủ các hướng dẫn gần như không bao giờ mention despite điều này đang verbatim trong Google own tài liệu.

Evidence for this claim Google says site-wide events such as site moves may temporarily increase crawl demand so content can be reprocessed under new URLs. Scope: large or rapidly changing websites Confidence: high · Verified: Optimize your crawl budget

Cách demand và host load interact: queue order so với. capacity gate

Ở đây một mental model đó làm đó toàn bộ topic click — và I muốn to là upfront đó đây là my synthesis, không điều gì đó Google documents as một literal algorithm. đây là được xây dựng on một Gary Illyes Q&MỘT, quoted ở đây qua Công cụ tìm kiếm Roundtable coverage: host load “sets a bucket of URLs in importance order and GoogleBot will crawl in that order based on the schedule the host load decided. If Google thinks your server can handle it, it will crawl the whole bucket, if not, it will stop.” (bản dịch) «sets một bucket of URLs trong importance order và GoogleBot sẽ crawl trong đó order dựa trên đó schedule đó host load decided. Nếu Google thinks máy chủ của bạn có thể xử lý điều này, điều này sẽ crawl đó toàn bộ bucket, nếu không, điều này sẽ dừng.» Notably, theo đó giống nhau Q&MỘT, host load tracks đó importance of của bạn các trang — không đó thô number of URLs bạn có hoặc cách nhiều bạn muốn được crawl. Đó trang blocks automated fetching, so I haven’t đã able to re-xác nhận đó chính xác wording trực tiếp so với đó trực tiếp nguồn này truyền — treat điều này as một well-corroborated paraphrase, không một verbatim chính-nguồn quote.

đọc đó cẩn thận và mối quan hệ falls out — again, Đây là Cách I connect pieces, không mechanism Google có spelled out end để end:

  • Demand sets đó order. Đó “bucket of URLs in importance order” (bản dịch) «bucket of URLs trong importance order» crawl demand — popularity và staleness deciding mà URLs sit tại đó top of đó queue.
  • Capacity sets cách far Google nhận. Host load / tốc độ crawl decides cách deep vào đó ordered bucket Googlebot thực ra crawl on một được cho day. “If your server can handle it, it crawls the whole bucket; if not, it stops.” (bản dịch) «Nếu máy chủ của bạn có thể xử lý điều này, điều này crawl đó toàn bộ bucket; nếu không, điều này dừng.»

So đó hai không chỉ multiplied together — they play khác nhau vai trò. MỘT healthy, fast máy chủ không manufacture demand (điều này chỉ raises đó ceiling on cách nhiều of của bạn existing demand nhận realized), và cao demand có thể vẫn là capacity-throttled (một chậm hoặc lỗi-prone máy chủ dừng Googlebot part-way xuống đó queue không quan trọng cách nhiều điều này wants to crawl). Này là vì sao “I bought a faster server and Google still isn’t crawling my new pages” (bản dịch) «I bought một nhanh hơn máy chủ và Google vẫn không crawling my new các trang» là such một phổ biến, frustrating kết quả: capacity đã là không bao giờ đó constraint — demand đã là.

Bạn có thể’t đặt demand trực tiếp — nhưng scheduler listens

Ngắn câu trả lời: bạn không đặt demand trực tiếp. bạn earn nó, indirectly, qua thực links và thực quality improvements đó hiển thị up trong lập chỉ mục các tín hiệu — không có gì khác moves nó.

Có không “crawl hơn” yêu cầu cho demand any hơn có cho rate. Nhưng demand là dynamic, và Google đã được unusually candid về cách điều này moves. Gary Illyes: “If you want to increase how much we crawl, then you somehow have to convince search that your stuff is worth fetching, which is basically what the scheduler is listening to.” (bản dịch) «Nếu bạn muốn to increase cách nhiều we crawl, thì bạn somehow có to convince tìm kiếm đó của bạn stuff là worth fetching, mà là basically điều gì đó scheduler là listening to.» Và đó feedback loop là real-time-ish: “Scheduling is very dynamic. As soon as we get the signals back from search indexing that the quality of the content has increased across this many URLs, we would just start turning up demand.” (bản dịch) «Scheduling là very dynamic. As soon as we nhận đó các tín hiệu back từ tìm kiếm lập chỉ mục đó đó quality of đó nội dung có increased across này nhiều URLs, we sẽ chỉ bắt đầu chuyển thành up demand.» Đó flip side cũng: “If search demand goes down, then that also correlates to the crawl limit going down.” (bản dịch) «Nếu tìm kiếm demand goes xuống, thì đó cũng correlates to đó crawl limit going xuống.» (I re-checked những three lines so với Search Engine Journal’s coverage này truyền và they match verbatim; I vẫn haven’t tracked xuống Google own podcast audio/transcript to xác nhận them as một chính nguồn, so treat them as well-corroborated phụ quotes.)

Đó reframes “how do I increase crawl demand” (bản dịch) «cách làm I increase crawl demand» away từ tricks. Fake lastmod timestamps, sitemap pings, và xuất bản volume không convince đó scheduler. Đó hai điều đó làm là đó hai hard điều: real popularity (links) và real quality improvements đó cho thấy up trong các tín hiệu lập chỉ mục và feed back vào đó scheduler. Mọi thứ khác là theater.

Google là trying để crawl ít hơn, không nhiều hơn

Đó single freshest angle ở đây, và một older các hướng dẫn miss hoàn toàn: Google own stated goal là to reduce total crawl volume, không grow điều này. Trong an April 2024 LinkedIn post, Illyes wrote: “My mission this year is to figure out how to crawl even less, and have fewer bytes on wire.” (bản dịch) «My mission này năm là to hình out cách crawl even ít hơn, và có ít hơn bytes on wire.» He pushed back on đó idea đó Google đã có slashed crawling — “we’re crawling roughly as much as before, however scheduling got more intelligent” (bản dịch) «chúng ta là crawling roughly as nhiều as trước, tuy nhiên scheduling đã nhận hơn intelligent» — và được diễn đạt đó goal as một shared win: “Decreasing crawling without sacrificing crawl-quality would benefit everyone.” (bản dịch) «Decreasing crawling không có sacrificing crawl-quality sẽ benefit mọi người.» Đó mechanisms he pointed tại đã là tốt hơn bộ nhớ đệm, bộ nhớ đệm sharing across người dùng-agents, và ít hơn bytes transferred — không “crawl my site more.” (bản dịch) «crawl my site hơn.»

takeaway: crawl demand là không bao giờ điều gì đó để maximize. Google là actively optimizing cho ít hơn total crawling với equal hoặc tốt hơn crawl quality, routing demand nó làm có toward các URL nhiều hơn có khả năng để deserve nó. của bạn goal không phải nhiều hơn demand — nó đúng prioritization của demand bạn’ve earned.

Làm bạn even có demand vấn đề?

Hầu hết các trang không, và không nên spend một minute on này. Google own de-escalation áp dụng squarely to demand: “If your site doesn’t have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don’t need to read this guide.” (bản dịch) «Nếu trang web của bạn không có một lớn number of các trang đó thay đổi rapidly, hoặc nếu của bạn các trang seem to là được crawl đó giống nhau day đó they là published, bạn không cần to đọc này hướng dẫn.» John Mueller đã được similarly blunt về scale — theo Công cụ tìm kiếm Roundtable coverage of his tweet, 100k URLs là thường không đủ to ảnh hưởng ngân sách crawl, since đó hoạt động out to well under một crawl theo minute over three months.

Nếu bạn big đủ to care, ở đây đó diagnostic đó separates một demand vấn đề từ một capacity một — với một caveat up front: điều này produces một hypothesis to kiểm thử, không một diagnosis. Pull up GSC Crawl Số liệu báo cáo và máy chủ của bạn logs. Nếu host status là healthy và average thời gian phản hồi là fine nhưng một set of URLs là barely getting được crawl — và họ là stuck trong “Discovered – currently not indexed” (bản dịch) «Discovered – hiện tại không được lập chỉ mục» — đó pattern là evidence pointing toward demand, không proof of điều này. Crawl Số liệu cho thấy crawl activity (một capacity-side view), không một demand score — ở đó là không công khai theo-site “crawl demand score” (bản dịch) «crawl demand score» anywhere, so bạn là luôn inferring demand từ activity plus lập chỉ mục status, không bao giờ reading điều này off một dashboard.

Trước khi bạn act on “it’s a demand problem,” (bản dịch) «đây là một demand vấn đề,» rule out đó other điều đó produce đó giống nhau healthy-máy chủ-nhưng-không-được crawl symptom: Google có thể không có discovered đó URLs tuy vậy (không path trong, không sitemap entry), kết xuất có thể là hiding đó nội dung Googlebot cần to see, canonicalization có thể point Google nơi nào đó khác hoàn toàn, real quality các vấn đề (thin, duplicate, thấp-giá trị) có thể nhận một trang được crawl nhưng có chủ ý left out of đó chỉ mục, và Google own lập chỉ mục selection có thể sit một trang out even khi crawling và quality là cả hai fine. Chỉ sau những là checked và không giải thích điều này làm “thấp demand” become đó hoạt động lời giải thích — và even thì, treat điều này as đó best-supported hypothesis, không một confirmed nguyên nhân. Nhanh hơn hardware sẽ không cách sửa một genuine demand vấn đề; đó cách sửa là on đó demand inputs: links to những các trang, genuine reasons to recrawl them, và ít hơn junk inventory drowning them out.

cho ground-truth, theo-URL crawl dữ liệu, log analysis là câu trả lời. I’ll flag một hiện tại tool I có thể speak để firsthand trong Tools tab.

Crawl demand so với. rate so với. budget so với. frequency

giữ family straight:

  • Crawl demand — cách nhiều Google wants to crawl (popularity + staleness + perceived inventory). Này trang.
  • Tốc độ crawl — cách fast điều này có thể (capacity / host load). Của nó own trang.
  • Ngân sách crawl — đó hai together: “the number of URLs Googlebot can and wants to crawl.” (bản dịch) «đó number of URLs Googlebot có thể và wants to crawl.»
  • Crawl frequency — cách thường một được cho URL nhận recrawled, mà là một demand output (mostly staleness).
Evidence for this claim Google's crawl-demand guidance discusses perceived inventory, popularity, and staleness. Scope: These inputs influence crawling but do not provide a user-controlled demand setting. Confidence: high · Verified: Google: Large site crawl budget guide

Bing không dùng đó term “crawl demand” (bản dịch) «crawl demand» — điều này reframes đó toàn bộ điều as crawl efficiency: Fabrice Canel defines điều này as “how often we crawl and discover new and fresh content per page crawled,” (bản dịch) «cách thường we crawl và discover new và fresh nội dung theo trang được crawl,» và Bing philosophy là inventory-reduction-đầu tiên, đó demand-side tương đương of perceived inventory. IndexNow là Bing way of signaling demand-relevant thay đổi events thay vì đang chờ cho đó scheduler to infer staleness — nhưng note Google không dùng IndexNow, so điều này sẽ không move Google demand.

Evidence for this claim Google's crawl-demand guidance discusses perceived inventory, popularity, and staleness. Scope: These inputs influence crawling but do not provide a user-controlled demand setting. Confidence: high · Verified: Google: Large site crawl budget guide

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.