Ngân sách crawl

Điều gì ngân sách crawl thực ra là — crawl capacity plus crawl demand — điều gì wastes điều này, và đó honest kiểm thử cho liệu trang web của bạn là big đủ to cần to care tại all.

Xuất bản lần đầu: 22 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Ngân sách crawl là cách nhiều một công cụ tìm kiếm có thể và wants to crawl trang web của bạn — crawl capacity (điều gì máy chủ của bạn có thể take) times crawl demand (popularity và staleness). đây là không phải là yếu tố xếp hạng: hơn crawling sẽ không lift positions. Hầu hết các trang không bao giờ cần to manage điều này — Mueller says 100k URLs thường sẽ không move đó needle, và Google own hướng dẫn tells nhỏ hoặc giống nhau-day-được crawl các trang không to bother. Điều này matters mainly tại 1M+ các trang, 10k+ các trang thay đổi daily, hoặc khi lots of URLs sit trong 'Discovered – hiện tại không được lập chỉ mục.' Đó biggest lever là removing waste (faceted nav, duplicates, soft 404s, infinite spaces) so đó budget lands on URLs đó quan trọng.

TL;DR — Ngân sách crawl = crawl capacity limit (điều gì máy chủ của bạn có thể take) × crawl demand (popularity + staleness + perceived inventory). đây là an efficiency concern, không phải là tín hiệu xếp hạng. Internally đây là scheduling by importance gated by host load, không một flat theo-site quota. Hầu hết các trang có thể bỏ qua điều này — Mueller pegs 100k URLs as “usually not enough,” (bản dịch) «thường không đủ,» và Google tells giống nhau-day-được crawl các trang to skip đó hướng dẫn. Điều này bites tại ~1M+ các trang (weekly thay đổi), 10k+ các trang (daily thay đổi), hoặc khi “Discovered – currently not indexed” (bản dịch) «Discovered – hiện tại không được lập chỉ mục» balloons. Đó highest-leverage move là cutting waste — faceted nav, duplicates, soft 404s, infinite spaces — so đó budget consolidates on URLs đó quan trọng.

Evidence for this claim Google defines crawl budget using crawl capacity limit and crawl demand. Scope: Google Search crawling for larger sites. Confidence: high · Verified: Google: Large site crawl budget guide

hai-factor model

Google defines điều này cleanly: “The amount of time and resources that Google devotes to crawling a site is commonly called the site’s crawl budget and it’s determined by two main elements: crawl capacity limit and crawl demand.” (bản dịch) «Đó amount of time và các tài nguyên đó Google devotes to crawling một site là commonly called đó site ngân sách crawl và đây là determined by hai main elements: crawl capacity limit và crawl demand.» Đó 2017 framing từ Gary Illyes là đó một-liner I vẫn reach cho: ngân sách crawl là “the number of URLs Googlebot can and wants to crawl.” (bản dịch) «đó number of URLs Googlebot có thể và wants to crawl.» Evidence for this claim Google defines crawl budget using crawl capacity limit and crawl demand. Scope: Google Search crawling for larger sites. Confidence: high · Verified: Google: Large site crawl budget guide Giữ này scoped: ngân sách crawl governs fetching, không lập chỉ mục. MỘT được crawl URL vẫn goes qua một tách biệt lập chỉ mục decision — folding đó hai together overstates điều gì ngân sách crawl controls.

Crawl capacity limit (đó supply side). Này là “the maximum number of simultaneous parallel connections that Google can use to crawl a site, as well as the time delay between fetches.” (bản dịch) «đó maximum number of simultaneous parallel connections đó Google có thể dùng to crawl một site, cũng như đó time delay giữa fetches.» Điều này moves với máy chủ của bạn health. Respond fast và sạch và đó limit rises; serve chậm các phản hồi, 5xx các lỗi, hoặc 429s và Googlebot backs off. Trong my Cách Tìm kiếm Hoạt động deck I list đó giống nhau rate-limit triggers: máy chủ stability, chậm các phản hồi, 5xx máy chủ các lỗi, và 429 (cũng nhiều các yêu cầu). Này là đó “có thể.”

Crawl demand (đó demand side). Driven by popularity (cách linked-to / quan trọng một URL là) và staleness (cách dài since điều này đã là cuối cùng được crawl, cách thường điều này thay đổi). Đó giống nhau deck breaks demand vào PageRank, cách frequently đó trang thay đổi, time since cuối cùng crawl, và major site thay đổi. Critically, Google flags perceived inventory as đó lever bạn control hầu hết: “Without guidance from you, Google tries to crawl all or most of the URLs that it knows about on your site. If many of these URLs are duplicates, or you don’t want them crawled for some other reason… this wastes a lot of Google crawling time on your site. This is the factor that you can positively control the most.” (bản dịch) «Không có hướng dẫn từ bạn, Google tries to crawl all hoặc hầu hết of đó URLs đó điều này knows về on của bạn site. Nếu nhiều of những URLs là duplicates, hoặc bạn không muốn them được crawl cho some other reason… này wastes một lot of Google crawling time trên trang web của bạn. Này là đó factor đó bạn có thể positively control đó hầu hết.»

Demand sets the priority order; capacity determines how much of that ordered queue Googlebot can actually crawl. Nguồn: Google Search Central

Crawl demand comes from popularity, genuine change, and the value of the URL inventory. It orders URLs in a priority queue. Crawl capacity comes from server response speed, stability, and error behavior. It limits how far Googlebot proceeds through that queue. Their interaction is the site's realized crawl budget, not a fixed daily URL quota.

© Patrick Stox LLC · CC BY 4.0 ·

couple của structural facts đó catch mọi người out:

  • Budget là theo hostname. https://www.example.com/ and https://code.example.com/ are two different hostnames, and therefore have separate crawl budgets.” (bản dịch) «https://www.example.com/https://code.example.com/ là hai khác nhau hostnames, và do đó có tách biệt crawl budgets. » Subdomains không share.
  • Đó khác nhau Googlebot types có khả năng draw từ một pool. Trong my own reporting, image, news, video, quảng cáo, và đó rest xuất hiện to pull từ đó giống nhau theo-site budget — I không có một hiện tại chính Google nguồn pinning này xuống chính xác, so treat điều này as một practitioner observation thay vì được ghi lại policy. Either way, đây là worth kiểm tra đó Crawl Số liệu báo cáo by-crawler-loại breakdown nếu bạn suspect một loại là elbowing out đó rest.

Điều gì nó thực sự là internally: scheduling by importance

“Crawl budget” (bản dịch) «Ngân sách crawl» là an SEO-coined umbrella term. Internally đây là closer to scheduling gated by host load. As Illyes có described điều này, Google scheduler “sets a bucket of URLs in importance order and GoogleBot will crawl in that order based on the schedule the host load decided. If Google thinks your server can handle it, it will crawl the whole bucket, if not, it will stop.” (bản dịch) «sets một bucket of URLs trong importance order và GoogleBot sẽ crawl trong đó order dựa trên đó schedule đó host load decided. Nếu Google thinks máy chủ của bạn có thể xử lý điều này, điều này sẽ crawl đó toàn bộ bucket, nếu không, điều này sẽ dừng.»

Đó reframes đó toàn bộ topic. Điều này không một flat “you get N pages a day” (bản dịch) «bạn nhận N các trang một day» quota — đây là một prioritized queue, và crawling tracks tìm kiếm demand. Illyes again: “If search demand goes down, then that also correlates to the crawl limit going down,” (bản dịch) «Nếu tìm kiếm demand goes xuống, thì đó cũng correlates to đó crawl limit going xuống,»“if you want to increase how much we crawl, then you somehow have to convince search that your stuff is worth fetching, which is basically what the scheduler is listening to.” (bản dịch) «nếu bạn muốn to increase cách nhiều we crawl, thì bạn somehow có to convince tìm kiếm đó của bạn stuff là worth fetching, mà là basically điều gì đó scheduler là listening to.» Đó Tìm kiếm Relations team có explicitly called đó “fixed daily page quota” (bản dịch) «fixed daily trang quota» idea một misconception.

Làm của bạn trang web thực ra có ngân sách crawl vấn đề?

Này là đó hầu hết valuable section, so I’ll là blunt: hầu hết các trang không cần to worry về ngân sách crawl. Google own hướng dẫn opens với đó de-escalation: “If your site doesn’t have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don’t need to read this guide. For Google Search specifically, merely keeping your sitemap up to date and checking your index coverage regularly is adequate.” (bản dịch) «Nếu trang web của bạn không có một lớn number of các trang đó thay đổi rapidly, hoặc nếu của bạn các trang seem to là được crawl đó giống nhau day đó they là published, bạn không cần to đọc này hướng dẫn. Cho Google Search cụ thể, merely giữ của bạn sitemap up to date và kiểm tra của bạn chỉ mục coverage regularly là adequate.» Evidence for this claim Google says sites without many rapidly changing pages, or whose pages are crawled the day they publish, generally do not need crawl-budget guidance. Scope: Google's rough applicability guidance, not a guarantee for every site. Confidence: high · Verified: Google: Large site crawl budget guide

John Mueller gave đó concrete number: “100k URLs is usually not enough to affect crawl budget (it’s <1/minute over 3 months).” (bản dịch) «100k URLs là thường không đủ to ảnh hưởng ngân sách crawl (đây là <1/minute over 3 months).» Nếu bạn là under six figures of URLs và các trang nhận được crawl promptly, move on.

Khi nó làm quan trọng, Google rough thresholds là:

  • Lớn các trang — 1 million+ unique các trang với nội dung đó thay đổi moderately thường (về weekly).
  • Medium hoặc lớn hơn các trang — 10 000+ unique các trang với very rapidly thay đổi nội dung (daily).
  • Các trang với một lớn portion of URLs classified as “Discovered – currently not indexed” (bản dịch) «Discovered – hiện tại không được lập chỉ mục» trong Search Console — đó là đó warning light đó Google knows về URLs điều này không getting to.

Google adds đó disclaimer đó “the numbers given here are a rough estimate… not exact thresholds.” (bản dịch) «đó numbers được cho ở đây là một rough estimate… không chính xác thresholds.» Và even on big các trang, đó nuance từ my own hoạt động holds: đây là thường new, poorly-linked, hoặc static các trang đó lag — không của bạn popular ones.

Làm ngân sách crawl ảnh hưởng thứ hạng? Không.

Crawling là necessary to xếp hạng, nhưng điều này là không một tín hiệu xếp hạng. Google trong 2017: “An increased crawl rate will not necessarily lead to better positions in Search results. Google uses hundreds of signals to rank the results, and while crawling is necessary for being in the results, it’s not a ranking signal.” (bản dịch) «An increased tốc độ crawl sẽ không nhất thiết lead to tốt hơn positions trong Tìm kiếm kết quả. Google dùng hundreds of các tín hiệu to xếp hạng đó kết quả, và trong khi crawling là necessary cho đang trong đó kết quả, đây là không phải là tín hiệu xếp hạng.» I put điều này đó giống nhau way trong my Ahrefs hướng dẫn: “More crawling doesn’t mean you’ll rank better, but if your pages aren’t crawled and indexed they aren’t going to rank at all.” (bản dịch) «Hơn crawling không có nghĩa là bạn’ll xếp hạng tốt hơn, nhưng nếu của bạn các trang không được crawl và được lập chỉ mục they không going to xếp hạng tại all.» Treat crawl budget as an efficiency vấn đề, đầy đủ dừng.

Điều gì wastes ngân sách crawl

Illyes published đó canonical list of thấp-giá trị-thêm URLs “in order of significance” (bản dịch) «trong order of significance»:

  1. Faceted navigation và session identifiers — #1 culprit, especially ecommerce filter/sort combinations đó multiply các URL combinatorially.
  2. On-trang web duplicate nội dung — kinh điển kỹ thuật variants: HTTP so với HTTPS, non-www so với www, trailing slash so với không, uppercase so với lowercase, default/chỉ mục các trang, và URL parameters. (Roughly 60% của web là duplicate nội dung, by Google own internal estimate.)
  3. Soft lỗi các trang — soft 404s đó trả về 200 giữ getting được crawl.
  4. Hacked các trang.
  5. Infinite spaces và proxies — calendars, infinite-scroll pagination đó duplicates, faceted combinations; kinh điển spider-trap patterns.
  6. Thấp-quality và spam nội dung.

Đó cost là concrete: “Wasting server resources on pages like these will drain crawl activity from pages that do actually have value, which may cause a significant delay in discovering great content on a site.” (bản dịch) «Wasting máy chủ các tài nguyên on các trang như những sẽ drain crawl activity từ các trang đó làm thực ra có giá trị, mà có thể nguyên nhân một significant delay trong discovering great nội dung on một site.» On top of đó list, dài chuyển hướng chains “have a negative effect on crawling,” (bản dịch) «có một negative effect on crawling,» và chậm, nặng các trang làm mỗi fetch hơn expensive.

Cách optimize nó

toàn bộ game là consolidating budget onto các URL đó quan trọng:

  • Consolidate duplicates. Google: “Consolidate duplicate content to focus crawling on unique content rather than unique URLs.” (bản dịch) «Consolidate duplicate nội dung to focus crawling on unique nội dung thay vì unique URLs.» Pick một host, một giao thức, một trailing-slash convention; canonicalize; xử lý parameters.
  • Block truly worthless paths với robots.txt — nhưng chỉ paths bạn không bao giờ muốn được crawl. Cho faceted navigation, đó thông thường options là blocking đó parameter paths trong robots.txt hoặc dùng một # thay vì một ? so đó URLs không crawlable ngay từ đầu.
  • không dùng noindex to save budget. Google: “Don’t use noindex, as Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time.” (bản dịch) «không dùng noindex, as Google sẽ vẫn yêu cầu, nhưng thì drop đó trang khi điều này sees một noindex meta tag hoặc header trong đó HTTP phản hồi, wasting crawling time.» Đó yêu cầu vẫn costs bạn. Block tại robots.txt nếu bạn không bao giờ muốn điều này fetched.
  • không expect robots.txt to reallocate budget. “Google won’t shift this newly available crawl budget to other pages unless Google is already hitting your site’s serving limit.” (bản dịch) «Google sẽ không shift này newly khả dụng ngân sách crawl to other các trang trừ khi Google là đã hitting của bạn site serving limit.» Blocking junk là good hygiene, nhưng điều này không hand đó freed-up crawl to của bạn good các trang trừ khi bạn đã là capacity-bound.
  • Cách sửa soft 404s; trả về real 404/410 cho đã biến mất các trang. “A 404 status code is a strong signal not to crawl that URL again.” (bản dịch) «MỘT 404 mã trạng thái là một mạnh tín hiệu không to crawl đó URL again.»
  • Shorten chuyển hướng chains, giữ sitemaps hiện tại (honest lastmod), và improve máy chủ speed.
  • Strengthen liên kết nội bộ to quan trọng và new các trang — easier hơn bất cứ điều gì khác vì bạn fully control điều này.

Và đó hai — chỉ hai — ways Google says bạn có thể thực ra increase budget: “Add more server resources… [and] optimize your content’s quality.” (bản dịch) «Thêm hơn máy chủ các tài nguyên… [và] optimize nội dung của bạn quality.» Note đó trap ở đó: một nhanh hơn máy chủ lifts đó capacity ceiling, nhưng nếu demand là thấp Google vẫn crawl ít hơn. Bạn cần cả hai.

Cách đo lường nó

  • GSC > Settings > Crawl Số liệu báo cáo — total các yêu cầu crawl theo thời gian, average thời gian phản hồi, host status, và breakdowns by phản hồi code, file loại, và Googlebot loại. Này là Google own view of cách điều này crawl bạn.
  • Máy chủ log file analysis — đó ground truth. Real Googlebot hits by URL pattern, so bạn có thể see crawl waste và uncrawled quan trọng các trang. Verify đó bot là genuinely Googlebot qua reverse + forward DNS hoặc Google published IP ranges (plenty of fake bots spoof người dùng-agent).
  • “Discovered – currently not indexed” (bản dịch) «Discovered – hiện tại không được lập chỉ mục» trong GSC — treat một growing pile ở đây as một crawl-budget warning light: Google knows về đó URLs nhưng không getting to them.
Crawl waste appears in the mismatch: facet URLs occupy 45% of the inventory and 61% of requests, but only 8% of useful 200 responses.

In a synthetic cohort, product pages are 28 percent of the URL inventory, 24 percent of Googlebot requests, and 52 percent of useful 200 responses. Category pages are 7, 10, and 21 percent. Facet URLs are 45, 61, and 8 percent. Gone URLs are 20, 5, and 0 percent. The figures illustrate comparison logic, not a live log sample.

Bing và other engines: “crawl efficiency” (bản dịch) «hiệu quả crawl»

Bing reframes đó topic as hiệu quả crawl thay vì budget. Fabrice Canel definition: “The crawl efficiency is how often we crawl and discover new and fresh content per page crawled.” (bản dịch) «Đó hiệu quả crawl là cách thường we crawl và discover new và fresh nội dung theo trang được crawl.» Đó goal là to “crawl an URL only when the content has been added (URL not crawled before), updated (fresh on-page context or useful outbound links).” (bản dịch) «crawl an URL chỉ khi đó nội dung đã được đã thêm (URL không được crawl trước), đã cập nhật (fresh on-trang context hoặc hữu ích outbound links).» Bing blunt philosophy: “Less is more for SEO. Never forget that. Less URLs to crawl, better for SEO.” (bản dịch) «Ít hơn là hơn cho SEO. Không bao giờ forget đó. Ít hơn URLs to crawl, tốt hơn cho SEO.»

Bing được ưu tiên khắc phục là IndexNow — push changed các URL so bingbot không cần exploratory crawl — và Crawl Control trong Bing Quản trị viên web Tools, mà lets bạn schedule Khi bingbot crawl by hour để bảo vệ máy chủ load. Đây là thực Google/Bing divergence worth noting: Google deprecated của nó old crawl-rate limiter trong Search Console, trong khi Bing vẫn lets bạn actively shape crawl schedule.

Ngân sách crawl myths, corrected

  • “Every site should optimize crawl budget.” (bản dịch) «Mỗi site nên optimize ngân sách crawl.» Không — hầu hết không nên. Giống nhau-day crawling và sub-100k URLs có nghĩa là bạn là fine.
  • “More crawling = better rankings.” (bản dịch) «Hơn crawling = tốt hơn thứ hạng.» Không. Crawling là necessary nhưng không một tín hiệu xếp hạng.
  • “It’s a fixed daily page quota.” (bản dịch) «đây là một fixed daily trang quota.» Không — đây là importance-driven scheduling gated by host load.
  • “Use noindex to save budget.” (bản dịch) «Dùng noindex to save budget.» Không — Google vẫn các yêu cầu đó trang đầu tiên.
  • “Block pages in robots.txt to give other pages more budget.” (bản dịch) «Block các trang trong robots.txt to cho other các trang hơn budget.» Generally không, trừ khi bạn là đã tại của bạn serving limit.
  • “A faster server alone raises your budget.” (bản dịch) «MỘT nhanh hơn máy chủ alone raises của bạn budget.» Điều này lifts đó capacity ceiling chỉ; thấp demand vẫn có nghĩa là ít hơn crawl.

cho rộng hơn pipeline điều này sits bên trong — discovery, crawl scheduler, kết xuất, và Cách crawling differs từ lập chỉ mục — see crawling hub. sibling topics (tốc độ crawl, crawl frequency, spider traps, và log file analysis) mỗi go deeper on một piece của điều này.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.