Crawl Demand
The "want" side of crawl budget — what makes Google want to crawl your pages (popularity, staleness, perceived inventory), how demand meets host-load capacity, and why you can't force it.
Crawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, as opposed to how fast it can (that's crawl rate/capacity). Google names popularity (links/PageRank), staleness (how often a page changes), and perceived inventory (how many URLs Google thinks exist, junk included — 'the factor you can positively control the most') as significant general demand factors — not a closed three-item formula; Google also points to site size, update frequency, page quality, and how a site compares to similar ones. Site moves temporarily spike demand. Host load acts as a ceiling on realized demand, not a driver — my synthesis: demand sets the priority order of URLs, capacity decides how far down that queue Googlebot gets. You can't set demand directly — you move its inputs by earning links, keeping content genuinely fresh, and cutting junk-URL inventory so quality signals feed back into the scheduler. Google is actually trying to crawl less overall while routing demand more precisely, so the real goal was never volume — it's correct prioritization. Most sites never need to manage this.
TL;DR — Crawl budgetThe number of URLs an engine will crawl in a timeframe. has two sides: how fast a search engine can fetch your pages (crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor.) and how much it wants to (crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side.). Crawl demand is the “want” side. Google wants to crawl a page more when it’s popular (lots of links), when it changes often, and when Google thinks the site is worth its time. You can’t push a button to raise demand — you earn it with links, genuine freshness, and by not burying your good pages under a pile of junk URLs.
What crawl demand is
When people say “crawl budgetThe number of URLs an engine will crawl in a timeframe.,” they’re really talking about two separate things squished together. One is your server’s ability to handle being crawled — how fast, how many pages at once. That’s crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor., and it’s covered on its own page. The other is how much the search engine actually wants to crawl you in the first place. That’s crawl demand — the topic here.
Think of it as supply and demand. Crawl rate is supply: how much crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. your site can support. Crawl demand is demand: how much Google feels like doing. Evidence for this claim Google describes crawl demand as one of the two main elements of crawl budget, alongside crawl capacity limit. Scope: Google Search crawling. Confidence: high · Verified: Google: Large site crawl budget guide
What makes Google want to crawl a page
Google points to a handful of significant factors, not one exhaustive checklist. The three it explains in the most depth:
- Popularity. Pages with more links pointing at them get crawled more often, so Google keeps its copy fresh.
- Staleness / freshness. If a page changes a lot, Google wants to check it more often. If it never changes, Google learns to check it less.
- How many URLs Google thinks you have. If your site is full of junk, duplicate, or low-value URLs, Google wastes its crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. on those instead of your real pages. This is the one you control the most. Evidence for this claim Google's crawl-demand guidance discusses perceived inventory, popularity, and staleness. Scope: These inputs influence crawling but do not provide a user-controlled demand setting. Confidence: high · Verified: Google: Large site crawl budget guide
Google also names a few site-level factors that shape demand alongside those three: how big your site is, how often you update it, page quality, and how your site compares to others covering similar ground. Don’t treat “popularity, staleness, perceived inventory” as a complete formula — it’s the biggest, most actionable levers, not the whole list.
There’s also a temporary one: if you move your site to a new domain, Google has to re-crawl everything to process it under the new URLs, so demand spikes for a while.
Why you can’t just “increase crawl demand”
There’s no dial for it — no more than there’s a button to make Google crawl faster. Publishing ten posts a day won’t do it if nobody links to them and they’re not genuinely useful. The things that actually work are slow and real: earn links, keep content genuinely fresh, and clean out the junk URLs so Google’s crawling lands on the pages that matter.
And here’s the surprise: more crawling isn’t even the goal. Getting crawled a lot doesn’t make you rank higher. What you want isn’t more demand — it’s the demand you have pointed at the right pages.
Want the deeper version — how demand and your server’s capacity interact, how site quality feeds back into the crawl scheduler, and how to tell a demand problem from a capacity problem? Switch to the Advanced tab.
TL;DR — Crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side. is the want side of crawl budgetThe number of URLs an engine will crawl in a timeframe.; crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor./capacity is the can side. Google names popularity (links / PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems.), staleness (how often a page changes), and perceived inventory (how many URLs Google thinks exist, junk included — “the factor you can positively control the most”) as significant general demand factors — not a closed formula; site size, update frequency, page quality, and comparative relevance also factor in. Site moves spike demand temporarily. My synthesis for tying it together: demand sets the priority order of URLs; host-load capacity decides how far down that queue GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. gets — Google doesn’t document this as a literal algorithm, but it’s the model that fits the evidence. A healthy server doesn’t manufacture demand, and high demand can still be capacity-throttled. You can’t set demand directly — the only levers are its inputs, and the scheduler turns demand up when quality signals from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. improve. Meanwhile Google is actively trying to crawl less while routing demand more precisely, so the real goal was never volume — it’s correct prioritization. Most sites never need to manage any of this.
Crawl demand is the “want,” crawl rate is the “can”
Google is explicit that crawl budgetThe number of URLs an engine will crawl in a timeframe. has two halves: the amount of time and resources Google devotes to crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. a site “is determined by two main elements: crawl capacity limit and crawl demand.” The way I frame it in my Ahrefs crawl-budget guide: crawl budget is “made up crawl demand which is how many pages a search engine wants to crawl on your site and crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor. which is how fast they can crawl.” Demand is the want; rate is the can. Evidence for this claim Google describes crawl demand as one of the two main elements of crawl budget, alongside crawl capacity limit. Scope: Google Search crawling. Confidence: high · Verified: Google: Large site crawl budget guide
This page is only about the want. The can side — the crawl capacity limit, the
GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. rate slider that was removed in January 2024, how 5xx/429 responses slow
GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer., and Bing’s manual Crawl Control grid — all lives on the crawl rate
page. I’m not going to re-derive it here; when the two interact I’ll link across.
The three demand inputs
Google’s current guidance calls perceived inventory, popularity, and staleness significant general factors that drive how much it wants to crawl — it doesn’t present them as an exhaustive, closed formula. The same guidance also names site size, how often it’s updated, page quality, and how it compares to similar sites as factors Googlebot weighs. The three below are the ones Google explains in the most depth and the ones you can act on directly. Evidence for this claim Google's crawl-demand guidance discusses perceived inventory, popularity, and staleness. Scope: These inputs influence crawling but do not provide a user-controlled demand setting. Confidence: high · Verified: Google: Large site crawl budget guide
Crawl demand orders URLs using popularity, genuine change, and the perceived value of the site's URL inventory. Crawl capacity, based on server response speed, stability, and errors, determines how far Googlebot can proceed through that ordered queue. Faster infrastructure raises the capacity ceiling but does not create demand for low-priority URLs.
© Patrick Stox LLC · CC BY 4.0 ·
Popularity
“URLs that are more popular on the Internet tend to be crawled more often to keep them fresher in our systems.” More links and more PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. pointing at a URL is a demand signal — it’s why your homepage gets crawled constantly and a deep, unlinked page barely at all. As I put it in my crawl-budget guide, “Popular pages, or those with more links and PageRank, will generally receive priority over other pages.” Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. count here too: a page nothing links to (an orphan) has almost no demand working for it.
Staleness
“Our systems want to recrawl documents frequently enough to pick up any changes.” Google learns each page’s rhythm. A page that changes constantly earns frequent recrawls; a page that never changes gets checked less and less. In my crawl-budget guide I describe the backoff Google applies to a static page: “if they crawl a page and see no changes after a day, they may wait three days before crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. again, ten days the next time, 30 days, 100 days, etc.” That per-URL recrawl cadence is really a crawl frequencyCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. question — I cover the mechanics there — but the underlying force is demand, and specifically staleness.
Perceived inventory (the one you control most)
This is the headline lever. Google: “Without guidance from you, Google tries to crawl all or most of the URLs that it knows about on your site. If many of these URLs are duplicates, or you don’t want them crawled for some other reason (removed, unimportant, and so on), this wastes a lot of Google crawling time on your site. This is the factor that you can positively control the most.”
The subtle part is what it does to demand, not just capacity. It’s easy to think of junk URLs as “wasting crawl budget” — spending fetches on copies instead of new content. True. But there’s a demand-side effect too: a site whose knowable inventory is mostly low-value duplicates and parameter sprawl looks, to Google, like a lower-value site to crawl. Cutting perceived inventory doesn’t just free capacity; over time it concentrates demand on the URLs that deserve it. Faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals., session IDs, infinite calendar spaces, and other spider trapsA spider trap (also called a crawler trap) is a site structure that generates an effectively infinite number of URLs — from faceted filters, calendars, session IDs, or redirect loops — so crawlers waste their budget on low-value, near-duplicate pages instead of your real content. are the classic inventory inflators — and the classic demand suppressors.
Site moves and other demand spikes
One demand driver isn’t about any single URL: “Additionally, site-wide events like site moves may trigger an increase in crawl demand in order to reprocess the content under the new URLs.” If you do a domain migration or a big replatform and notice Googlebot hitting you far harder than usual for a few weeks, that’s expected — Google has to re-fetch and reprocess everything under the new addresses. It’s a temporary spike, not a new baseline, and it’s a demand event competitor guides almost never mention despite it being verbatim in Google’s own documentation.
How demand and host load interact: queue order vs. capacity gate
Here’s a mental model that makes the whole topic click — and I want to be upfront that it’s my synthesis, not something Google documents as a literal algorithm. It’s built on a Gary Illyes Q&A, quoted here through Search Engine Roundtable’s coverage: host load “sets a bucket of URLs in importance order and GoogleBot will crawl in that order based on the schedule the host load decided. If Google thinks your server can handle it, it will crawl the whole bucket, if not, it will stop.” Notably, per that same Q&A, host load tracks the importance of your pages — not the raw number of URLs you have or how many you want crawled. That page blocks automated fetching, so I haven’t been able to re-confirm the exact wording directly against the live source this pass — treat it as a well-corroborated paraphrase, not a verbatim primary-source quote.
Read that carefully and a relationship falls out — again, this is how I connect the pieces, not a mechanism Google has spelled out end to end:
- Demand sets the order. The “bucket of URLs in importance order” is crawl demand — popularity and staleness deciding which URLs sit at the top of the queue.
- Capacity sets how far Google gets. Host load / crawl rate decides how deep into that ordered bucket Googlebot actually crawls on a given day. “If your server can handle it, it crawls the whole bucket; if not, it stops.”
So the two aren’t just multiplied together — they play different roles. A healthy, fast server doesn’t manufacture demand (it only raises the ceiling on how much of your existing demand gets realized), and high demand can still be capacity-throttled (a slow or error-prone server stops Googlebot part-way down the queue no matter how much it wants to crawl). This is why “I bought a faster server and Google still isn’t crawling my new pages” is such a common, frustrating result: capacity was never the constraint — demand was.
You can’t set demand directly — but the scheduler listens
Short answer: you don’t set demand directly. You earn it, indirectly, through real links and real quality improvements that show up in indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. signals — nothing else moves it.
There is no “crawl more” request for demand any more than there is for rate. But demand is dynamic, and Google has been unusually candid about how it moves. Gary Illyes: “If you want to increase how much we crawl, then you somehow have to convince search that your stuff is worth fetching, which is basically what the scheduler is listening to.” And the feedback loop is real-time-ish: “Scheduling is very dynamic. As soon as we get the signals back from search indexing that the quality of the content has increased across this many URLs, we would just start turning up demand.” The flip side too: “If search demand goes down, then that also correlates to the crawl limit going down.” (I re-checked these three lines against Search Engine Journal’s coverage this pass and they match verbatim; I still haven’t tracked down Google’s own podcast audio/transcript to confirm them as a primary source, so treat them as well-corroborated secondary quotes.)
That reframes “how do I increase crawl demand” away from tricks. Fake lastmod
timestamps, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. pings, and publishing volume don’t convince the scheduler.
The two things that do are the two hard things: real popularity (links) and
real quality improvements that show up in indexing signals and feed back into
the scheduler. Everything else is theater.
Google is trying to crawl less, not more
The single freshest angle here, and one older guides miss entirely: Google’s own stated goal is to reduce total crawl volume, not grow it. In an April 2024 LinkedIn post, Illyes wrote: “My mission this year is to figure out how to crawl even less, and have fewer bytes on wire.” He pushed back on the idea that Google had slashed crawling — “we’re crawling roughly as much as before, however scheduling got more intelligent” — and framed the goal as a shared win: “Decreasing crawling without sacrificing crawl-quality would benefit everyone.” The mechanisms he pointed at were better caching, cache sharing across user-agents, and fewer bytes transferred — not “crawl my site more.”
The takeaway: crawl demand was never something to maximize. Google is actively optimizing for less total crawling with equal or better crawl quality, routing the demand it does have toward URLs more likely to deserve it. Your goal isn’t more demand — it’s correct prioritization of the demand you’ve earned.
Do you even have a demand problem?
Most sites don’t, and shouldn’t spend a minute on this. Google’s own de-escalation applies squarely to demand: “If your site doesn’t have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don’t need to read this guide.” John Mueller has been similarly blunt about scale — per Search Engine Roundtable’s coverage of his tweet, 100k URLs is usually not enough to affect crawl budget, since that works out to well under one crawl per minute over three months.
If you are big enough to care, here’s the diagnostic that separates a demand problem from a capacity one — with a caveat up front: it produces a hypothesis to test, not a diagnosis. Pull up GSC’s Crawl Stats reportA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root). and your server logs. If host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity. is healthy and average response time is fine but a set of URLs is barely getting crawled — and they’re stuck in “Discovered – currently not indexed” — that pattern is evidence pointing toward demand, not proof of it. Crawl Stats shows crawl activity (a capacity-side view), not a demand score — there is no public per-site “crawl demand score” anywhere, so you’re always inferring demand from activity plus indexing status, never reading it off a dashboard.
Before you act on “it’s a demand problem,” rule out the other things that produce the same healthy-server-but-not-crawled symptom: Google may not have discovered the URLs yet (no path in, no sitemap entry), renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. may be hiding the content Googlebot needs to see, canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. may point Google somewhere else entirely, real quality issues (thin, duplicate, low-value) can get a page crawled but deliberately left out of the index, and Google’s own indexing selection can sit a page out even when crawling and quality are both fine. Only after those are checked and don’t explain it does “low demand” become the working explanation — and even then, treat it as the best-supported hypothesis, not a confirmed cause. Faster hardware won’t fix a genuine demand problem; the fix is on the demand inputs: links to those pages, genuine reasons to recrawl them, and less junk inventory drowning them out.
For ground-truth, per-URL crawl data, log analysis is the answer. I’ll flag one current tool I can speak to firsthand in the Tools tab.
Crawl demand vs. rate vs. budget vs. frequency
Keep the family straight:
- Crawl demand — how much Google wants to crawl (popularity + staleness + perceived inventory). This page.
- Crawl rate — how fast it can (capacity / host load). Its own page.
- Crawl budget — the two together: “the number of URLs Googlebot can and wants to crawl.”
- Crawl frequency — how often a given URL gets recrawled, which is a demand output (mostly staleness).
Bing doesn’t use the term “crawl demand” — it reframes the whole thing as crawl efficiency: Fabrice Canel defines it as “how often we crawl and discover new and fresh content per page crawled,” and Bing’s philosophy is inventory-reduction-first, the demand-side equivalent of perceived inventory. IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. is Bing’s way of signaling demand-relevant change events instead of waiting for the scheduler to infer staleness — but note Google does not use IndexNow, so it won’t move Google’s demand.
AI summary
A condensed take on the Advanced version:
- Crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side. = the “want” side of crawl budgetThe number of URLs an engine will crawl in a timeframe.; crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor./capacity is the “can” side. Budget is “the number of URLs GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. can and wants to crawl.”
- Significant demand factors — not a closed formula: popularity (links / PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems.), staleness (how often a page changes), and perceived inventory (how many URLs Google thinks exist — junk included; “the factor you can positively control the most”), plus site size, update frequency, page quality, and comparative relevance.
- Site moves spike demand temporarily while Google reprocesses content under new URLs.
- Demand vs. capacity model (my synthesis, not a documented Google algorithm): demand sets the priority order of URLs; host-load capacity decides how far down that queue GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. gets. A fast server doesn’t create demand; high demand can still be capacity-throttled.
- You can’t set demand directly. The scheduler “turns up demand” when quality
signals from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. improve — so the only real levers are earning links and
making genuine quality/freshness improvements. Fake
lastmod, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. pings, and publishing volume don’t move it. - Google is trying to crawl less, not more (Illyes: “crawl even less… fewer bytes on wire”), routing demand more precisely. The goal is correct prioritization, not volume.
- Diagnose demand vs. capacity: healthy host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity. + low crawl volume + stuck in “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” is a demand hypothesis, not proof — rule out discovery, renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it., and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.-selection causes first. No public “crawl demand score” exists.
- Bing reframes it as crawl efficiency; IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. signals change to Bing (not Google). Most sites never need to manage this.
Official documentation
Primary-source documentation from the search engines.
- Optimize your crawl budget — the source that defines crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side., its three inputs (popularity, staleness, perceived inventory), and site-move demand spikes.
- Crawl Budget Management — the capacity/demand split, with the crawl-capacity mechanics that demand runs into.
- What crawl budget means for Googlebot (2017) — Gary Illyes’ original post defining crawl budgetThe number of URLs an engine will crawl in a timeframe. as what GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. “can and wants to crawl,” and the low-value-URL categories that suppress effective demand.
- Myths and facts about crawling — confirms crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor. isn’t a ranking signal and that server health, not desire, sets the capacity ceiling.
- Crawling December series (2024) — GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer., HTTP cachingCaching stores a copy of a page or resource — in a browser, a CDN edge node, or a search crawler's own cache — so it can be served again without regenerating or re-downloading it. It isn't a direct ranking factor, but it feeds page speed and crawl efficiency., faceted nav, and the efficiency thinking behind crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. less.
Bing / Microsoft
- bingbot Series: Maximizing Crawl Efficiency — Bing’s “crawl efficiency” framing, its demand-side analogue.
- bingbot Series: Optimizing Crawl Frequency — Bing’s take on recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. cadence driven by how often content changes (its “staleness” analogue).
- IndexNow / indexnow.org — signal changed URLs to Bing and others (not Google) instead of waiting on inferred staleness.
Quotes from the source
On-the-record statements from Google and Bing. Each link is a deep link that jumps to the quoted passage on the source page.
Google — the demand definition and its inputs
- “The amount of time and resources that Google devotes to crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. a site is commonly called the site’s crawl budgetThe number of URLs an engine will crawl in a timeframe. and it’s determined by two main elements: crawl capacity limitCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor. and crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side..” — Large Site Owner’s Guide to Managing Crawl BudgetThe number of URLs an engine will crawl in a timeframe.. Jump to quote
- “Each crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. has its own ‘demand’ when it comes to crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. the web.” Jump to quote
- “URLs that are more popular on the Internet tend to be crawled more often to keep them fresher in our systems.” (popularity) Jump to quote
- “Our systems want to recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. documents frequently enough to pick up any changes.” (staleness) Jump to quote
Google — perceived inventory and site moves
- “If many of these URLs are duplicates, or you don’t want them crawled for some other reason (removed, unimportant, and so on), this wastes a lot of Google crawling time on your site. This is the factor that you can positively control the most.” Jump to quote
- “Additionally, site-wide events like site moves may trigger an increase in crawl demand in order to reprocess the content under the new URLs.” Jump to quote
- “If your site doesn’t have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don’t need to read this guide.” Jump to quote
Gary Illyes, Google — how demand actually moves (via Search Engine Journal’s coverage of his podcast appearance)
- “If you want to increase how much we crawl, then you somehow have to convince search that your stuff is worth fetching, which is basically what the scheduler is listening to.” Read the coverage
- “Scheduling is very dynamic. As soon as we get the signals back from search indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. that the quality of the content has increased across this many URLs, we would just start turning up demand.” Read the coverage
- “If search demand goes down, then that also correlates to the crawl limit going down.” Read the coverage
Gary Illyes, Google — the “crawl even less” mission (LinkedIn, April 2024 — primary source, verified verbatim)
- “My mission this year is to figure out how to crawl even less, and have fewer bytes on wire.” Read the post
- “we’re crawling roughly as much as before, however scheduling got more intelligent” and “Decreasing crawling without sacrificing crawl-quality would benefit everyone.” Read the post
Bing / Microsoft — crawl efficiency (the demand-side analogue)
- “The crawl efficiency is how often we crawl and discover new and fresh contentContent freshness is how recent or up-to-date a page is — by its original publish date, its last substantive revision, or the currency of the facts inside it. It only helps rankings when the query itself benefits from recent results (Query Deserves Freshness), and cosmetic date changes with no real update don't count. per page crawled.” — Fabrice Canel. Jump to quote
Is it a crawl-demand problem — and should you care?
Work top-down. Most sites exit early with “leave it alone.”
Crawl demand myths and mistakes
The conflations that send people down the wrong path.
“Crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side. and crawl budgetThe number of URLs an engine will crawl in a timeframe. are the same thing.” Why it’s wrong: demand is one of two components; budget is capacity × demand together — “the number of URLs GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. can and wants to crawl.” Do instead: keep the vocabulary straight. Budget is the outcome; demand and capacity are its two inputs. A budget problem is always really a demand problem, a capacity problem, or both.
“A faster server increases crawl demand.” Why it’s wrong: server speed raises the capacity ceiling only. It lets Google realize more of the demand you already have; it doesn’t make Google want to crawl more. Demand is set by popularity, staleness, and perceived inventory — none of which your hardware touches. Do instead: if a healthy server still isn’t getting your pages crawled, stop buying hardware and work the demand inputs (links, freshness, less junk inventory).
“Publishing more often increases demand.” Why it’s wrong: publishing volume without genuine importance or change doesn’t convince the scheduler. Ten thin posts a day that nobody links to move nothing. Do instead: publish things that earn links and genuinely change/improve — that’s what feeds the quality signals the scheduler listens to. (This is a crawl frequency point too; more detail lives there.)
“Blocking junk URLs in robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. instantly redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. that demand to my good pages.” Why it’s wrong: cutting perceived inventory helps demand concentrate over time as Google reassesses the site — but it isn’t an instant reallocation. Google won’t automatically pour freed-up crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. onto your good pages the moment you disallow the junk. Do instead: reduce junk inventory for the long-term concentration benefit, and be patient. It’s a trend, not a switch.
“There’s a crawl-demand score I can check in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results..” Why it’s wrong: no public per-site demand score exists. Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root). shows crawl activity — a capacity-side report — not a demand metric. Do instead: infer demand from activity plus indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. status (e.g., healthy host + low crawl volume + “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” = a demand signal). Use log data for per-URL ground truth.
“IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. or sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. pings raise Google’s crawl demand.”
Why it’s wrong: Google doesn’t use IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it., and it ignores changefreq/priority
in sitemaps. Pinging Google doesn’t make it want to crawl you more.
Do instead: use IndexNow for Bing and other participating engines. For Google, an
accurate lastmod helps it schedule, but the demand levers remain links, freshness,
and inventory.
“More crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. is always better — for me and for Google.” Why it’s wrong: getting crawled more doesn’t raise rankings, and Google itself is trying to crawl less (“My mission this year is to figure out how to crawl even less, and have fewer bytes on wire”) while routing demand more precisely. Do instead: aim for correct prioritization, not raw volume. The goal is the demand you have landing on the right URLs.
Runbook: “Google isn’t crawling my important pages, and my server is fine”
A linear path for a large-site owner who suspects a demand problem. Stop as soon as a step resolves it.
-
Confirm you’re big enough to care. If pages are usually crawled the day they’re published, or the site is normal-sized, stop — you don’t have a crawl-demand problem. This runbook is for 1M+-page or fast-changing sites, or sites with a large “Discovered – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” pile.
-
Rule out capacity first. Open GSC Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root).. Check host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity. and average response time over the last 90 days. If you see
5xx/timeout spikes or climbing response times, this is a capacity problem — fix server health (see the crawl rate page) and re-run this runbook afterward. If the server looks healthy, continue. -
Pull per-URL ground truth from logs. Get server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. (or a botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.-analytics feed) for the affected URL set. Confirm the pattern: real GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. hits are sparse or absent on pages you care about, while junk/parameter URLs are eating hits. Sparse hits on healthy-server pages is a demand hypothesis, not a confirmed cause yet — before you commit to “demand,” rule out the look-alikes: Google not having discovered the URL, a renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. failure hiding the content, canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. pointing away from the page, real quality/duplication issues, and Google’s own indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.-selection choices. Check Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance.’s URL Inspection and rendered-HTML view for each.
-
Check for perceived-inventory inflation. Count how many low-value URLs Google could be discovering: faceted-navigation combinations, session IDs, sort/filter parameters, calendar/infinite spaces, on-site duplicates. If these dwarf your real pages, demand is likely being spread across junk.
-
Cut the junk inventory. Reduce the low-value URL space at the source (parameter handling,
robots.txtdisallow of infinite spaces, fixing spider traps, consolidating duplicates via canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it.). Expect concentration over time, not an instant reallocation. -
Work the popularity input. Add internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. from strong pages to the under-crawled ones (kill orphans), and pursue external links. Popularity is a primary demand driver.
-
Work the quality/freshness input. Genuinely improve and update the pages so the signals that come back from indexing tell the scheduler to “turn up demand.” Accurate
lastmodhelps Google schedule; fake freshness does not. -
Give it time, then re-measure. Re-check Crawl Stats and logs after Google has had time to reassess. Demand shifts are gradual. If the pages are now crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexing, done. If not, revisit whether the pages are genuinely worth crawling — sometimes the honest answer is that they aren’t, and thin pages shouldn’t be forced into the index.
Crawl-demand checklist
Diagnose: is it demand or capacity?
- Confirmed the site is actually large/fast-changing enough to care (else stop).
- GSC Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root).: host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity. healthy, average response time stable, no
5xx/timeout spikes. - Server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. (or botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. analytics) reviewed for real per-URL GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. hits.
- Pattern identified: healthy server + low crawl volume on good pages + “Discovered – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” = a demand hypothesis, not proof — and not a capacity issue.
- Ruled out the look-alikes before blaming demand: discovery, renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it., quality/duplication, and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.-selection issues.
Work the demand inputs (there’s no direct dial)
- Popularity: important pages have internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. pointing at them (no orphans); external link-building underway where it matters.
- Staleness/freshness: pages that should be recrawled often are genuinely
updated;
lastmodis accurate (not faked). - Perceived inventory: faceted-nav, parameter, session-ID, and infinite URL spaces are controlled; spider trapsA spider trap (also called a crawler trap) is a site structure that generates an effectively infinite number of URLs — from faceted filters, calendars, session IDs, or redirect loops — so crawlers waste their budget on low-value, near-duplicate pages instead of your real content. fixed; duplicates consolidated.
Reality checks
- Not expecting a faster server to raise demand (it only raises the capacity ceiling).
- Not expecting robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. disallow of junk to instantly reroute demand to good pages (it concentrates over time).
- Not relying on IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it./sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. pings to move Google’s demand (Google ignores
IndexNow and
changefreq/priority). - After a site move, treating the temporary crawl spike as expected, not a problem.
- Remembering more crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. isn’t the goal — correct prioritization is; crawl volume isn’t a ranking factor.
Crawl demand — cheat sheet
The two sides of crawl budgetThe number of URLs an engine will crawl in a timeframe.
| Crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side. (this page) | Crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor. / capacity | |
|---|---|---|
| What it is | How much Google wants to crawl | How fast it can |
| Drivers | Popularity, staleness, perceived inventory (+ site moves) | Server health / host load |
| Role in the queue | Sets the order of URLs (importance) | Sets how far down Google gets |
| Your lever | Links, real freshness, cut junk inventory | Faster/healthier server |
| Direct dial? | No | No |
The three most-actionable demand inputs (Google names these plus site size, update frequency, page quality, and comparative relevance as significant general factors — not a closed formula)
| Input | What raises it | What it isn’t |
|---|---|---|
| Popularity | More links / PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. (internal + external) | Not publishing volume |
| Staleness | Genuine, frequent content change | Not fake lastmod |
| Perceived inventory | Fewer junk/duplicate URLs (cutting it helps) | Not a faster server |
Fast facts
- Crawl budgetThe number of URLs an engine will crawl in a timeframe. = “the number of URLs GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. can and wants to crawl.”
- Perceived inventory is “the factor you can positively control the most.”
- Site moves temporarily spike demand (reprocessing under new URLs).
- Demand sets the priority order; host-load capacity decides how deep Google crawls — a healthy server doesn’t create demand.
- The scheduler “turns up demand” when indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. quality signals improve — that’s the only real lever besides links.
- Google’s goal is to crawl less overall, not more (Illyes, 2024). Crawl volume is not a ranking factor.
- No public “crawl demand score” — Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root). shows activity (capacity side).
- Bing has no “crawl demand” term — it’s crawl efficiency; IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. signals change to Bing, not Google.
Patrick's relevant free tools
- Log File Analyzer — Drop a server access log and see crawl budget by bot and section, status-code waste, an AI-vs-search breakdown, and a spoofer report that names impostors faking a crawler user-agent. Parses nginx, Apache, IIS/W3C, and JSON logs entirely in your browser — nothing is uploaded.
- Googlebot Verifier — Check whether an IP claiming to be Googlebot, Bingbot, GPTBot, ClaudeBot, or another crawler is genuine — published IP ranges plus forward-confirmed reverse DNS, with the real network owner named for spoofers. IPs are checked in memory and never stored.
- Raw vs. Rendered HTML Checker — See what's in your page's initial HTML versus after JavaScript runs — headless-Chrome rendering only when the page actually needs it, a rendering-strategy verdict (SSR / prerendered / CSR / hybrid), ~15 calibrated JavaScript-SEO checks (noindex, canonicals, robots.txt blocking, links, soft 404s), a side-by-side raw-vs-rendered diff, and shareable reports.
Tools for diagnosing crawl demand
There’s no “demand meter,” so diagnosing demand means reading crawl activity and comparing it against what you know about the pages.
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. — Crawl Stats reportA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root). — total crawl requests over time, host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity., average response time, and breakdowns by response code, file type, purpose, and GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. type. Read it as a capacity view: healthy host statusThe at-a-glance availability indicator at the top of Google Search Console's Crawl Stats report. It shows whether Googlebot hit significant problems reaching your site in the last 90 days across three checks — robots.txt fetching, DNS resolution, and server connectivity. + low crawl volume on good pages is your demand tell.
- GSC — Page indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. — the “Discovered – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” bucket is the classic footprint of a demand shortfall (Google knows the URLs, doesn’t care enough to crawl them yet).
- URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. (GSC) — check when a specific URL was last crawled and whether it’s indexed; useful to confirm a single page’s demand story.
- Server log file analysisLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. — the ground truth for real, per-URL GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. hits: which URLs bots actually fetch, how often, and where crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. is being wasted on junk inventory. (See log file analysisLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened..)
- Ahrefs Bot Analytics — a tool I can speak to firsthand. As I described it when we launched it, “Have y’all checked out Bot Analytics in Ahrefs yet? We released a new tool that shows how bots crawl your website. Bot Analytics collects data server-side via Cloudflare integration.” It shows every bot crawling your site and the pages they hit across 12 categories — exactly the per-URL, per-bot ground truth you need to check Google’s “perceived inventory” and “popularity” story against reality. Ahrefs’ own framing of the problem it solves: “uncontrolled bot traffic wastes crawl budgetThe number of URLs an engine will crawl in a timeframe. — bots crawling 404 pages or low-value URLs aren’t crawling the pages you need indexed,” and it cites the estimate that over half of all crawler traffic is wasted effort.
- Ahrefs Site Audit / Screaming Frog SEO Spider — simulate a crawl to surface the parameter sprawl, duplicates, and trap-like patterns that inflate perceived inventory and suppress demand.
The importance × change × inventory framework
Use three questions to explain changes in crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side.:
- Importance: did internal or external signals make the URL more or less important?
- Change: did the page change meaningfully, and did truthful sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. signals communicate that?
- Inventory: did the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.’s known set of duplicates, parameters, or low-value URLs expand?
Host health is the ceiling, not a fourth demand input. If logs show errors or timeouts, diagnose crawl capacityThe number of URLs an engine will crawl in a timeframe. separately. If the server is healthy but valuable URLs lose crawl share, work through importance, change, and inventory in that order.
Compare crawl share by directory
This shell pipeline summarizes verified crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. requests by the first URL-path directory in a common access log:
awk 'BEGIN{IGNORECASE=1} /Googlebot/ {split($7,p,"/"); print "/" p[2] "/"}' access.log | sort | uniq -c | sort -nrOn PowerShell:
Select-String .\access.log -Pattern 'Googlebot' | ForEach-Object { if ($_.Line -match '"(?:GET|HEAD)\s+https?://[^/]+/([^/?\s]*)|"(?:GET|HEAD)\s+/([^/?\s]*)') { '/' + (($Matches[1],$Matches[2] | Where-Object { $_ })[0]) + '/' } } | Group-Object | Sort-Object Count -DescendingCompare the same-length windows before and after a change. A directory gaining share is a clue about scheduler allocation, not proof of higher quality or rankings.
Metrics for crawl demand
Valuable-template crawl share
Metric: crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. requests to important templates divided by verified crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. requests. What it tells you: whether demand is reaching the inventory you care about. How to pull it: classify access-log URLs by template. Benchmark / realistic range: define the desired mix from your own valuable inventory and update cadence; there is no universal percentage. Cadence: weekly for large changing sites, monthly otherwise.
Recrawl lag after meaningful changes
Metric: time from a real page update to the next verified crawler fetch. What it tells you: whether the scheduler recognizes the page’s importance and change pattern. How to pull it: join deployment or content timestamps to access logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened.. Benchmark / realistic range: baseline by template; news and stable reference pages should not share a target. Cadence: monthly.
Low-value inventory share
Metric: known and crawled parameter, duplicate, empty, or soft-404 URLs relative to useful URLs. What it tells you: whether perceived inventory is diluting attention. How to pull it: combine crawl exports, sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., indexability rules, and logs. Benchmark / realistic range: trend downward from the site’s baseline without blocking required resources. Cadence: monthly and after faceted-navigation or platform changes.
Resources worth your time
My related writing
- When Should You Worry About Crawl Budget? — where I frame crawl budgetThe number of URLs an engine will crawl in a timeframe. as demand (“how many pages a search engine wants to crawl”) plus rate, and cover popularity, the staleness backoff, and who actually needs to care.
- What Is Googlebot & How Does It Work? — how GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. decides what and how much to crawl.
- The Beginner’s Guide to Technical SEO — where crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and crawl budgetThe number of URLs an engine will crawl in a timeframe. fit in the bigger picture.
My speaking
- How Search Works (SlideShare) — my walkthrough of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., including the demand factors (PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems., freshness, time since last crawl, major site changes) and the separate capacity/host-load slide. (Standing disclaimer: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From around the industry
- Google’s Crawling Priorities: Insights From Gary Illyes (Search Engine Journal) — the “convince search your stuff is worth fetching,” “turning up demand,” and “search demand goes down” quotes on how demand actually moves.
- Gary Illyes on crawling even less (LinkedIn, April 2024) — the primary source for “crawl even less… fewer bytes on wire” and “scheduling got more intelligent.”
- Google Has Two Types Of Crawling: Discovery & Refresh (Search Engine Journal) — John Mueller on discovery vs. refresh crawls; refresh cadence is a pure demand output.
- Google’s Gary Illyes On Crawl Budget, Scheduling & Host Load (Search Engine Roundtable) — the “bucket of URLs in importance order” host-load framing (paraphrased in this article; the page blocks automated fetch — confirm against the live page).
- Google: 100k URLs Won’t Impact Crawl Budget (Search Engine Roundtable) — John Mueller’s scale gut-check (paraphrased here; confirm against the live page).
- What Is Crawl Budget? How It Works + Optimization Tips (Search Engine Land) — a solid crawl-budget mega-guide; useful background on the three demand factors.
- Ahrefs Bot Analytics — the product page with the “uncontrolled botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. traffic wastes crawl budget” and “over half… wasted effort” framing for diagnosing demand vs. capacity from real bot data.
Test yourself: Crawl Demand
Five quick questions on the “want” side of crawl budgetThe number of URLs an engine will crawl in a timeframe.. Pick an answer for each, then check.
Crawl Demand
Crawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side.
Related: Crawl Budget, Crawl rate
Crawl Demand
Crawl demand is one of the two components of crawl budgetThe number of URLs an engine will crawl in a timeframe. — the “want” side. Where crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor. (crawl capacity) is how fast a search engine is able to fetch pages from your server, crawl demand is how much it wants to crawl them in the first place. Google puts it simply: a site’s crawl budgetThe number of URLs an engine will crawl in a timeframe. “is determined by two main elements: crawl capacity limit and crawl demand.”
Google calls out three inputs it explains in the most depth — not an exhaustive formula; site size, update frequency, page quality, and how a site compares to similar ones also factor in:
- Popularity — URLs with more links and PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. tend to get crawled more often to keep them fresh.
- Staleness — Google’s systems try to recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. documents often enough to pick up changes, so a page that changes a lot earns more demand than one that never does.
- Perceived inventory — how many URLs Google believes exist on your site. Junk, duplicate, and low-value URLs inflate perceived inventory and waste crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor.; Google calls this “the factor that you can positively control the most.”
Site-wide events like a domain migration temporarily spike crawl demand because Google has to reprocess your content under the new URLs.
You cannot set crawl demand directly — there’s no dial for it, just as there’s no button to force a higher crawl rateCrawl rate is how fast a search engine crawler fetches pages from your site — the number of simultaneous requests it makes and the delay between them. Google sets it automatically based on your server's health; it's the supply side of crawl budget, not a ranking factor.. You move it indirectly by earning real popularity (links), keeping content genuinely fresh, and cutting the perceived inventory of low-value URLs so the demand you do have concentrates on the pages that matter.
Related: Crawl Budget, Crawl rate
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 17, 2026.
Editorial summary and recorded change details.Summary
Qualified the three-factor demand formula, labeled the queue-order model as Patrick's synthesis, and reframed the demand-vs-capacity diagnostic as a hypothesis with alternative causes to rule out first.
Change details
-
The 'three demand inputs' framing now reads as significant general factors Google names alongside site size, update frequency, page quality, and comparative relevance, not a closed three-item formula, across the beginner/advanced lenses, TL;DR, ai-summary, cheat sheet, glossary, and quiz.
-
The queue-order-vs-capacity-gate model is now explicitly labeled as Patrick's synthesis (not a documented Google algorithm), and its host-load source is flagged as an unverified secondary paraphrase.
-
The 'Do you even have a demand problem?' diagnostic, the runbook, checklist, and decision tree now frame a healthy-server-plus-low-crawl-volume pattern as a demand hypothesis to test, listing discovery, rendering, canonicalization, quality, and indexing-selection as alternative causes to rule out first.