Discovery

How search engines discover URLs — internal links, backlinks, sitemaps, RSS/Atom, and push protocols like IndexNow and the Indexing API. The hub for everything discovery-related.

First published: Jun 22, 2026 · Last updated: Jul 25, 2026 · Advanced
demand #49 in How Search Works#257 in Technical SEO#347 on the site
1 evidence signal on this page

Discovery is the find step before search ever fetches a page: the engine has to learn a URL exists before it can crawl, index, or rank it. It's this hub's teaching name for the opening move of what Google's official model calls the crawling stage. URLs are discovered by pull (internal links, backlinks, sitemaps, RSS/Atom) and by push (IndexNow, the Indexing API, lastmod, WebSub, URL Inspection). Google calls it 'URL discovery,' and it's distinct from crawling and indexing — which is exactly why 'Discovered – currently not indexed' means found-but-not-yet-crawled, a status with several possible causes rather than one deterministic trigger. Links and sitemaps are the two routes Google names directly, orphan pages are the classic failure, and submitting a URL never guarantees indexing. This hub maps it all and points you to the sitemap deep dives.

TL;DR — Discovery is this hub’s name for the first move in search (discover → crawl → indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. → serve): the engine has to learn a URL exists before it fetches anything. Google’s own official model nests this inside its three-stage crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. stage rather than naming a separate fourth stage — same sequence, different vocabulary. URLs are discovered by pull (internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., backlinks, sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., RSSAn RSS or Atom feed is an XML file listing a site's most recently published or updated URLs. Search engines accept it as a sitemap-style discovery signal for fresh content — not a replacement for a full XML sitemap./Atom) and push (IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it., the Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. lastmod, WebSubWebSub (formerly PubSubHubbub) is a push notification protocol for RSS/Atom feeds that lets a site instantly broadcast new or updated content to subscribing search engines instead of waiting to be re-crawled., URL Inspection). Google calls the find step “URL discoveryURL discovery is how search engines find URLs to crawl — by pull (following links and reading sitemaps) and by push (you notify them via IndexNow, the Indexing API, or WebSub). It's the find step that comes before a page is ever fetched.,” and it’s distinct from crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — which is why “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” means found-but-not-yet-crawled, a state with several possible causes rather than one guaranteed diagnosis. Links and sitemaps are the two routes Google names directly, orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. are the canonical failure, IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. is Bing/others not Google, the Indexing API is scope-limited to JobPosting/BroadcastEvent, and submitting a URL never guarantees indexing.

Evidence for this claim Google discovers URLs through links, sitemaps, and other previously known signals before crawling and possible indexing. Scope: Current Google crawling and indexing overview. Confidence: high · Verified: Google Search Central: How Search works Evidence for this claim Sitemaps and indexing requests can aid discovery but do not guarantee crawling, indexing, or serving in results. Scope: Current Google sitemap and indexing-request behavior. Confidence: high · Verified: Google Search Central: Learn about sitemaps

What URL discovery is

Discovery is the find step. Before a page can be crawled, indexed, or served, the engine has to know it exists and add it to its list of known pages. Google names this explicitly: “This process is called ‘URL discovery’.” In Google’s own words, “Some pages are known because Google has already visited them. Other pages are discovered when Google extracts a link from a known page to a new page… Still other pages are discovered when you submit a list of pages (a sitemap) for Google to crawl.”

That single passage is the whole hub in miniature: pages are found through links (internal and external) and through sitemaps. Everything else is a variation on those two ideas — or a way to push a notification instead of waiting to be pulled.

A note on the stage model. This hub treats discover → crawl → index → serve as four distinct, named steps, because separating “found” from “fetched” from “stored” is the clearest way to diagnose a stuck page. Google’s own official model names three stages — crawling, indexing, serving — and defines URL discovery as the opening move within crawling, not as its own fourth stage. The sequence of events is identical either way; this hub just gives the find step its own name and its own section because it behaves differently enough (and gets misdiagnosed often enough) to deserve one.

Discovery vs crawling vs indexing

Keep these three ideas separate and most discovery confusion disappears — using our four-step teaching breakdown of Google’s three official stages:

  • Discovery — the engine learns a URL exists. Nothing has been fetched yet.
  • Crawling — the engine downloads the discovered URL (and renders it). That’s the CrawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. hub’s job.
  • Indexing — the engine analyzes and stores the crawled page so it can be served.

Google frames the back half of this as three stages: “Google Search works in three stages, and not all pages make it through each stage” — crawling, indexing, and serving. Discovery is the front door of stage one. And none of these steps is a ranking factor on its own; they’re prerequisites to being eligible to rank. A URL that’s never discovered is never crawled, indexed, or served — full stop.

How search engines discover URLs: pull vs push

The cleanest mental model is pull vs push. Pull is the engine finding URLs on its own; push is you notifying it.

Pull — the engine finds it on its own

Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. and backlinks. Google: “Other pages are discovered when Google extracts a link from a known page to a new page: for example, a hub page, such as a category page, links to a new blog post.” Google’s own docs also say a properly linked site can usually get most of its pages discovered this way, with new sites that have few external links and large sites with unlinked pages at greater risk of gaps — links aren’t the only route in, but they’re the one Google names first and the one that scales automatically as you publish. That’s why internal linkingLinks between pages on the same site. is one of the highest-leverage discovery levers you have, and why orphan pages struggle (more below). Backlinks from other sites work identically — an external link from a page the engine already knows surfaces your new URL.

Sitemaps. Google: “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” That caveat is the whole point — a sitemap is a discovery aid, not an indexing guarantee. The full mechanics live in the nested sitemaps deep dives (XML sitemap, sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file., image sitemapAn image sitemap is a sitemap (or an extension added to an existing sitemap) that lists the images on your pages using Google's image namespace, helping Google discover images it might otherwise miss., video sitemapA video sitemap is an XML sitemap (or mRSS feed) that uses the video:video extension to tell Google about videos hosted on your pages — the landing page, a thumbnail, a title, a description, and a link to the video file or player. It helps Google discover and surface video content, but it doesn't guarantee indexing.).

RSS / Atom feeds. Google accepts feeds as a sitemap format — “Google accepts RSS 2.0 and Atom 1.0 feeds.” The catch is that a feed only surfaces your recently changed URLs, so it complements a full sitemap rather than replacing it.

Push — you notify the engine something changed

WebSubWebSub (formerly PubSubHubbub) is a push notification protocol for RSS/Atom feeds that lets a site instantly broadcast new or updated content to subscribing search engines instead of waiting to be re-crawled. (PubSubHubbub). For RSS/Atom feeds, Google supports WebSub: “If you use Atom or RSS, you can use WebSub to broadcast your changes to search engines, including Google.” Instead of waiting to be re-pulled, your feed broadcasts the change.

Sitemap lastmod. A freshness signal that tells engines a URL changed and may be worth re-crawling. Bing leans on it harder than Google does — Bing has said the lastmod field remains a key signal that helps it prioritize URLs for recrawling and reindexing. (Detail lives in the sitemaps cluster.)

Google Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. — scope-limited. This is the one that gets misused most. The Indexing API can only be used to crawl pages with either JobPosting or BroadcastEvent embedded in a VideoObject. It is not a general “submit any URL to Google” endpoint, no matter how often it’s pitched that way. If your page isn’t a job posting or a live-stream broadcast event, the Indexing API isn’t your tool.

IndexNow — Bing and others, not Google. IndexNow is a push protocol: it notifies enabled search engines the instant a URL is added, updated, or deleted, and per its own FAQ, “Search engines adopting the IndexNow protocol agree that submitted URLs will be automatically shared with all other participating search engines.” As of this writing (per indexnow.org’s current documentation, checked July 2026), the participating engines are Microsoft Bing, Yandex, Naver, Seznam.cz, and Yep — check the live IndexNow docs for the current list, since it can change. Google is not on that list. There’s no Google help page that says “we don’t support IndexNow” — Google’s non-participation is inferred from its absence from the current partner list plus Googlers describing it as something they only tested, not from a first-party Google confirmation, so treat it as the best available reading rather than an official statement. And regardless of engine: notifying IndexNow only tells it a URL changed — each participating engine still decides independently whether and when to crawl and index it, the same “notification isn’t a guarantee” rule that applies to sitemaps and every other discovery channel. For Google, you fall back to links, sitemaps, and URL Inspection.

Manual submission — URL Inspection. For one-off Google submission: “To request a crawl of individual URLs, use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version..” But don’t expect repeats to force speed — Google is explicit that “there’s a quota for submitting individual URLs and requesting a recrawl multiple times for the same URL won’t get it crawled any faster.” Bing has its own URL submission channel (up to 10,000 URLs/day), separate from IndexNow.

Pull and push can both make a URL known, but neither route skips the crawl queue or guarantees indexing. Source: /technical-seo/how-search-works/discovery/

Two discovery routes feed one crawl queue. Pull discovery includes following links and reading sitemaps. Push discovery includes IndexNow for participating engines, the scope-limited Google Indexing API, and freshness notifications such as sitemap lastmod, RSS, or WebSub. Google does not use IndexNow for general pages, and its Indexing API supports only JobPosting and qualifying livestream pages.

© Patrick Stox LLC · CC BY 4.0 ·

How Google and Bing each describe discovery

Both engines frame discovery as the front of the pipeline. Bing puts it directly in its definition of crawling — Fabrice Canel: “Crawling is the process by which bingbotBingbot is Microsoft Bing's primary web crawler — the bot that discovers, fetches, and renders pages to build the Bing index. That index also powers Yahoo, DuckDuckGo, Ecosia, and Microsoft Copilot, so Bingbot's reach is far wider than Bing's own search-market share. discovers new and updated documents or content to be added to Bing’s searchable index.”

The scale is worth sitting with, even if the exact figure moves over time. Fabrice Canel has described Bing discovering on the order of tens of billions of normalized URLs it had never seen before, per day — he put it as ”12s of billions … never seen before” in an August 2022 post, a number I haven’t independently re-verified against a current primary source, so treat it as a directional sense of scale rather than a precise, current daily count. Whatever the exact figure is today, discovery operates as a firehose, and the engines deprioritize aggressively — which is the backdrop for “Discovered – currently not indexed” below.

Orphan pages: when discovery fails

An orphan page has no internal links pointing to it. Because links are one of the two discovery routes Google names directly (alongside sitemaps), an orphan can only be found via a sitemap, an external link, a redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., or a canonical/hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. reference — and if none of those exist, it may never be discovered at all.

The fix is internal linking, not push tricks. Link the page from a relevant hub/category page — exactly Google’s “category page links to a new blog post” example. Putting an orphan in your XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. helps Google discover it, but it doesn’t replace the internal-link signal that tells the engine the page matters. Orphans frequently show up as “Discovered – currently not indexed” for precisely this reason: Google found the sitemap entry but deprioritized crawling a page nothing links to.

”Discovered – currently not indexed” in Search Console

What’s certain about this status is the stage boundary — this hub’s job is to fix that boundary in your head, not to fully diagnose it (the dedicated deep diveA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal. does that). Google’s description: the page was found by Google but not crawled yet — Google wanted to crawl the URL but expected it would overload the site, so it rescheduled the crawl, which is why the last-crawl date is empty.

Contrast it with Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.,” which is a different, downstream problem: the page was crawled but not indexed, and it may or may not be indexed in the future. The stage boundary is the lesson most third-party write-ups miss:

  • Discovered – currently not indexed = Google knows the URL but hasn’t fetched it. A discovery/crawl-scheduling state (empty last-crawl date).
  • Crawled – currently not indexed = Google fetched it but chose not to index it. An indexing/quality state.

What causes “Discovered – currently not indexed” — work through it in this order, not as a single confirmed cause:

  1. Crawl scheduling / perceived server-load deprioritization — Google’s own stated reason for the status itself.
  2. The page is orphaned or only weakly internally linked, lowering how much Google wants to spend crawl effort on it.
  3. Site-wide quality or authority signals dragging on overall crawl demandCrawl demand is the 'want' side of crawl budget — how much a search engine wants to crawl a site or URL, driven by popularity, staleness, and perceived inventory (plus temporary spikes from site moves). It's distinct from crawl rate/capacity, the 'can' side. — a site-level pattern, not proof of a defect on this one page.
  4. A new or low-authority site with little crawl demand yet.

None of these is the universal cause — Google documents the crawl-scheduling explanation directly but doesn’t publish a single deterministic reason a given URL sits in this state, so treat the list above as a hypothesis ladder to work through, not a diagnosis to assume. It also doesn’t mean something is broken: John Mueller has said “it’s completely normal that we don’t index everything off of the website,” with the advice being to reconsider overall site quality rather than hunt for a per-page technical issue, and Gary Illyes has pointed out that the vast majority of websites don’t need to think about crawl budgetThe number of URLs an engine will crawl in a timeframe. at all. My own practitioner shorthand, from my indexing guide at Ahrefs (worth a re-read before treating it as gospel, since I haven’t re-verified the live page in this pass): Discovered means Google knows the URL but hasn’t crawled it; Crawled means it was fetched but not indexed and usually points to a quality issue — and the fixes overlap. Make the content unique, valuable, and intent-matched; clear any stray noindex/robots/canonical blockers; keep the server fast and stable; build a logical hierarchy with strong internal linking; then use URL Inspection to request a re-crawl and monitor. For the full breakdown of crawl capacity vs. crawl demand and a step-by-step fix path, see the dedicated “Discovered – currently not indexed” articleA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal..

How to help discovery

In rough order of leverage:

  1. Strengthen internal links to the page — kill the orphan. This is the highest- leverage lever, because links are one of the two discovery routes Google names directly.
  2. Put it in a clean XML sitemap (canonical, indexable URLs; accurate lastmod). Submit the sitemap in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility..
  3. Raise overall site quality so crawl demand rises — this is what moves “Discovered – currently not indexed” pages.
  4. Use URL Inspection → Request Indexing for a single important URL (remember the quota; repeats don’t speed it up).
  5. For Bing and other participating engines, use IndexNow to push changed URLs instantly. For Google, there’s no equivalent push for general pages — rely on links + sitemaps + URL Inspection.

Where to go next

This hub is the map for how a URL becomes known. The deep dives nested under it cover the mechanics of the sitemap side:

Sitemaps

  • Sitemaps — the overview: what a sitemap is, the formats Google accepts, and how submitting one fits into discovery (aid, not guarantee).
    • XML sitemap — the standard format, its fields, size limits, and lastmod done right.
    • Sitemap index — how to split and reference multiple sitemaps when you outgrow one file.
    • Image sitemap — surfacing images for discovery in image search.
    • Video sitemap — surfacing video content and its metadata.

For what happens after a URL is discovered — fetching, the crawl scheduler, renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., and crawl budget — see the CrawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. hub. Related push topics like IndexNow and the Google Indexing API also live near the crawling side. For the broader picture, see How Search WorksSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank..

Every nested topic above is its own deep dive under this hub — they’re in the sidebar too.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.