Discovery
How search engines discover URLs — internal links, backlinks, sitemaps, RSS/Atom, and push protocols like IndexNow and the Indexing API. The hub for everything discovery-related.
1 evidence signal on this page
- Related live toolXML Sitemap Generator
Discovery is the find step before search ever fetches a page: the engine has to learn a URL exists before it can crawl, index, or rank it. It's this hub's teaching name for the opening move of what Google's official model calls the crawling stage. URLs are discovered by pull (internal links, backlinks, sitemaps, RSS/Atom) and by push (IndexNow, the Indexing API, lastmod, WebSub, URL Inspection). Google calls it 'URL discovery,' and it's distinct from crawling and indexing — which is exactly why 'Discovered – currently not indexed' means found-but-not-yet-crawled, a status with several possible causes rather than one deterministic trigger. Links and sitemaps are the two routes Google names directly, orphan pages are the classic failure, and submitting a URL never guarantees indexing. This hub maps it all and points you to the sitemap deep dives.
Evidence for this claim Google discovers URLs through links, sitemaps, and other previously known signals before crawling and possible indexing. Scope: Current Google crawling and indexing overview. Confidence: high · Verified: Google Search Central: How Search works Evidence for this claim Sitemaps and indexing requests can aid discovery but do not guarantee crawling, indexing, or serving in results. Scope: Current Google sitemap and indexing-request behavior. Confidence: high · Verified: Google Search Central: Learn about sitemapsTL;DR — Discovery is how a search engine finds out your page exists — the step before it ever downloads the page. Engines mostly find pages by following links and reading sitemaps. If nothing links to a page and it’s not in a sitemap, it may never get found. Submitting a page helps it get discovered, but it doesn’t guarantee it’ll be indexed.
What discovery is
Before a search engine can show your page, three things have to happen in order:
- Discover — the engine learns your URL exists.
- Crawl — the engine downloads the page.
- Index — the engine stores it so it can show up in results.
Discovery is step zero of that pipeline. Google even has a name for it: URL discovery. A page that’s never discovered can’t be crawled, indexed, or ranked — the engine simply doesn’t know it’s there.
A quick note on the model: discover → crawl → index is this hub’s teaching breakdown, because it’s the clearest way to reason about a stuck page. Google’s own official stage list is three stages — crawling, indexing, serving — and it defines “URL discovery” as the first part of crawling, not as a separate fourth stage. Same sequence of events either way; we’re just naming the find step explicitly because it’s useful to talk about on its own.
Evidence for this claim Google's official three-stage model is crawling, indexing, and serving; it places URL discovery at the beginning of the crawling stage rather than defining discovery as a separate official fourth stage. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search worksHow search engines find your pages
There are two big ways:
- Following links. When a bot reads a page it already knows, it grabs the links on it and adds those URLs to its list. A category page linking to a new blog post is the classic example — that’s how the new post gets found. Links from other sites (backlinks) work the same way.
- Sitemaps. An XML sitemap is a list of your URLs that you hand straight to search engines, so they don’t have to find everything through links. RSS and Atom feeds do a similar job for your most recent pages.
There are also “push” options where you actively ping an engine that a page changed (like IndexNow for Bing) — more on those in the Advanced version — but links and sitemaps do most of the work.
Orphan pages: the thing that breaks discovery
An orphan page is a page with no internal links pointing to it. Because links are the main way engines find pages, an orphan can be hard or impossible to discover. The fix isn’t a clever submission trick — it’s linking to the page from somewhere relevant on your own site.
The thing most people get wrong
Submitting a page doesn’t mean it gets indexed. Adding a URL to your sitemap or hitting “Request Indexing” in Search Console helps the engine discover it — but crawling and indexing still aren’t guaranteed. If you’ve ever seen “Discovered – currently not indexed” in Search Console, that’s exactly this: Google knows the URL exists, it just hasn’t gotten around to crawling it yet.
Want the deeper version — the full pull-vs-push channel list, why Google doesn’t use IndexNow, and how to fix “Discovered – currently not indexed”? Switch to the Advanced tab.
Evidence for this claim Google discovers URLs through links, sitemaps, and other previously known signals before crawling and possible indexing. Scope: Current Google crawling and indexing overview. Confidence: high · Verified: Google Search Central: How Search works Evidence for this claim Sitemaps and indexing requests can aid discovery but do not guarantee crawling, indexing, or serving in results. Scope: Current Google sitemap and indexing-request behavior. Confidence: high · Verified: Google Search Central: Learn about sitemapsTL;DR — Discovery is this hub’s name for the first move in search (discover → crawl → index → serve): the engine has to learn a URL exists before it fetches anything. Google’s own official model nests this inside its three-stage crawling stage rather than naming a separate fourth stage — same sequence, different vocabulary. URLs are discovered by pull (internal links, backlinks, sitemaps, RSS/Atom) and push (IndexNow, the Indexing API, sitemap
lastmod, WebSub, URL Inspection). Google calls the find step “URL discovery,” and it’s distinct from crawling and indexing — which is why “Discovered – currently not indexed” means found-but-not-yet-crawled, a state with several possible causes rather than one guaranteed diagnosis. Links and sitemaps are the two routes Google names directly, orphan pages are the canonical failure, IndexNow is Bing/others not Google, the Indexing API is scope-limited toJobPosting/BroadcastEvent, and submitting a URL never guarantees indexing.
What URL discovery is
Discovery is the find step. Before a page can be crawled, indexed, or served, the engine has to know it exists and add it to its list of known pages. Google names this explicitly: “This process is called ‘URL discovery’.” In Google’s own words, “Some pages are known because Google has already visited them. Other pages are discovered when Google extracts a link from a known page to a new page… Still other pages are discovered when you submit a list of pages (a sitemap) for Google to crawl.”
Evidence for this claim Google does not guarantee that a discovered or submitted URL will be crawled, indexed, or served. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search worksThat single passage is the whole hub in miniature: pages are found through links (internal and external) and through sitemaps. Everything else is a variation on those two ideas — or a way to push a notification instead of waiting to be pulled.
A note on the stage model. This hub treats discover → crawl → index → serve as four distinct, named steps, because separating “found” from “fetched” from “stored” is the clearest way to diagnose a stuck page. Google’s own official model names three stages — crawling, indexing, serving — and defines URL discovery as the opening move within crawling, not as its own fourth stage. The sequence of events is identical either way; this hub just gives the find step its own name and its own section because it behaves differently enough (and gets misdiagnosed often enough) to deserve one.
Evidence for this claim Google's official three-stage model is crawling, indexing, and serving; it places URL discovery at the beginning of the crawling stage rather than defining discovery as a separate official fourth stage. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search worksDiscovery vs crawling vs indexing
Keep these three ideas separate and most discovery confusion disappears — using our four-step teaching breakdown of Google’s three official stages:
- Discovery — the engine learns a URL exists. Nothing has been fetched yet.
- Crawling — the engine downloads the discovered URL (and renders it). That’s the Crawling hub’s job.
- Indexing — the engine analyzes and stores the crawled page so it can be served.
Google frames the back half of this as three stages: “Google Search works in three stages, and not all pages make it through each stage” — crawling, indexing, and serving. Discovery is the front door of stage one. And none of these steps is a ranking factor on its own; they’re prerequisites to being eligible to rank. A URL that’s never discovered is never crawled, indexed, or served — full stop.
How search engines discover URLs: pull vs push
The cleanest mental model is pull vs push. Pull is the engine finding URLs on its own; push is you notifying it.
Pull — the engine finds it on its own
Internal links and backlinks. Google: “Other pages are discovered when Google extracts a link from a known page to a new page: for example, a hub page, such as a category page, links to a new blog post.” Google’s own docs also say a properly linked site can usually get most of its pages discovered this way, with new sites that have few external links and large sites with unlinked pages at greater risk of gaps — links aren’t the only route in, but they’re the one Google names first and the one that scales automatically as you publish. That’s why internal linking is one of the highest-leverage discovery levers you have, and why orphan pages struggle (more below). Backlinks from other sites work identically — an external link from a page the engine already knows surfaces your new URL.
Evidence for this claim Google says it can usually discover most pages on a properly linked site, while new sites with few external links and large sites with unlinked pages have a greater risk of pages not being discovered. Scope: sitemaps and internal links Confidence: high · Verified: What is a sitemap?Sitemaps. Google: “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” That caveat is the whole point — a sitemap is a discovery aid, not an indexing guarantee. The full mechanics live in the nested sitemaps deep dives (XML sitemap, sitemap index, image sitemap, video sitemap).
RSS / Atom feeds. Google accepts feeds as a sitemap format — “Google accepts RSS 2.0 and Atom 1.0 feeds.” The catch is that a feed only surfaces your recently changed URLs, so it complements a full sitemap rather than replacing it.
Push — you notify the engine something changed
WebSub (PubSubHubbub). For RSS/Atom feeds, Google supports WebSub: “If you use Atom or RSS, you can use WebSub to broadcast your changes to search engines, including Google.” Instead of waiting to be re-pulled, your feed broadcasts the change.
Sitemap lastmod. A freshness signal that tells engines a URL changed and may
be worth re-crawling. Bing leans on it harder than Google does — Bing has said the
lastmod field remains a key signal that helps it prioritize URLs for recrawling
and reindexing. (Detail lives in the sitemaps cluster.)
Google Indexing API — scope-limited. This is the one that gets misused most.
The Indexing API can only be used to crawl pages with either JobPosting or
BroadcastEvent embedded in a VideoObject. It is not a general “submit any URL
to Google” endpoint, no matter how often it’s pitched that way. If your page isn’t a
job posting or a live-stream broadcast event, the Indexing API isn’t your tool.
IndexNow — Bing and others, not Google. IndexNow is a push protocol: it notifies enabled search engines the instant a URL is added, updated, or deleted, and per its own FAQ, “Search engines adopting the IndexNow protocol agree that submitted URLs will be automatically shared with all other participating search engines.” As of this writing (per indexnow.org’s current documentation, checked July 2026), the participating engines are Microsoft Bing, Yandex, Naver, Seznam.cz, and Yep — check the live IndexNow docs for the current list, since it can change. Google is not on that list. There’s no Google help page that says “we don’t support IndexNow” — Google’s non-participation is inferred from its absence from the current partner list plus Googlers describing it as something they only tested, not from a first-party Google confirmation, so treat it as the best available reading rather than an official statement. And regardless of engine: notifying IndexNow only tells it a URL changed — each participating engine still decides independently whether and when to crawl and index it, the same “notification isn’t a guarantee” rule that applies to sitemaps and every other discovery channel. For Google, you fall back to links, sitemaps, and URL Inspection.
Manual submission — URL Inspection. For one-off Google submission: “To request a crawl of individual URLs, use the URL Inspection tool.” But don’t expect repeats to force speed — Google is explicit that “there’s a quota for submitting individual URLs and requesting a recrawl multiple times for the same URL won’t get it crawled any faster.” Bing has its own URL submission channel (up to 10,000 URLs/day), separate from IndexNow.
Two discovery routes feed one crawl queue. Pull discovery includes following links and reading sitemaps. Push discovery includes IndexNow for participating engines, the scope-limited Google Indexing API, and freshness notifications such as sitemap lastmod, RSS, or WebSub. Google does not use IndexNow for general pages, and its Indexing API supports only JobPosting and qualifying livestream pages.
© Patrick Stox LLC · CC BY 4.0 ·
How Google and Bing each describe discovery
Both engines frame discovery as the front of the pipeline. Bing puts it directly in its definition of crawling — Fabrice Canel: “Crawling is the process by which bingbot discovers new and updated documents or content to be added to Bing’s searchable index.”
The scale is worth sitting with, even if the exact figure moves over time. Fabrice Canel has described Bing discovering on the order of tens of billions of normalized URLs it had never seen before, per day — he put it as ”12s of billions … never seen before” in an August 2022 post, a number I haven’t independently re-verified against a current primary source, so treat it as a directional sense of scale rather than a precise, current daily count. Whatever the exact figure is today, discovery operates as a firehose, and the engines deprioritize aggressively — which is the backdrop for “Discovered – currently not indexed” below.
Orphan pages: when discovery fails
An orphan page has no internal links pointing to it. Because links are one of the two discovery routes Google names directly (alongside sitemaps), an orphan can only be found via a sitemap, an external link, a redirect, or a canonical/hreflang reference — and if none of those exist, it may never be discovered at all.
The fix is internal linking, not push tricks. Link the page from a relevant hub/category page — exactly Google’s “category page links to a new blog post” example. Putting an orphan in your XML sitemap helps Google discover it, but it doesn’t replace the internal-link signal that tells the engine the page matters. Orphans frequently show up as “Discovered – currently not indexed” for precisely this reason: Google found the sitemap entry but deprioritized crawling a page nothing links to.
”Discovered – currently not indexed” in Search Console
What’s certain about this status is the stage boundary — this hub’s job is to fix that boundary in your head, not to fully diagnose it (the dedicated deep dive does that). Google’s description: the page was found by Google but not crawled yet — Google wanted to crawl the URL but expected it would overload the site, so it rescheduled the crawl, which is why the last-crawl date is empty.
Contrast it with “Crawled – currently not indexed,” which is a different, downstream problem: the page was crawled but not indexed, and it may or may not be indexed in the future. The stage boundary is the lesson most third-party write-ups miss:
- Discovered – currently not indexed = Google knows the URL but hasn’t fetched it. A discovery/crawl-scheduling state (empty last-crawl date).
- Crawled – currently not indexed = Google fetched it but chose not to index it. An indexing/quality state.
What causes “Discovered – currently not indexed” — work through it in this order, not as a single confirmed cause:
- Crawl scheduling / perceived server-load deprioritization — Google’s own stated reason for the status itself.
- The page is orphaned or only weakly internally linked, lowering how much Google wants to spend crawl effort on it.
- Site-wide quality or authority signals dragging on overall crawl demand — a site-level pattern, not proof of a defect on this one page.
- A new or low-authority site with little crawl demand yet.
None of these is the universal cause — Google documents the crawl-scheduling
explanation directly but doesn’t publish a single deterministic reason a given URL
sits in this state, so treat the list above as a hypothesis ladder to work through,
not a diagnosis to assume. It also doesn’t mean something is broken: John Mueller has
said “it’s completely normal that we don’t index everything off of the website,”
with the advice being to reconsider overall site quality rather than hunt for a
per-page technical issue, and Gary Illyes has pointed out that the vast majority of
websites don’t need to think about crawl budget at all. My own practitioner shorthand,
from my indexing guide at Ahrefs (worth a
re-read before treating it as gospel, since I haven’t re-verified the live page in
this pass): Discovered means Google knows the URL but hasn’t crawled it; Crawled
means it was fetched but not indexed and usually points to a quality issue — and the
fixes overlap. Make the content unique, valuable, and intent-matched; clear any stray
noindex/robots/canonical blockers; keep the server fast and stable; build a logical
hierarchy with strong internal linking; then use URL Inspection to request a re-crawl
and monitor. For the full breakdown of crawl capacity vs. crawl demand and a
step-by-step fix path, see the
dedicated “Discovered – currently not indexed” article.
How to help discovery
In rough order of leverage:
- Strengthen internal links to the page — kill the orphan. This is the highest- leverage lever, because links are one of the two discovery routes Google names directly.
- Put it in a clean XML sitemap (canonical, indexable URLs; accurate
lastmod). Submit the sitemap in Google Search Console and Bing Webmaster Tools. - Raise overall site quality so crawl demand rises — this is what moves “Discovered – currently not indexed” pages.
- Use URL Inspection → Request Indexing for a single important URL (remember the quota; repeats don’t speed it up).
- For Bing and other participating engines, use IndexNow to push changed URLs instantly. For Google, there’s no equivalent push for general pages — rely on links + sitemaps + URL Inspection.
Where to go next
This hub is the map for how a URL becomes known. The deep dives nested under it cover the mechanics of the sitemap side:
Sitemaps
- Sitemaps — the overview: what a sitemap is, the formats Google accepts, and
how submitting one fits into discovery (aid, not guarantee).
- XML sitemap — the standard format, its fields, size limits, and
lastmoddone right. - Sitemap index — how to split and reference multiple sitemaps when you outgrow one file.
- Image sitemap — surfacing images for discovery in image search.
- Video sitemap — surfacing video content and its metadata.
- XML sitemap — the standard format, its fields, size limits, and
For what happens after a URL is discovered — fetching, the crawl scheduler, rendering, and crawl budget — see the Crawling hub. Related push topics like IndexNow and the Google Indexing API also live near the crawling side. For the broader picture, see How Search Works.
Every nested topic above is its own deep dive under this hub — they’re in the sidebar too.
AI summary
A condensed take on the Advanced version:
- Discovery = the find step in search (discover → crawl → index → serve). The engine has to learn a URL exists before anything is fetched. Google calls it “URL discovery” and nests it inside its official three-stage model (crawling, indexing, serving) rather than naming it a separate fourth stage.
- It’s distinct from crawling and indexing. Discovered = known. Crawled = fetched. Indexed = stored. None of the three is a ranking factor on its own.
- Pull channels (engine finds it): internal links and backlinks (the link-based route Google names directly), XML sitemaps (the other route Google names directly), RSS/Atom feeds (recent URLs only).
- Push channels (you notify): WebSub, sitemap
lastmod, the Google Indexing API (JobPosting/BroadcastEventonly — not general submission), IndexNow (Bing/Yandex/Naver/Seznam/Yep — not Google), and URL Inspection “Request Indexing” (quota-limited). - Sitemaps aid discovery, not indexing — submitting a URL never guarantees a crawl or an index.
- Orphan pages (no internal links) are the canonical discovery failure; fix with internal links, not push tricks.
- “Discovered – currently not indexed” = Google knows the URL but hasn’t crawled it (a scheduling/priority state with several possible contributing causes, not one guaranteed diagnosis), distinct from the quality-driven “Crawled – currently not indexed.”
- Help discovery with internal links, a clean sitemap, higher site quality, URL Inspection, and IndexNow for Bing.
Official documentation
Primary-source documentation from the search engines.
- In-Depth Guide to How Google Search Works — defines “URL discovery” and the discover → crawl → index → serve pipeline.
- Sitemaps overview — how sitemaps help discovery (and why they don’t guarantee indexing).
- Build and submit a sitemap — RSS/Atom feed support and WebSub for pushing feed changes.
- Indexing API quickstart — the scope limit:
JobPostingandBroadcastEventpages only. - Ask Google to recrawl your URLs — URL Inspection submission and the quota note.
- Page Indexing report — the “Discovered – currently not indexed” and “Crawled – currently not indexed” definitions.
Bing / Microsoft
- bingbot Series: Maximizing Crawl Efficiency — Bing’s definition of crawling as a discovery process.
- Submit up to 10,000 URLs/day to Bing — Bing’s own URL submission channel.
- Keeping Content Discoverable with Sitemaps in AI-Powered Search — sitemaps + IndexNow as the discovery foundation, and
lastmodas a recrawl signal. - IndexNow / indexnow.org — the push protocol for instantly signaling changed URLs.
Quotes from the source
On-the-record statements from Google and Bing. Each link is a deep link that jumps to the quoted passage on the source page.
Google — what discovery is
- “Google must constantly look for new and updated pages and add them to its list of known pages. This process is called ‘URL discovery’.” — Google Search Central docs. Jump to quote
- “Other pages are discovered when Google extracts a link from a known page to a new page: for example, a hub page, such as a category page, links to a new blog post.” Jump to quote
- “Google Search works in three stages, and not all pages make it through each stage.” Jump to quote
Google — sitemaps, feeds, and push
- “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” — Google sitemaps overview. Jump to quote
- “Google accepts RSS 2.0 and Atom 1.0 feeds.” Jump to quote
- “If you use Atom or RSS, you can use WebSub to broadcast your changes to search engines, including Google.” Jump to quote
- “To request a crawl of individual URLs, use the URL Inspection tool.” Jump to quote
- “There’s a quota for submitting individual URLs and requesting a recrawl multiple times for the same URL won’t get it crawled any faster.” Jump to quote
Google — not indexing everything is normal
- “It’s completely normal that we don’t index everything off of the website.” — John Mueller, Google. Read the coverage
Bing / Microsoft
- “Crawling is the process by which bingbot discovers new and updated documents or content to be added to Bing’s searchable index.” — Fabrice Canel, Microsoft Bing. Jump to quote
- “Instead of Bing continually monitoring RSS and similar feeds or frequently crawling websites to check for new pages, discover content changes and/or new outbound links, websites will notify Bing directly about relevant URLs changing on their website.” — Microsoft Bing, on Bing’s URL submission. Jump to quote
IndexNow
- “Search engines adopting the IndexNow protocol agree that submitted URLs will be automatically shared with all other participating search engines.” — indexnow.org documentation. Jump to quote
Discovery checklist
A quick pass to confirm search engines can actually find what matters:
- Every important page is linked from somewhere crawlable — no orphans.
- Important pages sit close to the homepage / main nav (the more linked, the easier to discover).
- An XML sitemap lists only canonical, indexable URLs, with accurate
lastmod, and is submitted in Google Search Console and Bing Webmaster Tools. - If you publish frequently, an RSS/Atom feed surfaces recent URLs (and WebSub broadcasts changes).
- You’re not relying on the Google Indexing API for general pages — it’s only
for
JobPosting/BroadcastEvent. - IndexNow is wired up for Bing and other participating engines (it does nothing for Google).
- For a single urgent URL, you use URL Inspection → Request Indexing (and don’t spam it — quota applies).
- Pages you want found aren’t accidentally blocked by
robots.txtor buried behind click-only navigation instead of real<a href>links. - You’ve checked the Page Indexing report for “Discovered – currently not indexed” and treated it as a linking/quality signal, not a per-page bug.
The mental models
1. Discovery = pull + push.
Pull is the engine finding URLs on its own (internal links, backlinks, sitemaps,
RSS/Atom). Push is you notifying it (IndexNow, the Indexing API, lastmod, WebSub,
URL Inspection). When a page isn’t being found, ask: which channel should be
carrying it — and is anything actually linking to it?
2. Orphan pages can’t be discovered. Links are one of the two routes Google names directly, so a page nothing links to has almost nothing carrying it into the engine’s known-pages list. A sitemap entry helps it get discovered but doesn’t replace the internal link that says the page matters. Fix discovery failures with links first.
3. Discovery ≠ crawling ≠ indexing. Three separate stages: Discovered (the engine knows the URL) → Crawled (it fetched the URL) → Indexed (it stored the URL). None is a ranking factor on its own — they’re prerequisites. Locate which stage a page is failing at before you change anything.
4. “Discovered – currently not indexed” = found, not yet (or never) crawled. It’s a discovery/scheduling state: Google knows the URL but hasn’t fetched it (empty last-crawl date). That’s different from “Crawled – currently not indexed,” which is a downstream quality/indexing state. The fix for Discovered is usually internal links + site quality + a clean sitemap, not repeated re-submission.
5. Submission ≠ indexing. Adding a URL to a sitemap, pinging IndexNow, or hitting Request Indexing only helps discovery. The engine still decides whether to crawl and index. Never promise a client that “submitting it” gets it indexed.
Discovery channels — cheat sheet
PULL — the engine finds it on its own
| Channel | Who uses it | What it’s for | Gotcha |
|---|---|---|---|
| Internal links | Google, Bing | One of the two discovery routes Google names directly; carries new pages into the known-pages list | A page nothing links to (orphan) may never be found |
| External links / backlinks | Google, Bing | Discovery via a third-party page the engine already knows | You don’t control when/whether others link |
| XML sitemap | Google, Bing | A direct list of your URLs as a discovery aid | Aids discovery only — never guarantees crawl or index |
| RSS / Atom feed | Google, Bing | Surfaces your recently changed URLs | Recent URLs only; not a substitute for a full sitemap |
PUSH — you notify the engine something changed
| Channel | Who uses it | What it’s for | Gotcha |
|---|---|---|---|
| IndexNow | Bing, Yandex, Naver, Seznam, Yep | Instantly notify participating engines a URL changed (submit once, shared with all) | Not Google — Google does not participate |
| Google Indexing API | Push for specific page types | Only JobPosting / BroadcastEvent (in VideoObject) — not general submission | |
Sitemap lastmod | Google, Bing | Freshness signal to prioritize recrawl | Bing weights it more than Google; only trusted if accurate |
| WebSub (PubSubHubbub) | Google (feeds) | Broadcast RSS/Atom feed changes | Only useful if you publish a feed |
| URL Inspection (manual) | Request a crawl of one URL | Quota-limited; repeats don’t speed it up |
Two myths to keep straight
- IndexNow = Bing/Yandex/others, NOT Google. There’s no Google IndexNow endpoint; for Google use links + sitemaps + URL Inspection.
- Indexing API = Google,
JobPosting/BroadcastEventonly. It is not a general “submit any URL to Google” tool.
Common discovery issues
Symptom-first fixes for the discovery problems that actually show up in Search Console and site audits.
”Discovered – currently not indexed” in the Page Indexing report
Symptom: A page sits in the “Discovered – currently not indexed” bucket, with an empty last-crawl date, sometimes for weeks.
Likely cause(s): Google knows the URL exists (from a link or a sitemap) but hasn’t scheduled a crawl for it yet — usually because of low crawl demand: the page is orphaned or weakly linked, or site-wide quality/authority signals are keeping overall crawl demand low.
Fix + check: Add internal links to the page from a relevant hub/category page,
confirm it’s in a clean XML sitemap with an accurate lastmod, and raise
site-wide quality if the site as a whole has low crawl demand. Confirm the fix by
re-checking the same URL in URL Inspection after a few days — a populated
last-crawl date means the state has moved (to indexed, or to “Crawled – currently
not indexed,” which is a different problem).
A new page never shows up anywhere in Search Console
Symptom: A URL you know exists doesn’t appear in the Page Indexing report at all — not discovered, not crawled, nothing.
Likely cause(s): It’s an orphan page with no internal links pointing to it, it was never added to the sitemap, and no external site links to it either.
Fix + check: Link to it from a relevant page on your own site (the single
biggest lever), add it to your XML sitemap, and submit the URL through URL
Inspection → Request Indexing for a one-off nudge. Check by searching
site:yourdomain.com/the-url — no results confirms it’s still undiscovered.
IndexNow pings aren’t doing anything for Google traffic
Symptom: IndexNow is wired up, changed URLs are being pinged, but Google rankings/traffic for those URLs don’t move any faster.
Likely cause(s): Google does not participate in IndexNow — only Bing, Yandex, Naver, Seznam.cz, and Yep do. Pinging IndexNow has no effect on Google discovery.
Fix + check: Keep IndexNow for the engines that support it, and rely on links, sitemaps, and URL Inspection for Google. Confirm by checking whether the affected URLs show movement in Bing Webmaster Tools (where IndexNow should help) versus Google Search Console (where it won’t).
The Indexing API call gets rejected or does nothing
Symptom: A call to Google’s Indexing API returns an error, or the URL never gets crawled faster despite a successful-looking call.
Likely cause(s): The Indexing API only accepts pages with JobPosting or
BroadcastEvent (in a VideoObject) structured data — it’s not a general
“submit any URL” endpoint, no matter how it gets pitched.
Fix + check: Confirm the page actually carries JobPosting or BroadcastEvent
markup before troubleshooting the API call itself. If it’s any other page type,
switch to internal links, a sitemap, and URL Inspection instead.
Playbook: a new page isn’t showing up in search
A linear runbook for the most common discovery complaint — “I published it, why can’t I find it in Google?”
1. Confirm it’s actually a discovery problem, not an indexing problem. Run the URL through URL Inspection in Search Console. If the report shows “Discovered – currently not indexed,” you’re in the right place — keep going. If it shows “Crawled – currently not indexed,” this is a downstream quality issue, not a discovery issue — stop here and work the content/quality angle instead.
2. Check whether the page is an orphan.
Crawl the site (or check your internal-link report) for any live <a href> links
pointing to the URL. If nothing links to it, that’s a strong candidate for the root
cause — links are one of the two discovery routes Google names directly.
3. If it’s an orphan, link to it from a relevant page. Add a real, crawlable link from a hub or category page that’s already indexed — exactly Google’s own example of a category page linking to a new blog post. Don’t substitute a sitemap entry for this step; a sitemap helps discovery but doesn’t carry the same weight as an internal link.
4. Check the sitemap regardless.
Confirm the URL is listed in your XML sitemap with the canonical form and an
accurate lastmod, and that the sitemap itself is submitted in Google Search
Console and Bing Webmaster Tools.
5. If you see “Discovered – currently not indexed” and the page is already well linked, look at site-wide quality. Google and Bing both deprioritize crawl demand for sites with weaker overall quality signals — this is a site-level issue, not a bug on this one page. Improving overall site quality is what actually moves pages out of this state at scale.
6. For a single urgent URL, request indexing manually. Use URL Inspection → Request Indexing once. Remember there’s a quota, and repeating the request doesn’t get it crawled any faster.
7. If the target engine is Bing (or Yandex, Naver, Seznam, Yep), push with IndexNow instead of waiting. IndexNow gives those engines an instant notification. It does nothing for Google — don’t expect it to move the Google needle.
8. Re-check after a few days, not a few hours. Discovery and crawl scheduling operate on a queue. Come back to URL Inspection later rather than repeatedly resubmitting.
Patrick's relevant free tools
- XML Sitemap Validator — Paste, upload, or fetch a sitemap by URL — errors, warnings, and a health score with line numbers. Pasted and uploaded sitemaps are validated entirely in your browser.
- IndexNow Submitter — Validate and explicitly submit a same-host URL list to IndexNow; Google does not use IndexNow.
- XML Sitemap Generator — Generate an XML sitemap from a capped, robots-respecting same-site crawl. Noindex, off-canonical, failed, and uncertain URLs remain visibly separate; lastmod dates are emitted only when the page provides evidence.
Tools for checking and fixing discovery
- Google Search Console — URL Inspection — check whether a specific URL is discovered, crawled, and indexed, and request a one-off crawl.
- Google Search Console — Page Indexing report — see how many URLs are stuck in “Discovered – currently not indexed” versus “Crawled – currently not indexed.”
- XML Sitemap Generator — build a clean sitemap of canonical, indexable URLs to hand straight to search engines.
- Sitemap Validator — check an existing sitemap for malformed entries, non-canonical URLs, or other issues that weaken it as a discovery aid.
- Robots.txt Tester — confirm
robots.txtisn’t accidentally blocking crawler access to pages you want discovered. - Canonical Checker — verify canonical tags aren’t quietly pointing discovery and indexing signals away from the page you want found.
- Scout Site Audit Free — crawl your own site to surface orphan pages and weak internal-linking patterns before they turn into “Discovered – currently not indexed.”
- Log File Analyzer — check server logs to see whether Googlebot and Bingbot are actually requesting the URLs you expect them to discover and crawl.
- Google Search Console CSV Analyzer — work with Search Console indexing and coverage data at scale instead of clicking through individual URLs.
- Bing Webmaster Tools — Bing’s equivalent of URL Inspection and the Page Indexing report, plus its own URL submission channel and IndexNow status.
- IndexNow (indexnow.org) — ping Bing, Yandex, Naver, Seznam.cz, and Yep the moment a URL changes.
Resources worth your time
My related writing
- Crawled – currently not indexed: 7 ways to fix it — the GSC status that sits right after discovery, and how the fixes overlap with “Discovered – currently not indexed.”
- Googlebot: What Is It & How Does It Work? — where Googlebot collects URLs from (pages, sitemaps, RSS feeds, Search Console, the Indexing API) and how it reprocesses pages to find more links.
- IndexNow: now in Ahrefs and powering Yep — what IndexNow does, which engines participate, and why Google sits it out.
- The Beginner’s Guide to Technical SEO — where discovery fits in the bigger picture.
My speaking
- How Search Works (SlideShare) — my walkthrough of discovery, crawling, rendering, indexing, and ranking, including the discovery sources (links, XML sitemaps, GSC requests, the Indexing API, RSS feeds, WebSub) and the crawl-demand vs crawl-rate model that explains why a page sits “Discovered – currently not indexed.” (My standing disclaimer applies: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From others
- r/TechSEO — the community for discovery, crawl, and index debugging.
- Google’s Crawling December series — official, and the best concentrated set of crawl/discovery explainers.
- Understanding and resolving “Discovered – currently not indexed” (Search Engine Land) — causes and remediation, including the Gary Illyes crawl-budget framing.
- It’s Normal for Pages of a Site to Not Be Indexed (Search Engine Journal) — covers John Mueller’s “completely normal” statement on Google not indexing everything, and the quality-over-submission advice.
- Search Engine Roundtable — Crawled, Currently Not Indexed Is A Quality Issue — Mueller on the downstream “Crawled – currently not indexed” status being site-wide quality, not a per-page bug.
- Bing Lets Webmasters Submit 10,000 URLs Per Day (Search Engine Land) — background on Bing’s own URL submission channel, distinct from IndexNow.
- Search Engine Crawling — How It Works (Lumar) — crawler mechanics, link-following, and sitemaps as a discovery foundation.
Podcasts
- Search Off the Record (Google Search Relations) — How Googlebot crawls the web. Gary Illyes and Martin Splitt on how Googlebot discovers and fetches the web: unified crawling infrastructure, conditional requests, and how discovery feeds the scheduler. Listen
Videos
- Google Search Central (YouTube) — the How Google Search Works series and Martin Splitt’s crawling/rendering explainers, which cover where URL discovery sits in the pipeline. Channel
Test yourself: Discovery
Five questions on how search engines find URLs, based on the hub above.
URL discovery
URL discovery is how search engines find URLs to crawl — by pull (following links and reading sitemaps) and by push (you notify them via IndexNow, the Indexing API, or WebSub). It's the find step that comes before a page is ever fetched.
URL discovery
URL discovery is the very first thing that has to happen for a page to appear in search: the engine has to learn the URL exists. It’s Google’s own name for this step — before anything is crawled, indexed, or ranked, a URL has to be discovered and added to the engine’s list of known pages.
Discovery happens two ways. Pull is the engine finding URLs on its own: it extracts links from pages it already knows (internal links and backlinks), and it reads the sitemaps and RSS/Atom feeds you publish — links and sitemaps are the two routes Google names directly, and links do especially heavy lifting on a well-linked site, which is why a page that nothing links to — an orphan page — struggles to get discovered at all. Push is you telling the engine directly that something changed: IndexNow (used by Bing, Yandex, and others — not Google, per IndexNow’s current documentation), Google’s Indexing API (officially only for JobPosting and BroadcastEvent pages), sitemap lastmod, WebSub for feeds, and Google’s URL Inspection “Request Indexing.”
Discovery is distinct from crawling and indexing. Discovery = the engine knows the URL. Crawling = the engine fetches it. Indexing = the engine analyzes and stores it. That’s why Search Console’s “Discovered – currently not indexed” status means Google knows about the URL but hasn’t crawled it yet — a discovery/scheduling state with several possible contributing causes (crawl scheduling, weak internal links, site-wide crawl demand), not a single deterministic content problem. Submitting a URL or a sitemap aids discovery; it never guarantees a crawl or an index.
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 25, 2026.
Editorial summary and recorded change details.Summary
Updated the public name of the bounded site-audit tool referenced in this hub.
Change details
-
Renamed the linked SEO Site Audit Crawler to Scout Site Audit Free.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 20, 2026.
Editorial summary and recorded change details.Summary
Corrected this hub's own sourceUrl after the how-search-works taxonomy reorg dropped the intermediate crawling segment from the discovery article's canonical path.
Change details
-
Updated sourceUrl from /technical-seo/how-search-works/crawling/discovery/ to /technical-seo/how-search-works/discovery/ to match the current taxonomy.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Fixed a dead Indexing API quickstart link in the Official Docs lens (old developers.google.com/search/docs/crawling-indexing/... path 404s); replaced with the current developers.google.com/search/apis/indexing-api/v3/quickstart URL.
Change details
-
Official Docs: replaced dead developers.google.com/search/docs/crawling-indexing/indexing-api/quickstart with the live developers.google.com/search/apis/indexing-api/v3/quickstart.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 17, 2026.
Editorial summary and recorded change details.Summary
Qualified overclaimed 'dominant channel' language, added Google's official stage-model framing, replaced a single-cause 'Discovered – currently not indexed' explanation with a hypothesis ladder linking to the dedicated deep dive, and dated/hedged the IndexNow participant list and Bing daily-discovery statistic.
Change details
- Advanced
Replaced categorical claims that internal links are the 'dominant' discovery channel and that 'Discovered – currently not indexed' is 'not a quality problem' with evidence-bounded phrasing (links/sitemaps are the two routes Google names directly; the status has several possible contributing causes), and linked to the dedicated discovered-currently-not-indexed article for full diagnosis.
- Beginner
Added a short note distinguishing this hub's four-step discover/crawl/index/serve teaching model from Google's official three-stage model, which nests URL discovery inside crawling.
Full comparison unavailable — no prior snapshot was archived for this revision.