How to Get Indexed by Google
The right way to get pages indexed by Google — sitemaps, internal links, Request Indexing, the Indexing API and IndexNow, and why quality is the real gate.
1 evidence signal on this page
- Related live toolrobots.txt Tester
Getting indexed means Google crawled your URL, judged it worthwhile, and stored it so it can appear in results — a prerequisite for ranking, never a ranking factor. The single most useful reframe: crawling and indexing are separate steps, and the usual blocker isn't a missing button-click, it's quality. What you actually control is discovery and priority: a clean XML sitemap, real internal links so nothing is orphaned, and — for a handful of individual high-priority URLs — Request Indexing in URL Inspection (which does not guarantee or speed a quality decision, and 'no need to resubmit'). The two 'fast-index' shortcuts people reach for are narrower than advertised: Google's Indexing API is officially only for JobPosting and BroadcastEvent-in-VideoObject pages (default 200 publish requests/day), and IndexNow is a Bing/Yandex protocol that Google does not use at all. Nothing guarantees indexing. This page is the actions companion to the Page Indexing report.
TL;DR — Getting indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. means Google has crawled your page, decided it’s worth keeping, and stored it so it can show up in search. You can’t rank without it. The things you actually control are helping Google find the page (internal links + a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.) and, for a few important URLs, clicking Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.. None of it forces indexing — if a page isn’t getting indexed, the reason is usually that Google doesn’t think it’s good enough yet.
What “getting indexed” means
When you publish a page, three things have to happen before it can appear in Google:
- Crawl — GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. finds the URL and downloads it.
- Index — Google processes the page and decides whether to store it in its index (the giant database it pulls search results from).
- Serve (rank) — when someone searches, Google picks the best matches from the index and orders them.
“Getting indexed” is step two. And here’s the part that trips almost everyone up: a page can be crawled and still not indexed. CrawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexing are separate steps. Google will happily fetch your page and then choose not to keep it — you’ll see that in Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. as “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error..” Evidence for this claim Google separates crawling, indexing, and serving and states that not every page makes it through each stage. Scope: Google Search; crawling does not guarantee indexing. Confidence: high · Verified: Google: How Search works
What actually helps
There’s no button that forces Google to index a page. What you can do is make the page easy to find and worth keeping:
- Link to it. Every page you care about should be linked from at least one other page on your site. Pages that nothing links to (orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site.) are a top reason things never get indexed.
- Put it in your XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. and submit that in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.. A sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. helps Google discover your URLs — it doesn’t guarantee they’ll be indexed, but it’s the right tool when you have more than a handful of pages.
- Request indexing for individual important URLs. In Search Console, use the URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. tool and click “Request Indexing.” Use this for a few high-priority pages — not in bulk, and not over and over on the same URL. Evidence for this claim Google's URL Inspection request is for individual URLs, has a quota, does not crawl faster when repeated, and does not guarantee inclusion. Scope: Google Search Console request indexing workflow. Confidence: high · Verified: Google: Ask Google to recrawl URLs
- Make the page genuinely good. This is the real gate. Thin, duplicate, or low-effort pages are exactly the ones Google crawls and then declines to index.
The things people get wrong
- Clicking “Request Indexing” ten times doesn’t help. Google says re-requesting the same URL won’t get it crawled any faster, and its own docs literally say “no need to resubmit.”
- The “Google Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content.” is not a universal fast-track. It officially works only for job-posting pages and certain live-video pages. For a normal blog post or product page, it’s not the tool.
- IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. does not tell Google anything. IndexNow is a Bing and Yandex protocol. Google doesn’t use it. Submitting via IndexNow will speed up Bing, not Google.
- There’s no secret paid service that forces Google to index a page. Anything legitimate is just doing what you can do yourself in Search Console — and none of it overrides Google’s quality bar.
Want the full version — exact quotas, when you actually need a sitemap, what the Indexing API is really for, and how to read “Crawled – currently not indexed”? Switch to the Advanced tab.
TL;DR — IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. is crawl → judge → store, and the judgment is the hard part: Google indexes a URL only if it clears a quality/usefulness bar, so “why isn’t my page indexed” is usually a content problem, not a mechanics problem. The levers you control affect discovery and crawl priority, not the indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. decision: internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. (no orphans), a clean XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. (discovery aid, not a guarantee), and Request Indexing in URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. for a few individual high-priority URLs — which explicitly does not guarantee or speed a quality decision, and “no need to resubmit.” The two “fast-index” APIs are narrower than their reputation: Google’s Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. is officially only for
JobPostingandBroadcastEvent-in-VideoObjectpages (default 200 publish requests/day), and IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. is a Bing/Yandex protocol that Google does not use. Nothing guarantees indexing.
Crawling ≠ indexing — the reframe the whole page hangs on
The pipeline moves from crawl, where a bot discovers and downloads a URL, to index, where the engine processes the page and decides whether to store it, to serve or rank, where indexed pages may be selected for a query. Indexing is highlighted as a separate decision between crawling and ranking.
© Patrick Stox LLC · CC BY 4.0 ·
Most “why isn’t my page indexed” frustration comes from collapsing two separate steps into one. Google’s pipeline is three stages — crawl, index, serve — and, in Google’s own words, “not all pages make it through each stage.” A URL can be crawled and then not indexed. Evidence for this claim Google separates crawling, indexing, and serving and states that not every page makes it through each stage. Scope: Google Search; crawling does not guarantee indexing. Confidence: high · Verified: Google: How Search works That’s not a bug; it’s Google exercising editorial judgment.
So indexing is not a mechanical queue you can force your way into. It’s a gate with a quality bar. Getting indexed is a prerequisite for ranking, not a ranking factor: an indexed page can still rank poorly or for nothing at all, but a non-indexed page can never appear. Everything below is organized around that: the parts you control (discovery, crawl priority) versus the part Google controls (the indexing decision).
Why a page isn’t indexed — a quick diagnostic
Before reaching for tools, rule out the mechanical blockers. A page won’t be indexed if:
- It carries a
noindexmeta tag orX-Robots-Tagheader. - It’s blocked in
robots.txt(Google can’t read the content — and a robots-blocked URL can still show up indexed without content if links point to it; see canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. for that failure mode). - Its
rel="canonical"points at a different URL, so Google consolidates onto the canonical instead (GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.: “Alternate page with proper canonical tagA Google Search Console Page Indexing status meaning a page is a duplicate or alternate version that correctly points its canonical at another, indexed page. It's normal, healthy behavior — Google says there is nothing you need to do.” or “Duplicate, Google chose different canonical than user”). - It’s an orphan — nothing internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to it.
- It returns a 4xx/5xx or soft-404s.
- It depends on JavaScript for its main content or links, and Google’s renderer never successfully executes it — a blocked script/CSS file, a render error, or an app-shell page with no content in the initial HTML. GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. queues renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. as a separate step after the initial crawl, so a page can be fetched but still leave Google unable to see the real content or links inside it.
- The content is thin, duplicate, or low-value — the most common cause of “Crawled – currently not indexed.”
If none of the first six apply, you’re almost certainly in quality territory, and no amount of resubmitting changes that.
Request Indexing (URL Inspection) — for individual, high-priority URLs
The URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. in Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. lets you ask Google to (re)crawl one URL. Google: “To request a crawl of individual URLs, use the URL Inspection tool. You must be an owner or full user of the Search Console property to be able to request indexing in the URL Inspection tool.”
Three things to internalize:
- It’s rate-limited, and spamming it does nothing. Google: “There’s a quota for submitting individual URLs and requesting a recrawl multiple times for the same URL won’t get it crawled any faster.” Google does not publish the exact daily number, and has historically kept it opaque on purpose — it removed the on-screen “submissions remaining” counter from the old Fetch-as-Google tool back in 2018 while keeping the underlying limit. Practitioners informally cite a range around 10–15/day for lower-trust properties, but that’s observed, not official — treat any specific number as anecdotal.
- It doesn’t guarantee anything. Google: “Requesting a crawl does not guarantee that inclusion in search results will happen instantly or even at all. Our systems prioritize the fast inclusion of high quality, useful content.” And the timeline is vague by design — crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. “can take anywhere from a few days to a few weeks.” Evidence for this claim Google's URL Inspection request is for individual URLs, has a quota, does not crawl faster when repeated, and does not guarantee inclusion. Scope: Google Search Console request indexing workflow. Confidence: high · Verified: Google: Ask Google to recrawl URLs
- It’s for a handful of URLs, not bulk. Google’s own guidance pushes you to sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. for anything beyond a few: “If you want many pages indexed, try submitting a sitemap to Google.” And when a page is already stuck in “Crawled – currently not indexed,” Google’s status glossary is explicit: “no need to resubmit this URL for crawling.”
The through-line, which the reps have made repeatedly over the years, is to favor the non-manual channels — sitemaps and internal links — over compulsively clicking a button.
Submit and maintain a clean XML sitemap
A sitemap is a discovery aid, full stop. Google is blunt about the limit of what it buys you: “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.”
When you actually need one. Google’s rough guidance: you likely need a sitemap if your site is large, is new with few external links (Googlebot leans on already-known pages and backlinks to discover new URLs, so new sites are structurally slower to get found), or is rich in media/video/news. You likely don’t need one for a site of about 500 pages or fewer where every important page is reachable by following links from the homepage and there’s no significant media/news content.
Size limits. A single sitemap file is capped at 50MB uncompressed / 50,000 URLs; past that, split into multiple files and reference them from a sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file. file.
Audit indexing rates from it. The most useful workflow: in the Page Indexing report you can filter by sitemap to see how many of a given sitemap’s URLs are actually indexed. A big gap between submitted and indexed is a quality/architecture signal, not a reason to resubmit. And keep sitemaps lean — stuffing them with non-canonical or low-value URLs wastes crawl attention and muddies your own reporting; list only canonical, indexable URLs.
Fix your internal linking
This is the most foundational and most overlooked lever. Google’s link best-practices doc states it directly: “Every page you care about should have a link from at least one other page on your site.” Orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. — those with no internal links pointing at them — are a top real-world cause of non-indexing on both brand-new and very large sites.
A few specifics that matter for indexing:
- Links must be crawlable. That means real
<a href>elements, notonClickhandlers or JS-only navigation. If a bot can’t extract the link, the page it points to may never be discovered. - Keep important pages shallow. CrawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. don’t use search boxes or interpret nav menus the way people do — important pages should be reachable within a few clicks of the homepage.
- Link to the canonical URL, not a duplicate. Consistent internal linkingLinks between pages on the same site. to the preferred version reinforces Google’s understanding of which URL to index, and it’s one of the canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. signals.
- Anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. helps. Google: “Paying more attention to the anchor text used for internal links can help both people and Google make sense of your site more easily and find other pages on your site.”
The Google Indexing API — what it’s actually for
This is the single most misunderstood tool in the indexing conversation, so be
precise: the Indexing API is not a general fast-index endpoint. Google’s docs
name the only two supported content types outright: “The Indexing API can only be
used to crawl pages with either JobPosting or BroadcastEvent embedded in a
VideoObject.”
- Scope: job-posting pages and pages with a livestream event (BroadcastEvent in VideoObject) only. A normal article, product, or category page is out of scope.
- Quota: the default is 200 publish requests/day per project (covering both
URL_UPDATEDandURL_DELETED), 180 read-only requests/minute, and 380 across all endpoints/minute. Google notes “the quota may increase or decrease based on the document quality,” and quota-increase requests are gated to the two supported use cases. - Why using it for other content is risky. People do point it at unsupported pages and sometimes see it “work” temporarily — but Google’s Search Relations team has been clear this is unsupported and can be revoked without notice: the API may stop supporting unsupported content formats at any time, and access for non-supported verticals could be cut off overnight. Building your indexing strategy on an unsupported hack is building on sand.
If your site genuinely publishes job listings or livestreams, the Indexing API is excellent and fast for exactly those. For everything else, it’s the wrong tool.
IndexNow — real, useful, and not a Google tool
IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. gets conflated with “fast Google indexing” constantly, and it’s worth being blunt: IndexNow does not include Google. It’s an open protocol launched in October 2021 by Microsoft Bing and Yandex; you ping a participating endpoint when a URL changes and the engines share that notification with each other. Bing’s own IndexNow page lists adopters like Yandex, LinkedIn, Yahoo, eBay, Etsy, GitHub, Wix, Cloudflare, Yoast, and RankMath — Google is not among them.
So submitting via IndexNow will help Bing, Yandex, and other participating engines discover changed URLs faster. It has zero direct effect on Google indexing — Google was never in scope. “I submitted via IndexNow but Google still hasn’t indexed my page” is a category error, not a bug.
Even for the engines that do participate, IndexNow is a discovery notification, not an indexing guarantee. Its own documentation is careful about what a successful response means: “The HTTP 200 response code only indicates that the search engine has received your URL.” Received, not indexed.
Should you still implement it? Yes — it’s low-effort (host a key file, ping an endpoint) and it genuinely speeds Bing/Yandex discovery. Just don’t expect it to move Google.
The real gatekeeper: content quality
Strip away the tooling and the actual bottleneck is almost always quality. Google’s own language around the not-indexed statuses is editorial, not mechanical:
- “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.: The page was crawled by Google but not indexed. It may or may not be indexed in the future; no need to resubmit this URL for crawling.”
- “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.: The page was found by Google, but not crawled yet. Typically, Google wanted to crawl the URL but this was expected to overload the site; therefore Google rescheduled the crawl.”
The reps have consistently framed “Discovered/Crawled – currently not indexed” as a worthiness question. As John Mueller put it on the topic, moving pages from not-indexed to indexed basically means convincing Google it’s worthwhile to index more — since Google doesn’t yet have an understanding of an un-indexed URL, it pulls in the rest of the site to judge the URL’s potential context, and there’s nothing special or new about “discovered / not indexed.” His shorthand for what it takes, across several public statements, has been some variation of “awesomeness” — the page needs to be genuinely good.
In practice that means: thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count., near-duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling., low-effort or mass-produced pages, and pages that overlap heavily with better pages you already have are the usual suspects behind non-indexing. Consolidate or improve them rather than resubmitting them.
What about crawl budget?
For most sites, crawl budgetThe number of URLs an engine will crawl in a timeframe. is a red herring for indexing problems. Gary Illyes has said the vast majority of sites don’t need to worry about crawl budget at all — it becomes a real constraint only at large scale (roughly 1M+ pages changing weekly, or 10k+ daily). If you’re a small-to-midsize site and pages aren’t indexing, the culprit is far more likely quality or architecture than budget. (See crawl budget for the full treatment.)
How long does indexing take, and how to monitor
Honestly: a few days to a few weeks is normal, and new/low-authority sites are slower because Google relies on already-known pages and backlinks to find and value new URLs. A couple of quality backlinks and a clean sitemap matter disproportionately early on. Monitor with the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (aggregate, property-wide) and URL Inspection (single URL) rather than re-clicking Request Indexing — patience and improvement beat the button.
Where this fits
This is the actions companion to the Page Indexing report, which explains how to read the not-indexed statuses; this page is what to do about them. It sits inside the broader indexing stage of how search worksSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank., and leans on its siblings: canonicalization decides which URL from a duplicate cluster gets indexed at all, XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. and internal linking are the discovery channels, and crawl budget is the efficiency concern that occasionally (rarely) intersects. For the whole pipeline — discovery, crawling, rendering, indexing, serving — see the How Search Works clusterSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank..
AI summary
A condensed take on the Advanced version:
- IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. = crawl → judge → store. Getting indexed is a prerequisite for ranking, never a ranking factor. CrawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. are separate steps — a page can be crawled and still not indexed.
- You control discovery and crawl priority, not the indexing decision. The levers: internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. (no orphans), a clean XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags., and Request Indexing for a few individual high-priority URLs.
- Request Indexing (URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.): rate-limited, doesn’t guarantee or speed a quality decision, and “no need to resubmit.” Google’s own exact daily number is unpublished; ~10–15/day is anecdotal. For bulk, use a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing..
- XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. aids discovery only — “doesn’t guarantee that all the items in your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. will be crawled and indexed.” 50MB / 50,000-URL cap per file; audit indexed-vs-submitted in the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason..
- Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.: “Every page you care about should have a link from at least one
other page.” Use real
<a href>, keep important pages shallow, link to canonicals. - Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. is officially only for
JobPostingandBroadcastEvent-in-VideoObjectpages; default 200 publish requests/day. Using it elsewhere is unsupported and can be cut off without notice — Google does not use it as a general fast-index tool. - IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. is Bing/Yandex/others — Google does not participate. HTTP 200 there means “received,” not indexed. Implement it for Bing, not Google.
- Quality is the real gate. “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” is usually a quality signal; convince Google the page is worthwhile. Crawl budgetThe number of URLs an engine will crawl in a timeframe. is a non-issue for most sites (Illyes). Indexing takes days to weeks; new sites are slower.
Official documentation
Primary-source documentation from the search engines.
- Ask Google to Recrawl Your Website — Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. via URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version., the quota, and why resubmitting doesn’t help.
- URL Inspection tool — Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Help: daily limit, non-guarantee, and “use a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. for many pages.”
- What Is a Sitemap — when you need one, and the “doesn’t guarantee indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” line.
- Build and Submit a Sitemap — the 50MB / 50,000-URL cap and how to submit.
- Manage Your Sitemaps With Sitemap Index Files — splitting large sites across a sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file..
- SEO Link Best Practices for Google — crawlable links and “every page you care about should have a link.”
- How to Use the Indexing API — the scope:
JobPostingandBroadcastEvent-in-VideoObjectonly. - Requesting Approval and Quota (Indexing API) — the 200/day publish quota and the quality-based quota note.
- Page indexing report — Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. Help: the not-indexed statuses, including “no need to resubmit.”
Bing / Microsoft / IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it.
- Why IndexNow — Bing’s IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. overview and its published list of adopters (Google not among them).
- IndexNow Documentation — the protocol, and “the HTTP 200HTTP 200 OK is the standard 2xx success status code, meaning the server received, understood, and fulfilled the request and is returning the resource. It's the code every page you want indexed should return — but a 200 alone doesn't guarantee Google will index the page. response code only indicates that the search engine has received your URL.”
Quotes from the source
On-the-record statements from Google and Bing/IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it.. Each link is a deep link that jumps to the quoted passage on the source page.
Google — Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. (URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.)
- “To request a crawl of individual URLs, use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.. You must be an owner or full user of the Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. property to be able to request indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. in the URL Inspection tool.” — Google Search Central docs. Jump to quote
- “There’s a quota for submitting individual URLs and requesting a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. multiple times for the same URL won’t get it crawled any faster.” Jump to quote
- “Requesting a crawl does not guarantee that inclusion in search results will happen instantly or even at all. Our systems prioritize the fast inclusion of high quality, useful content.” Jump to quote
Google — sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and links
- “A sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” Jump to quote
- “Every page you care about should have a link from at least one other page on your site.” Jump to quote
Google — the Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. scope (the accuracy spine)
- “The Indexing API can only be used to crawl pages with either
JobPostingorBroadcastEventembedded in aVideoObject.” Jump to quote
Google — “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” (don’t resubmit)
- “The page was crawled by Google but not indexed. It may or may not be indexed in the future; no need to resubmit this URL for crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor..” — Page indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason., Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. Help. Jump to quote
IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. — what a 200 means
- “The HTTP 200 response code only indicates that the search engine has received your URL.” — IndexNow documentation. Jump to quote
My page isn’t indexed — what do I do?
Work top-down. Rule out the mechanical blockers before you touch any “fast-indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.”
tool, because none of those tools override a noindex, a robots block, a canonical,
or a quality problem.
Diagnose a page that won't index
Get-it-indexed checklist
A pass to confirm a page can be foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore., crawled, and judged fairly:
- URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. shows the page is not blocked by
noindexorrobots.txt. - The page’s
rel="canonical"points at itself (or you genuinely intend it to consolidate elsewhere). - The page is linked internally from at least one other relevant page — not orphaned — with descriptive anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page..
- Those internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. are real
<a href>elements, not JS-only click handlers. - If the page’s main content or links depend on JavaScript, they actually render — check the rendered HTML in URL Inspection, not just the raw server response.
- The page is reachable within a few clicks of the homepage.
- The URL is in a clean XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. (canonical, indexable URLs only) submitted in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results..
- The server returns 200 fast and stably (no 4xx/5xx, no soft-404).
- The content is genuinely useful and unique — not thin, duplicate, or overlapping a better page you already have.
- For a single high-priority URL, you clicked Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. once — and did not resubmit repeatedly.
- You are not relying on the Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. (unless it’s JobPosting / BroadcastEvent) or on IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. (which Google doesn’t use) to force Google.
- You’re monitoring via the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. and URL Inspection, not by re-clicking the button.
SOP: getting a new (or stuck) page indexed
A repeatable procedure. Run it in order — earlier steps gate the later ones.
- Inspect first. Open URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and enter the exact URL. Note: is it indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.? What canonical did Google choose? Is it blocked?
- Clear mechanical blockers. If there’s a
noindex, remove it (if you want the page indexed). If it’s blocked inrobots.txt, unblock it so Google can read the content. If the chosen canonical is a different URL, decide which URL should be the indexed one and align the tag, redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., and internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to it. - Make it discoverable. Add at least one internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. from a relevant, crawled
page (ideally a strong one), using descriptive anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. and a real
<a href>. Confirm the page is within a few clicks of the homepage. - Add it to the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.. Ensure the URL appears in a clean XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. that lists only canonical, indexable URLs, and that the sitemap is submitted in Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance..
- Judge the content honestly. Is it unique and worthwhile, or thin/duplicate/ overlapping? If it’s weak, improve or consolidate it before expecting indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — this is the actual gate.
- Request indexing once (optional, for individual high-priority URLs). In URL Inspection, click “Request Indexing” a single time. Do not resubmit the same URL repeatedly — it won’t crawl faster.
- Wait and monitor. Give it days to a few weeks. Check the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (filter by sitemap to see indexed-vs-submitted) and re-run URL Inspection. Do not re-click Request Indexing as a reflex.
- If still stuck after weeks, treat it as a quality/architecture problem, not a tooling problem: strengthen internal links, earn a couple of relevant backlinks, and improve the content. Special cases (job postings, livestreams) can use the Indexing API; Bing/Yandex discovery can use IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. — neither affects Google indexing of normal pages.
Indexing anti-patterns
The mistakes that show up over and over — what’s wrong, and what to do instead.
Spam-clicking “Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” Why it’s wrong: Google explicitly says re-requesting a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. of the same URL won’t get it crawled any faster, and its status glossary says “no need to resubmit.” Repeated clicks accomplish nothing and burn your daily quota. Do instead: Request once for individual high-priority URLs; for anything beyond a handful, use a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and fix discovery/quality.
Treating the Google Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. as a universal fast-index hack.
Why it’s wrong: It’s officially only for JobPosting and BroadcastEvent-in-
VideoObject pages. Pointing it at normal pages is unsupported, can “work” briefly,
and can be cut off without notice.
Do instead: Use it only for those two content types. For everything else rely on
sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., and quality.
Expecting IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. to speed up Google. Why it’s wrong: IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. is a Bing/Yandex protocol; Google doesn’t participate. Submitting there does nothing for Google indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.. Do instead: Implement IndexNow for Bing/Yandex discovery, and use Google’s own tools (sitemaps, URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.) for Google.
Assuming a sitemap guarantees indexing. Why it’s wrong: Google says a sitemap “doesn’t guarantee that all the items in your sitemap will be crawled and indexed” — it only aids discovery. Do instead: Keep the sitemap clean (canonical, indexable URLs only) and fix the underlying quality/architecture; audit indexed-vs-submitted in the Page IndexingThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. report.
Buying “guaranteed indexing” services. Why it’s wrong: No legitimate mechanism forces Google to index a page outside its own tools, and none override the quality bar. These services do nothing you can’t do yourself. Do instead: Spend the effort on content quality, internal linkingLinks between pages on the same site., and a couple of real backlinks.
Blaming crawl budgetThe number of URLs an engine will crawl in a timeframe. on a small site. Why it’s wrong: Per Gary Illyes, the vast majority of sites don’t have crawl-budget problems; for small-to-midsize sites the real culprit is quality or architecture. Do instead: Diagnose quality and orphaning first; only investigate crawl budget at genuinely large scale.
Bloating the sitemap to look bigger. Why it’s wrong: Stuffing low-value or non-canonical URLs in wastes crawl attention and muddies your reporting — more URLs isn’t more authority. Do instead: List only the canonical, indexable pages you actually want indexed.
Getting indexed — cheat sheet
Which tool for which job
| Goal | Right tool | Not |
|---|---|---|
| Get a few important URLs looked at | URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. → Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. (once) | Bulk resubmitting |
| Get many URLs discovered | XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. | Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. |
| Fast-index a job posting or livestream page | Google Indexing APIThe Google Indexing API lets site owners notify Google directly when a URL is added, updated, or removed. Google officially supports it only for pages with JobPosting or BroadcastEvent (livestream) structured data — not for general content. | Using it for normal pages |
| Speed up Bing/Yandex discovery | IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. | Expecting it to reach Google |
| See why a page isn’t indexed | URL Inspection (single) / Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (aggregate) | Guessing |
What each channel actually does
| Channel | Effect on Google | Guarantee? |
|---|---|---|
| Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. | Discovery + crawl priority | No |
| XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. | Discovery aid | No — “doesn’t guarantee… crawled and indexed” |
| Request Indexing | Queues one URL for (re)crawl | No — doesn’t speed a quality decision |
| Indexing API | Fast crawl for JobPosting / BroadcastEvent only | No; out of scope elsewhere |
| IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. | Nothing (Google doesn’t use it) | N/A for Google |
| Content quality | The actual gate | Still no, but it’s the deciding factor |
Fast facts
- Indexing API scope: JobPosting + BroadcastEvent-in-VideoObject only; default 200 publish requests/day.
- Request Indexing: no official daily number published; ~10–15/day is anecdotal; “no need to resubmit.”
- SitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. file cap: 50MB / 50,000 URLs — split with a sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file. beyond that.
- IndexNow HTTP 200 = “received,” not indexed. Google does not participate.
- Typical indexing time: a few days to a few weeks; new/low-authority sites are slower.
Find orphan pages (pages nothing links to)
Orphans are a top cause of non-indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.. Compare the URLs in your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. against the URLs your own crawl actually reaches — anything in the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. that your crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. never foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. via internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. is effectively orphaned.
Shell — pull every URL out of an XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.
# Extract <loc> URLs from a sitemap (works for a single sitemap file)
curl -s https://example.com/sitemap.xml \
| grep -oE '<loc>[^<]+</loc>' \
| sed -E 's~</?loc>~~g' \
| sort -u > sitemap-urls.txt
wc -l sitemap-urls.txtExport the “crawled/found via links” URL list from Screaming Frog or Ahrefs Site
Audit as crawled-urls.txt, then:
# URLs in the sitemap that the crawl never reached = likely orphans
comm -23 sitemap-urls.txt <(sort -u crawled-urls.txt) Check a page’s indexability signals from the terminal
Before blaming discovery, confirm the page isn’t quietly telling Google to stay away.
# noindex in the response header?
curl -sI https://example.com/my-page/ | grep -i x-robots-tag
# noindex / canonical in the HTML <head>?
curl -s https://example.com/my-page/ \
| grep -iE '<meta[^>]+robots|rel=["'"'"']canonical'Chrome DevTools Console — indexability at a glance
Paste into the Console on the live page to surface the signals that decide indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.:
// Meta robots, canonical, and whether this URL matches its canonical
(() => {
const robots = document.querySelector('meta[name="robots"]')?.content || '(none)';
const canonical = document.querySelector('link[rel="canonical"]')?.href || '(none)';
console.log('meta robots:', robots);
console.log('canonical:', canonical);
console.log('this URL:', location.href);
console.log('self-canonical?', canonical === location.href);
})();Bookmarklet — jump straight to URL Inspection for the current page
Save as a bookmark; click it on any page to open that URL in your Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. property’s inspection tool (replace the property prefix with yours).
javascript:(function(){var p='https://example.com/';var u=location.href;window.open('https://search.google.com/search-console/inspect?resource_id='+encodeURIComponent('sc-domain:example.com')+'&id='+encodeURIComponent(u),'_blank');})();Adjust resource_id to your Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. property (a sc-domain: property or
a URL-prefix property). The bookmarklet just opens the inspection URL — it doesn’t
request indexing on its own. Patrick's relevant free tools
- IndexNow Submitter — Validate and explicitly submit a same-host URL list to IndexNow; Google does not use IndexNow.
- Google Index Checker — Check one URL’s observable indexability blockers, or reconcile sitemap, crawl, and supplied Search Console evidence across a URL set before verifying Google’s actual state in URL Inspection.
- Canonicalization Checker — Audit HTML and HTTP canonical signals, test the canonical target, and identify observable conflicts that can cause Google to choose a different URL.
Tools for getting (and checking) indexing
- URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. (Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.) — the single-URL source of truth: is it indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., what canonical did Google choose, and the Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. button.
- Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (GSC) — aggregate, property-wide view of indexed vs. not-indexed, grouped by reason; filter by sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. to audit indexed-vs-submitted.
- Sitemaps reportThe Google Search Console report where you submit sitemaps and watch how Google processes them — type, last read date, status, and how many URLs were discovered. It confirms Google read your list; it doesn't prove anything got indexed. (GSC) — submit and monitor your XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags..
- Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. — Bing’s index coverage and URL submission, plus IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it..
- IndexNow — one-time-setup protocol to ping Bing/Yandex on URL changes (not Google).
- Screaming Frog SEO Spider / Ahrefs Site Audit — crawl your site to find orphan pages, broken internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., and non-canonical/noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. URLs before Google does.
- Ahrefs Webmaster Tools — free crawl + audit for sites you verify.
Common indexing issues
Lookup table for what you’re actually seeing, in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. or in the wild — matched to the likely cause and the fix to try.
GSC shows “Discovered – currently not indexed”
Symptom: URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. or the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. shows the page as “Discovered – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” — Google knows the URL exists but hasn’t crawled it.
Likely cause(s): Nothing (or too little) links to it internally, so it isn’t prioritized; or Google expected the crawl to overload the site and rescheduled it.
Fix: Add internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. from relevant, already-indexed pages with descriptive anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page., and confirm the URL is in a clean XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.. Confirm by re-running URL Inspection after a few days — resubmitting via Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. doesn’t speed this up.
GSC shows “Crawled – currently not indexed”
Symptom: Google fetched the page (you can see it was crawled) but it’s still not indexed, sometimes for weeks.
Likely cause(s): Content is thin, duplicate, or overlaps heavily with a better page you already have — this is the most common cause of this exact status.
Fix: Improve or consolidate the page rather than resubmitting; Google’s own guidance is explicit that there’s “no need to resubmit this URL for crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor..” Confirm by checking whether the page’s unique value versus your other pages has actually changed, then re-check status after a few weeks.
”Request Indexing” seems to do nothing no matter how many times you click it
Symptom: Clicked Request Indexing multiple times on the same URL over days; status hasn’t changed.
Likely cause(s): Re-requesting the same URL doesn’t get it crawled any faster — you’re burning quota, not fixing anything. The real blocker is elsewhere (mechanical or quality).
Fix: Stop resubmitting. Run URL Inspection once to see what Google actually reports (blocked, canonicalized elsewhere, or a not-indexed status), then fix that root cause instead.
Page looks fine in “View Source” but content/links are missing from the index
Symptom: The raw HTML response looks thin or empty, but the page looks normal in a browser; internally-linked pages that only appear in JS-rendered navigation aren’t getting discovered or indexed.
Likely cause(s): The page relies on client-side JavaScript to inject its main content or links (an app-shell pattern), and Google’s renderer either hasn’t gotten to it yet or failed — commonly because a script, stylesheet, or API call the render depends on is blocked or erroring. RenderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. happens as a separate, queued step after the initial crawl, so a 200 response doesn’t mean Google saw the rendered content.
Fix: Check the rendered HTML in URL Inspection (not just “View Source”),
confirm no JS/CSS resources the page needs are blocked by robots.txt, and fix any
console errors that stop the render from completing. Prefer server-rendered or
statically-generated markup for content and links you need indexed quickly.
A robots.txt-blocked URL shows up in search results anyway
Symptom: A URL you blocked in robots.txt still appears in Google’s index —
usually as a bare URL with no title/snippet.
Likely cause(s): Blocking crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. in robots.txt doesn’t remove a URL from
the index if other pages still link to it — Google can index the URL itself
without reading its content.
Fix: If you want it fully out, use noindex (which requires Google to be able
to crawl the page to see the tag) or remove the incoming links, not just a
robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. block. Confirm with the robots-txt-tester
and a fresh URL Inspection.
Page is indexed under a different URL than expected
Symptom: URL Inspection shows “Alternate page with proper canonical tagA Google Search Console Page Indexing status meaning a page is a duplicate or alternate version that correctly points its canonical at another, indexed page. It's normal, healthy behavior — Google says there is nothing you need to do.” or “Duplicate, Google chose different canonical than userA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead.” for a URL you wanted indexed as-is.
Likely cause(s): Your rel="canonical", internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., or both point at a
different URL, so Google consolidated onto that one instead.
Fix: Decide which URL should be canonical and align the tag, internal links, and any redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. to point at it consistently. Confirm with the canonical-checker and re-run URL Inspection.
Sitemap submitted, but the Page Indexing report shows a big gap between submitted and indexed
Symptom: Hundreds or thousands of URLs in your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., but the indexed count is much lower.
Likely cause(s): The sitemap itself doesn’t force indexing — it’s a discovery aid. A large gap usually points to quality or architecture problems across many URLs, not a submission problem.
Fix: Filter the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. by that sitemap to see the specific not-indexed reasons, then treat it as a quality/architecture audit — not a reason to resubmit the sitemap.
Validation tests
Proof that an indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. fix actually took effect — run these after making a change, not instead of one.
Confirm a noindex was actually removed
Test to run: curl -sI https://example.com/page/ | grep -i x-robots-tag and
check the page’s <head> for <meta name="robots"> (or use the
render-gap tool if the tag is added client-side).
Expected result: No noindex in either the HTTP header or the rendered HTML.
Failure interpretation: If noindex still appears in the rendered HTML but
not the raw response (or vice versa), a template, CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. setting, or CDN rule is
still injecting it — check both source and rendered output.
Monitoring window: Immediate — this is a status check, not something that needs to propagate.
Rollback trigger: If removing noindex was intentional and the page still
doesn’t index after several weeks, that’s not a rollback signal — move to the
quality diagnosis instead.
Confirm robots.txt no longer blocks the URL
Test to run: robots-txt-tester against the exact URL path.
Expected result: The tool reports the URL as allowed for GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer..
Failure interpretation: Still blocked means a Disallow rule (yours or an
inherited one) still matches the path — check for overly broad wildcard rules.
Monitoring window: Immediate for the robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. check itself; allow a few days
for GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. to re-fetch robots.txt (it’s cached) before expecting a crawl.
Rollback trigger: N/A — this is a one-way fix; re-block only if you deliberately want the URL excluded again.
Confirm the canonical is self-referencing (or points where you intend)
Test to run: canonical-checker on the URL, and cross-check “Google-selected canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead.” in URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version..
Expected result: The declared canonical matches the URL you want indexed, and Google’s selected canonical agrees with it.
Failure interpretation: If Google’s selected canonical still differs from your declared one, internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. or duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. are overriding your signal — Google treats the canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. as a hint, not a directive.
Monitoring window: 2–4 weeks — canonical selection can take time to update after you fix the signals.
Rollback trigger: If, after 4+ weeks, Google’s selected canonical still disagrees, revisit whether the URLs are different enough that they shouldn’t be consolidated at all.
Confirm a requested URL actually gets indexed
Test to run: URL Inspection → Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. (once) on the specific URL, then re-run URL Inspection later.
Expected result: Status moves from “not indexed” to “URL is on Google,” and a
site: search (or the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.) shows the page.
Failure interpretation: Still not indexed after a reasonable wait, with no mechanical blockers found, means it’s a quality/worthiness decision — resubmitting again won’t change the outcome.
Monitoring window: Days to a few weeks per Google’s own guidance; check again before assuming it failed.
Rollback trigger: Not applicable — there’s nothing to undo; if it stays not-indexed, address content quality instead.
Confirm a sitemap fix improved the indexed-vs-submitted rate
Test to run: Page Indexing report in GSC, filtered by the specific sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing..
Expected result: The indexed count for that sitemap trends up relative to submitted URLs over successive checks.
Failure interpretation: A flat or worsening ratio after cleanup means the remaining not-indexed URLs share a quality or architecture problem, not a discovery problem.
Monitoring window: 2–4 weeks for a meaningful trend — Page Indexing data itself lags by a few days.
Rollback trigger: N/A — if the sitemap change was cleanup (removing non-canonical URLs), there’s nothing to roll back even if the ratio doesn’t move quickly.
How to measure indexing health
The standing KPIs for whether your discovery and quality work is actually landing — not one-time checks, but numbers to watch over time.
Indexed vs. submitted ratio
What it tells you: How much of your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. is actually making it into the indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — the single best proxy for “is my discovery/quality pipeline working.”
How to pull it: Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. in GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., filtered by sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. (IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. → Pages → filter by sitemap).
Benchmark / realistic range: Depends heavily on site type and age — a small, well-linked site can sit near 90–100%; large or newer sites commonly run lower. There’s no universal “good” number; establish your own baseline and watch the trend rather than chasing a fixed target.
Cadence: Monthly, or after any sitemap/content cleanup.
Not-indexed reason breakdown
What it tells you: Why pages aren’t indexed — separates mechanical problems (noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., blocked, canonicalized elsewhere) from quality problems (“Crawled – currently not indexed”), which need completely different fixes.
How to pull it: Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. in GSC — the “Why pages aren’t indexed” table breaks URLs down by reason.
Benchmark / realistic range: No universal benchmark; the useful signal is the mix shifting toward mechanical (fixable in minutes) versus quality (needs real content work) over time.
Cadence: Monthly.
Time-to-index for new pages
What it tells you: How quickly new content clears discovery + the quality bar — a proxy for overall site authority and internal-linking health.
How to pull it: Note publish date, then check URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. periodically (or
watch for the URL appearing in the Page Indexing report / a site: search) to log
days-to-index for a sample of new pages.
Benchmark / realistic range: Days to a few weeks is normal per Google’s own guidance; new or low-authority sites run slower because Google leans on existing backlinks and known pages to prioritize crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor.. Depends on site authority — track your own trend rather than comparing across sites.
Orphan page count
What it tells you: How many pages have no internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. pointing at them — a leading real-world cause of non-indexing that’s entirely within your control.
How to pull it: Crawl the site with the site-audit-lite
tool or Screaming Frog/Ahrefs Site Audit and compare against your sitemap URL
list (see the Scripts tab for a comm-based diff).
Benchmark / realistic range: The honest target is zero for any page you want indexed — this isn’t situational the way the other metrics are.
Cadence: After every content push, or monthly for larger sites.
Prompts for indexing diagnosis
Use these with exported Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. data, crawl data, or page details. Do not ask the model to guess whether a URL is indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.; verify that in Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance..
Triage a set of excluded URLs
Classify these URLs by Search Console exclusion reason. For each class, separate discovery, crawl-access, canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it., and content-quality causes. Recommend the smallest test that would confirm each suspected cause. Do not treat Request IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. as a sitewide fix. Data: [paste URL, status, canonical, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., internal-link count].
Review one page before requesting indexing
Review this page’s indexability evidence: HTTP status, robots access, meta/X-Robots directives, rendered canonical, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. presence, and internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. Return a pass/fail table, then identify anything that would make an indexing request premature. Evidence: [paste checks].
Turn a recurring issue into a template fix
These URLs share an indexing exclusion. Group them by page template and identify which fixes belong in the CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms./template rather than page-by-page edits. Rank the fixes by affected URLs and confidence, and specify how to validate each deployment. Sample: [paste rows].
Resources worth your time
My related writing
- How to Get Google to Index Your Website — Ahrefs’ practical guide to request indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., and content quality (a good companion to this page).
- What “Crawled – Currently Not Indexed” Means In Google Search Console — why that status is fundamentally a quality signal.
- How to Fix “Discovered – currently not indexed” — the Ahrefs guide I reviewed: orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site., crawl prioritization, backlinks as a value signal.
- The Beginner’s Guide to Technical SEO — where indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. fits in the bigger picture.
- My author archive on the Ahrefs blog — more of my technical-SEO writing on crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexing.
My speaking
- How Search Works (SlideShare) — my walkthrough of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., indexing, and ranking, and why crawled ≠ indexed. (My standing disclaimer applies: “This is my understanding of systems… not going to be 100% complete or accurate.”)
Official
- Ask Google to Recrawl Your Website and the URL Inspection tool — Request Indexing, its quota, and the non-guarantee.
- How to Use the Indexing API — the JobPosting/BroadcastEvent-only scope.
- Why IndexNow (Bing) and IndexNow documentation.
From around the industry
- Google: Sites Need To Be Worthwhile To Be Indexed (Search Engine Journal) — John Mueller on “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” as a worthiness question.
- Google Search Console request indexing limit removed from interface (Search Engine Land, 2018) — the history of Google keeping the manual quota opaque.
- Should I Worry About Crawl Budget? (Search Off the Record, Aug 2022) — Gary Illyes on why most sites don’t need to worry about crawl budgetThe number of URLs an engine will crawl in a timeframe..
- What is “Crawled – currently not indexed” in Search Console? (Yoast) — a plain-language take on the same status.
- How To Fix “Crawled – Currently Not Indexed” in GSC (Onely) — a thorough troubleshooting walkthrough.
- r/TechSEO — the community for crawl/index debugging.
Test yourself: getting indexed
Five questions on how pages actually get indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. by Google. Pick an answer for each, then check.
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Added JavaScript rendering as a documented reason pages get crawled but not indexed, verified directly against Google's JavaScript SEO basics doc.
Change details
- Advanced
Added a JS-rendering/app-shell cause to the "why a page isn't indexed" diagnostic list, a matching checklist item, and a new troubleshooting entry for content/links missing because a render never completed.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 16, 2026.
Editorial summary and recorded change details.Summary
Added a visual that separates crawling, indexing, and serving so the diagnostic flow is easier to scan.
Change details
- Advanced
Added the crawl-index-serve pipeline diagram.