Page Indexing Report (GSC)
How Google Search Console's Page Indexing report (formerly Index Coverage) works — indexed vs. not indexed, the Source column, every status, Validate fix, and report lag.
The Page Indexing report (formerly Index Coverage, now labeled 'Pages' under 'Indexing' in Google Search Console) shows how many of your URLs are indexed vs. not indexed, then groups the not-indexed ones by reason in the 'Why pages aren't indexed' table. It's an aggregate, property-wide report — for a single page you use URL Inspection instead. The single most useful mindset: not indexed is not necessarily bad — canonical/duplicate, noindex, robots, and intentional 404s are all correct outcomes, and Google says you should only expect your canonical pages to be indexed. Filter by Source = Website to find what you can actually fix, work the pre-sorted table top-down, and use Validate fix (~2 weeks) to ask for a recrawl. Sites under 500 pages probably don't need it. This hub explains how to read the report and links down to every individual status.
Evidence for this claim The Page indexing report shows indexed and not-indexed pages known to Google and groups non-indexing by reason. Scope: Current Page indexing report terminology and behavior. Confidence: high · Verified: Google Search Console: Page indexing report Evidence for this claim The Page indexing report is for site-wide patterns; URL Inspection provides the indexed and live-test details for an individual URL. Scope: Current distinction between Page indexing and URL Inspection. Confidence: high · Verified: Google Search Console: URL Inspection toolTL;DR — The Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. tells you how many of your pages Google has indexed (can show in search) versus not indexed, and for the not-indexed ones it tells you why. It used to be called the Index Coverage report. “Not indexed” sounds scary but usually isn’t — a lot of those URLs are supposed to be left out.
What the Page Indexing report is
When you open Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and click Pages under IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. in the left menu, you land on the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.. Google describes it simply: it lets you “see which pages Google can find and index on your site, and learn about any indexing problems encountered.”
It splits all the URLs Google knows about on your site into two buckets:
- Indexed — these pages can appear in Google Search.
- Not indexed — these pages aren’t in the index, either because something’s broken or because there’s a perfectly good reason (the page is a duplicate, you blocked it, you told Google not to index it, and so on).
Below that, a table called “Why pages aren’t indexed” lists the reasons and how many URLs fall under each one. That’s the part you actually work from.
Why “not indexed” usually isn’t a crisis
Here’s the thing most people miss the first time they open this report: not indexed doesn’t mean broken. Google says it plainly in the report’s own help — “Not indexed is not necessarily bad.” You shouldn’t expect every URL on your site to be indexed. Tag pages, filtered versions of a product listing, old redirected URLs, duplicates — Google leaving those out is the report working correctly.
So when you see a big “not indexed” number, don’t panic. Read which reasons are behind it before you change anything.
It’s the wrong tool for checking one page
This is an overview of your whole site. If you want to know whether one specific URL is indexed, this isn’t the report for it — you use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. (the search bar at the top of Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance.). Google is explicit about this. The report is for trends and groupings; URL Inspection is for one page at a time.
Do you even need it?
If your site has fewer than 500 pages, Google says you “probably don’t need to use
this report” — a site:yoursite.com search in Google is enough to spot-check
what’s indexed. The report earns its keep on bigger sites where you can’t eyeball
everything.
Want the full version — how to read every column, what all 16 statuses mean, how “Validate fix” works, and how to diagnose a sudden drop in indexed pages? Switch to the Advanced tab.
Evidence for this claim The Page indexing report shows indexed and not-indexed pages known to Google and groups non-indexing by reason. Scope: Current Page indexing report terminology and behavior. Confidence: high · Verified: Google Search Console: Page indexing report Evidence for this claim The Page indexing report is for site-wide patterns; URL Inspection provides the indexed and live-test details for an individual URL. Scope: Current distinction between Page indexing and URL Inspection. Confidence: high · Verified: Google Search Console: URL Inspection toolTL;DR — The Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report (formerly Index Coverage; “Pages” under “IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” in GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.) is an aggregate view of every URL Google knows about on your property, split into Indexed and Not indexed, with the not-indexed URLs grouped by reason in the “Why pages aren’t indexed” table. It’s not for single pages — that’s URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.. The mindset that matters: not indexed is not necessarily bad — canonical/duplicate, noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., robots, and intentional 404s are correct outcomes; Google says you should only expect your canonical pages to be indexed. Filter by Source = Website to find what you can fix, work the pre-sorted table top-down, and use Validate fix (~2 weeks typical) to request a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial.. The report lags reality. Sites under 500 pages probably don’t need it.
What the report actually is (and its old name)
Google’s one-liner is the cleanest definition: the report lets you “see which pages Google can find and index on your site, and learn about any indexing problems encountered.” More precisely, it “shows the Google indexing status of all URLs that Google knows about in your property.” It’s an aggregate, property-wide report — not a per-URL tool.
If you’ve been doing SEO for more than a couple of years, you knew this as the Index Coverage reportThe Google Search Console report (renamed from \"Index Coverage\" in 2022) that shows which URLs Google has indexed, which it hasn't, and why. It splits your known URLs into Indexed and Not indexed, grouping the not-indexed ones by reason.. Google renamed it to Page indexing in 2022 (it shows as “Pages” in the left nav), so a lot of older guides — and a lot of people’s muscle memory — still call it “Coverage.” Same report. I’m spelling that out because people searching the old name need to know they’re in the right place. (The rename was first spotted in a Google I/O 2022 demo and rolled out later that year; before that, a January 2021 data-quality update had already reworked several of the statuses, which is why some old screenshots won’t match what you see today.)
How to read the report
The summary view has a few parts worth understanding before you start clicking.
Indexed vs. Not indexed. The two totals above the chart. Google notes these are “complete and accurate from Google’s perspective, but small discrepancies can occur for various reasons” — so don’t expect them to perfectly equal your URL count. You can click “View data about indexed pages” for the historical indexed count and a sample of up to 1,000 indexed URLs.
The “Why pages aren’t indexed” table. This is the heart of it. It “shows issues that prevented URLs from being indexed on your site,” sorted by what Google thinks are the most important issues to address. Start at the top — the ordering is already a priority list.
The Source column — your fix filter. Each issue is sourced to either Google or the website, and Google’s guidance is direct: “The Source value in the table shows whether the source of the issue is Google or the website. In general, you can fix only issues where the source is listed as ‘Website’.” So Source = Website plus a validation state of “failed” or “not started” is your real to-do list. Google-source statuses (like a duplicate it chose to consolidate) usually aren’t yours to “fix.”
The “Improve page experienceGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data.” table. A separate table for “issues that didn’t prevent page indexing, but we recommend that you fix them.” These are warnings, not blockers.
The sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. filter. Above the chart you can scope the report to All known pages, All submitted pages, Unsubmitted pages only, or a specific sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.. (Note: “A URL is considered to be submitted by a sitemap even if it was also discovered through some other mechanism.”) This filter is genuinely useful — more on that under validation.
The example URL lists are capped. When you open a status, the sample of affected URLs is “limited to 1,000 items, and isn’t guaranteed to show all URLs in a given status, even when less than 1,000 items.” Treat the examples as a sample, not a complete export of everything in that bucket.
Triage order, if you only have ten minutes. Scope the report to the sitemap that represents your intended URL inventory, so you’re reading it against the pages you actually want indexed. Scan the totals for any unexpected change first — a drop or spike is more urgent than a steady-state number. Filter to Source = Website for the issues you can actually act on. Within that filtered list, prioritize whatever’s hitting a business-critical template or URL pattern over an isolated one-off. Then fix the whole pattern — not just one URL — before you validate.
Report vs. URL Inspection — use the right one
The biggest practical distinction on this page. Google: “This report isn’t used to investigate the index status of specific pages. To find the index status of a specific page, use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version..”
- Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. = aggregate trends, grouped by reason, property-wide.
- URL Inspection = one URL’s live/indexed state, the canonical Google chose, and a “Test live URL” option.
When a status confuses you, pull a sample URL out of it and inspect that URL. The two tools are meant to be used together.
A note on lag — the report is behind reality
This one saves a lot of false alarms. The report doesn’t update in real time; it reflects the last time Google crawled each URL, and the counts move when Google recrawls (Google: it “updates your instance count whenever it crawls a page with known issues, whether or not you explicitly requested fix validation”). Google’s John Mueller has described the indexing report as lagging behind — his read is that it’s more a matter of timing, where URLs show up in the report and then get indexed over time, so the report is essentially catching up to a state that’s already changed. Don’t over-react to a number that may just be stale.
One limitation worth knowing before you lean on URL Inspection’s “Test live URL” to settle an argument: the live test confirms whether Google can currently crawl and index that URL right now — it doesn’t tell you which canonical Google will choose among a set of duplicates. Canonical selectionHow search engines pick one canonical URL among duplicates and consolidate signals onto it. is a separate indexing decision made against the indexed data, not something the live test evaluates. For canonical/duplicate statuses specifically, the indexed result in URL Inspection (not the live test) is the signal to trust, and it can still lag your latest change.
Reconcile sitemap intent, a current crawl, and an exported GSC state before treating a lagging status as current with my free Indexation Reconciler Free
- Load the current sitemap and crawl data alongside the relevant Search Console export.
- Separate evidence that agrees from conflicts caused by a recent directive change or stale report data.
- Inspect representative URLs and wait for recrawling before concluding that the current page state failed.
The focused Indexation Reconciler row is for https://example.com/b. It says Sitemap: Yes, Current crawl: noindex, GSC: indexed. The action is to review the new blocking directive and whether the export is stale.
Where to go next — every status, grouped
The “Why pages aren’t indexed” table is the map to the rest of this section. Each status below is its own deep dive (they’re in the sidebar too). I’ve grouped them by what’s actually going on, because the right reaction is completely different depending on the group — some you fix, some you confirm and leave alone.
Not indexed — Google’s choice (often nothing to fix)
- Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal. — Google found the URL but hasn’t crawled it yet. Usually a crawl-demand or site-quality signal, not a one-page bug.
- Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error. — Google crawled the page and chose not to index it (yet). A large bucket here often points at quality or duplication, site-wide.
Blocked by you (intentional — confirm it’s deliberate)
- Blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing. — your robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. told Google not to crawl it.
- URL marked ‘noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.’ — Google found a
noindexdirective when it tried to index the page.
HTTP errors (these you usually fix)
- Blocked due to unauthorized request (401) — the page asked GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. to authenticate.
- Blocked due to access forbidden (403) — access wasn’t granted to the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
- Blocked due to other 4xx issueA Google Search Console Page Indexing status: Googlebot got a 4xx client error not covered by another issue type (e.g. 400, 405, 410, 429, 451) when crawling the URL, so the page can't be indexed. — a 4xx that isn’t one of the named ones.
- Server error (5xx) — your server returned a 500-level error when the page was requested.
- Not found (404) — the URL returned a 404 when requested.
Canonical & duplicates (mostly correct — verify the chosen canonical)
- Alternate page with proper canonical tagA Google Search Console Page Indexing status meaning a page is a duplicate or alternate version that correctly points its canonical at another, indexed page. It's normal, healthy behavior — Google says there is nothing you need to do. — this page correctly points at its canonical, which is indexed. Working as intended.
- Duplicate without user-selected canonicalA Google Search Console Page Indexing status: Google found this page to be a duplicate, you didn't declare a canonical, so Google chose a different page as the canonical — and this URL isn't indexed. — Google clustered it as a duplicate and chose a different page as canonical, and you never declared one.
- Duplicate, Google chose different canonical than userA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead. — you declared a canonical, but Google picked a different URL. (See canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. — the fix is aligning your signals, not adding a stronger tag.)
RedirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. (one is normal, one is a bug)
- Page with redirectA Google Search Console Page Indexing status for a URL that redirects elsewhere. It's not indexed by design because it's a redirect — the destination is a separate question, and Google says the target may or may not end up indexed — usually expected, not an error (unlike the separate \"Redirect error\"). — a non-canonical URL that redirects to another page. This is normal; the destination is what gets indexed.
- Redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\" — a broken redirect: a chain that’s too long, a loop, an empty or bad target URL, or one that exceeds the max URL length. This one you fix.
Warnings (indexed, but worth attention)
- Indexed, though blocked by robots.txtA Google Search Console Page Indexing warning: Google indexed the URL anyway despite your robots.txt disallowing crawling it — Google names other pages linking to it as the likely path. robots.txt blocks crawling, not indexing. — the page got indexed despite your
robots.txt block (links to it were enough). Google can’t read its content or any
noindexyou put there — a classic crawl-vs-index gotcha. - Page indexed without contentA Google Search Console status meaning the URL is in Google's index, but Googlebot couldn't read any usable content from it — possible causes include a server/CDN block, cloaking, an empty render, or an unsupported format. Not the same as a page-level robots.txt block. — indexed, but Google couldn’t read meaningful content (often a renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. or cloaking-like issue).
I’m naming each status the way Google labels it so the per-status articles auto-link as they publish. The recurring theme: the canonical/duplicate, noindex, robots, and intentional 404 buckets are frequently correct — Google literally says “You should not expect all URLs on your site to be indexed, only the canonical pages.” The buckets to actually chase are the HTTP errors, redirect errors, and unexpectedly large “crawled/discovered not indexed” piles.
Fixing and validating
Once you’ve fixed the real issues (Source = Website), you tell Google to re-check:
- Fix every instance of the issue on your site first.
- Open the issue details and click “Validate fix.”
- Don’t click it again until validation succeeds or fails.
Timing: “Validation typically takes up to about two weeks, but in some cases can take much longer, so please be patient.” The request moves through states (Not started → Started → Looking good → Passed, or Failed, or N/A if Google found the issue fixed on its own before you ever started). You don’t strictly have to click validate at all — Google can detect fixes on its own — but validation gives you a tracked result.
The speed trick: validate against a subset. Submit a sitemap of just your most important pages, filter the report to that sitemap, then request validation — Google notes “a validation request against a subset of your affected URLs can complete faster.”
Diagnosing drops, spikes, and “more not-indexed than indexed”
A few patterns Google itself calls out, and that I look for first:
- Indexed pages dropped with no errors showing. You probably blocked existing
pages — a new robots.txt rule, a
noindex, or a login requirement. Look for a matching spike in a non-indexed status. - More not-indexed than indexed. Usually a robots.txt rule blocking a big section,
or a flood of duplicates from filter/sort parameters (
?type=dress,?color=green,?sort=price). That’s a faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals. and canonicalization problem more than an indexing one. - A sudden error spike. Often a template change that introduced an error across many URLs, or a submitted sitemap full of URLs that are blocked or noindexed.
A few things to keep straight
- Indexed ≠ ranking. Indexed only means eligible to appear in Search. Whether it ranks depends on the query and many other factors.
- 100% coverage isn’t the goal. Per Google, only your canonical pages should be indexed — a healthy site has plenty of intentionally not-indexed URLs.
- Validate fix doesn’t re-index instantly. It queues a recrawl of known affected URLs; budget ~2 weeks.
- GSC data has limits. Example lists cap at 1,000; totals can drift slightly; and the report lags. (My own GSC data study — Anonymized Queries Make Up Nearly Half of GSC Traffic — is a good reminder that Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. is a powerful but incomplete picture.)
A one-liner on Bing
Bing has no single “Page Indexing” equivalent. The closest aggregate view is Site Explorer in Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. (browse your site as a folder tree broken down by indexed, error, warning, and excluded pages), and URL Inspection in Bing covers the single-URL case. Use Site Explorer for the bird’s-eye view, URL Inspection for one page.
This page is the hub for the Page Indexing report; the per-status deep dives sit under it. For the whole pipeline — discovery, crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., rendering, indexing, and serving — see the How Search Works clusterSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank..
AI summary
A condensed take on the Advanced version:
- What it is: the Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report (formerly Index Coverage; “Pages” under “IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.”). It shows the indexing status of every URL Google knows about on your property, split into Indexed and Not indexed.
- It’s aggregate, not per-URL. For a single page, use URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. — Google says the report “isn’t used to investigate the index status of specific pages.”
- Not indexed is not necessarily bad. Canonical/duplicate, noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., robots, and intentional 404s are correct outcomes; Google says you should only expect your canonical pages to be indexed.
- The “Why pages aren’t indexed” table is pre-sorted by importance — work it top-down. Filter by Source = Website to find what you can actually fix (Google-source issues generally aren’t yours).
- 16 statuses, grouped: Google’s-choice not-indexed (discovered/crawled – currently not indexed); blocked-by-you (robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere., noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.); HTTP errors (401, 403, other 4xx, 5xx, 404); canonical/duplicate (alternate, duplicate w/o canonical, Google chose different canonical); redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. (page with redirectA Google Search Console Page Indexing status for a URL that redirects elsewhere. It's not indexed by design because it's a redirect — the destination is a separate question, and Google says the target may or may not end up indexed — usually expected, not an error (unlike the separate \"Redirect error\"). = normal vs. redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\" = broken); warnings (indexed though blocked by robots.txtA Google Search Console Page Indexing warning: Google indexed the URL anyway despite your robots.txt disallowing crawling it — Google names other pages linking to it as the likely path. robots.txt blocks crawling, not indexing., page indexed without content).
- Fix → Validate fix (~2 weeks typical). Speed trick: filter to a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. of your top pages and validate the subset.
- The report lags reality — Mueller has described it as a report that’s catching up over time. Don’t over-react to a stale number.
- Diagnostics: indexed drop + no errors = you blocked something; more not-indexed than indexed = robots block or parameter duplicates; error spike = template change or a bad sitemap.
- Small sites (<500 pages) probably don’t need it — a
site:search suffices. - Bing: no direct equivalent — Site Explorer for the aggregate view, URL Inspection for one page.
Official documentation
Primary-source documentation from the search engines.
- Page indexing report — Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Help: the buckets, the Source column, every status, the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. filter, and the Validate fix flow.
- Index Coverage Data Improvements (Jan 2021) — the data-quality update that reworked several statuses (removed “crawl anomaly,” added “indexed without contentA Google Search Console status meaning the URL is in Google's index, but Googlebot couldn't read any usable content from it — possible causes include a server/CDN block, cloaking, an empty render, or an unsupported format. Not the same as a page-level robots.txt block.,” changed “submitted but blocked” to “indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. but blocked”).
- URL Inspection Tool — the per-URL companion to this report.
Bing / Microsoft
- Bing Webmaster Tools — URL Inspection — Bing’s per-URL index/SEO/markup view.
- Bing Webmaster Tools — Site Explorer (Refreshed Webmaster Tools) — the closest aggregate analogue to GSC’s Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason..
Quotes from the source
On-the-record statements from Google. Each link is a deep link that jumps to the quoted passage on the source page.
Google — what the report is and shows
- “See which pages Google can find and indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. on your site, and learn about any indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. problems encountered.” — Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Help. Jump to quote
- “The Page indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. shows the Google indexing status of all URLs that Google knows about in your property.” Jump to quote
Google — report vs. single-page, and who can fix
- “This report isn’t used to investigate the index status of specific pages. To find the index status of a specific page, use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version..” Jump to quote
- “The Source value in the table shows whether the source of the issue is Google or the website. In general, you can fix only issues where the source is listed as “Website”.” Jump to quote
Google — “not indexed” isn’t necessarily bad
- “Not indexed is not necessarily bad.” Jump to quote
- “You should not expect all URLs on your site to be indexed, only the canonical pages.” Jump to quote
Google — who needs it, validation timing
- “If your site has fewer than 500 pages, you probably don’t need to use this report.” Jump to quote
- “Validation typically takes up to about two weeks, but in some cases can take much longer, so please be patient.” Jump to quote
Gary Illyes, Google — on “crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” and site quality (SERP Conf 2024, relayed by Search Engine Journal)
- “And the general quality of the of the site, that can matter a lot of how many of these crawled but not indexed you see in search consoleGoogle's free tool for monitoring crawling, indexing, and search performance.. If the number of these URLs is very high that could hint at general quality issues.” Read the coverage
John Mueller, Google — on report lag (relayed by Search Engine Journal)
- “It’s just a report that’s kind of lagging behind.” Read the coverage
#:~:text= fragments resolve in your browser. The Illyes and
Mueller lines are quoted through Search Engine Journal’s contemporaneous coverage
(the SEJ pages use curly quotes, so the deep-link fragments target shorter ASCII
substrings) and should be confirmed against the originals before being treated as
final. The Illyes “general quality” line is attributed to Gary Illyes, not
Mueller. Working the Page Indexing report — checklist
A repeatable pass for triaging the report instead of staring at it:
- Note the IndexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. vs. Not indexed totals and the trend — is anything moving?
- Open the “Why pages aren’t indexed” table and work it top-down (it’s pre-sorted by importance).
- Filter your attention to Source = Website — those are the issues you can actually fix.
- For each real issue, pull a sample URL and run it through URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. to confirm what’s happening on that page.
- Separate the intentional statuses (noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., robots-blocked, alternate w/ proper canonical, page with redirectA Google Search Console Page Indexing status for a URL that redirects elsewhere. It's not indexed by design because it's a redirect — the destination is a separate question, and Google says the target may or may not end up indexed — usually expected, not an error (unlike the separate \"Redirect error\").) from the broken ones (5xx, 404, redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\", unexpected 4xx) — only chase the broken ones.
- For canonical/duplicate statuses, verify the Google-chosen canonical in URL Inspection before “fixing” anything (see canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it.).
- If crawled / discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal. is large, treat it as a site-quality / crawl-demand signal, not a per-page bug.
- Cross-check a drop in indexed pages against a spike in a not-indexed status (you may have blocked something).
- Fix all instances, then click Validate fix — and don’t click it again until it resolves.
- To validate faster, filter to a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. of your top pages and validate that subset.
- Remember the report lags — give changes ~2 weeks before judging them.
The mental models
1. Report = aggregate; URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. = single page. The Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report tells you patterns across the site. URL Inspection tells you the truth about one URL (live status, chosen canonical, render). Always pair them: spot a pattern in the report, confirm the cause with Inspection.
2. Not indexed ≠ broken. The default reaction to a big not-indexed number should be “let me read the reasons,” not “let me fix everything.” Canonical/duplicate, noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., robots, and intentional 404s are correct. Google only expects your canonical pages to be indexed.
3. Source = Website is your to-do list. The Source column splits issues into “Google’s call” and “your call.” You can only move the website ones. Filtering to Source = Website (plus validation state failed/not started) turns a scary table into an actionable list.
4. Sort order is a priority list. Google pre-sorts the “Why pages aren’t indexed” table by importance. Don’t invent your own order — start at the top and go down.
5. Validate fix is a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. request, not a switch. Clicking it asks Google to re-check known affected URLs; it takes ~2 weeks and can fail. The trick to speed it up is scoping it to a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. of your most important pages.
6. The report is always a little behind. It reflects the last crawl per URL and updates on recrawl. Treat any single reading as a lagging snapshot, not live truth.
Page Indexing report — cheat sheet
The two reports people confuse
| Question | Use |
|---|---|
| How is my whole site indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and why aren’t pages indexed? | Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. |
| Is this one URL indexed? What canonical did Google pick? | URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. |
The 16 statuses, grouped
| Group | Statuses | Usual reaction |
|---|---|---|
| Google’s choice (not indexed) | Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal. · Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error. | Quality / crawl-demand signal; not a one-page fix |
| Blocked by you | Blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing. · URL marked ‘noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.’ | Confirm it’s intentional |
| HTTP errors | 401 · 403 · other 4xx · 5xx · 404 | Fix these |
| Canonical & duplicates | Alternate page w/ proper canonical · Duplicate without user-selected canonicalA Google Search Console Page Indexing status: Google found this page to be a duplicate, you didn't declare a canonical, so Google chose a different page as the canonical — and this URL isn't indexed. · Duplicate, Google chose different canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead. than user | Mostly correct; verify chosen canonical |
| RedirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. | Page with redirectA Google Search Console Page Indexing status for a URL that redirects elsewhere. It's not indexed by design because it's a redirect — the destination is a separate question, and Google says the target may or may not end up indexed — usually expected, not an error (unlike the separate \"Redirect error\"). (normal) · Redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\" (broken) | Fix only the redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\" |
| Warnings | Indexed, though blocked by robots.txtA Google Search Console Page Indexing warning: Google indexed the URL anyway despite your robots.txt disallowing crawling it — Google names other pages linking to it as the likely path. robots.txt blocks crawling, not indexing. · Page indexed without contentA Google Search Console status meaning the URL is in Google's index, but Googlebot couldn't read any usable content from it — possible causes include a server/CDN block, cloaking, an empty render, or an unsupported format. Not the same as a page-level robots.txt block. | Indexed but worth attention |
Fast facts
- Old name: Index Coverage reportThe Google Search Console report (renamed from \"Index Coverage\" in 2022) that shows which URLs Google has indexed, which it hasn't, and why. It splits your known URLs into Indexed and Not indexed, grouping the not-indexed ones by reason. (renamed to Page indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. in 2022).
- Source = Website = the issues you can fix.
- Validate fix ≈ 2 weeks; scope it to a small sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. to go faster.
- Example URL lists cap at 1,000; totals can drift slightly; the report lags.
- Indexed ≠ ranking — indexed only means eligible to appear.
- <500 pages? You probably don’t need this report — use a
site:search.
Common issues
Symptom → likely cause → fix, for the patterns that actually send people back to this report.
Indexed count dropped, no new errors showing
- Symptom: the “IndexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” total falls between visits, but the “Why pages aren’t indexed” table doesn’t show a matching spike in an error status.
- Likely cause: you blocked pages that were previously indexed — a new
robots.txtdisallow rule, anoindexshipped in a template, or a login/paywall requirement added to a section of the site. - Fix: check for a matching spike in a non-error, not-indexed status (blocked by
robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere., noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., or discovered/crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.) rather than the
HTTP-error statuses. Confirm on a sample URL with URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version., and check the
live robots directives with the robots-txt-tester (
/tools/robots-txt-tester).
”More not indexed than indexed”
- Symptom: the not-indexed total is larger than the indexed total, and it isn’t a brand-new site.
- Likely cause: either a
robots.txtrule blocking a large section, or a flood of near-duplicateThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. URLs from filter/sort parameters (?color=green,?sort=price) that Google is folding into “duplicate” statuses. - Fix: check the parameter pattern against your canonical tagsA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. with the
canonical-checker (
/tools/canonical-checker), then compare representative page templates in your crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.. This is a faceted-navigation/canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. problem wearing an indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. costume, not a report bug.
A sudden spike in one HTTP-error status
- Symptom: 401, 403, 404, or 5xx counts jump across many URLs at once, usually after a deploy.
- Likely cause: a template change that broke a shared component (auth check, redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. rule, error page), or a newly submitted sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. full of URLs that are blocked, noindexed, or already gone.
- Fix: pull a sample of affected URLs from the status detail and re-check their
live status codes with the http-status-checker (
/tools/http-status-checker). If it’s sitemap-sourced, audit the sitemap itself with the sitemap-validator (/tools/sitemap-validator) before resubmitting.
”Redirect error” (not the normal “Page with redirect”)
- Symptom: the report shows Redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\", distinct from the expected Page with redirectA Google Search Console Page Indexing status for a URL that redirects elsewhere. It's not indexed by design because it's a redirect — the destination is a separate question, and Google says the target may or may not end up indexed — usually expected, not an error (unlike the separate \"Redirect error\"). status.
- Likely cause: a redirect chainA → B → C instead of A → C. Each hop loses link equity and adds latency. that’s too long, a redirect loopA redirect loop is a chain of redirects that circles back on itself instead of ever reaching a live page — URL A redirects to B and B redirects back to A (or a longer cycle). No page ever returns a 200, so browsers show ERR_TOO_MANY_REDIRECTS and crawlers can't index anything., an empty or malformed target URL, or a target URL that exceeds Google’s max URL length.
- Fix: trace the actual hop sequence with the redirect-chain-mapper
(
/tools/redirect-chain-mapper) and shorten it to a single hop to the final URL.
”Validate fix” comes back Failed
- Symptom: you fixed the issue, clicked Validate fix, and it resolves to Failed instead of Passed.
- Likely cause: either the fix didn’t actually reach every affected URL (a CDN cache, a staging-only deploy, or a template fix that missed some URL patterns), or Google recrawled before the fix propagated.
- Fix: spot-check a handful of the originally affected URLs directly (view source or an HTTP-status/robots check, not just the browser render) to confirm the fix is live everywhere, then re-request validation — you don’t need to wait for the full cycle to restart if the underlying issue is confirmed fixed.
Validation tests
Proof that a fix you made actually changed the report’s underlying reality — not just that the report says so.
Confirm a blocked page’s live status matches expectations
- Test to run: http-status-checker (
/tools/http-status-checker) orcurl -Iagainst the affected URL. - Expected result: the status code you intended —
200if you meant to unblock it, or a deliberate301/404if you meant to remove it. - Failure interpretation: if it still returns the old error (401/403/5xx), the fix hasn’t shipped everywhere (check CDN cache, load balancer, or environment).
- Monitoring window: immediate — HTTP status is live, not delayed.
- Rollback trigger: status doesn’t match intent after a cache purge and a few minutes’ wait — revert the change and re-diagnose before touching production again.
Confirm a robots.txt fix actually unblocks the URL
- Test to run: robots-txt-tester (
/tools/robots-txt-tester) against the specific URL path. - Expected result: the tester reports the URL as allowed for GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer..
- Failure interpretation: still blocked means the rule wasn’t removed, a more specific rule is still matching, or the wrong robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. (staging vs. production) was checked.
- Monitoring window: immediate for the file itself; allow a normal crawl cycle (days, not hours) before expecting the GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. status to update.
- Rollback trigger: the tester still shows the URL blocked after confirming you
edited the live, production
robots.txt— treat as unresolved and keep debugging.
Confirm a canonical fix resolved the way you intended
- Test to run: canonical-checker (
/tools/canonical-checker) on the affected URL, followed by URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. in GSC for the Google-selected canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead.. - Expected result: the
<link rel="canonical">tag matches the URL you want indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and URL Inspection shows Google agreeing with your declared canonical. - Failure interpretation: if Google still shows a different “Google-selected canonical” than your declared one, your fix addressed the tag but not the underlying duplication signals (internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. entries, redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. all still need to agree).
- Monitoring window: 2–4 weeks — canonical selection is one of the slower signals to re-settle.
- Rollback trigger: Google’s selected canonical still disagrees with yours after 4+ weeks and a Validate fix cycle — treat the underlying signals (not just the tag) as still conflicting.
Confirm a redirect fix removed the chain/loop
- Test to run: redirect-chain-mapper (
/tools/redirect-chain-mapper) on the affected URL. - Expected result: a single hop straight to the final destination URL with a
200. - Failure interpretation: more than one hop, or a loop back to an earlier URL, means the redirect rule wasn’t consolidated — you’re likely chaining a new rule on top of an old one instead of replacing it.
- Monitoring window: immediate for the chain itself; a normal recrawl cycle before the GSC status updates.
- Rollback trigger: the chain is still longer than one hop after the fix — revert and consolidate the redirect rules at the source.
Confirm the “Validate fix” request actually passed
- Test to run: the Validate fix button on the specific status in GSC’s Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report (scope to a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. of your priority URLs to get a faster read).
- Expected result: the validation state moves to Passed (or N/A, meaning Google confirmed the fix independently).
- Failure interpretation: Failed means Google rechecked and still found the issue on at least some URLs — go back to the underlying fix, don’t just re-click Validate.
- Monitoring window: “typically up to about two weeks,” per Google — can run longer.
- Rollback trigger: two consecutive Failed results on the same status after confirming the underlying fix is live — escalate to checking whether the fix actually reaches every affected URL pattern, not just the sample you tested.
How to measure
The standing KPIs for keeping an eye on this report over time, not one-off checks.
Indexed page count (trend, not a snapshot)
- What it tells you: whether Google’s indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. footprint of your site is growing, flat, or shrinking.
- How to pull it: the “Indexed” total at the top of the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason., or the historical chart under “View data about indexed pages.”
- Benchmark / realistic range: no universal number — depends entirely on how many URLs you want indexed (your canonical page count), which varies by site. Track the trend against your own canonical URLHow search engines pick one canonical URL among duplicates and consolidate signals onto it. count rather than chasing 100% coverage; Google explicitly says only canonical pages should be indexed.
- Cadence: monthly, or immediately after a large content or URL-structure change.
”Why pages aren’t indexed” total, by Source = Website
- What it tells you: the size of your actual, actionable to-do list — issues Google says you (not Google) can fix.
- How to pull it: filter the “Why pages aren’t indexed” table by Source = Website and sum the affected-URL counts across statuses with a “Failed” or “Not started” validation state.
- Benchmark / realistic range: depends on site size and history — there’s no honest universal target. Establish your own baseline on your next full pass through the table, then track whether that number goes down over successive checks.
- Cadence: monthly for most sites; weekly during an active cleanup or after a large migration.
HTTP-error status count (401/403/404/5xx)
- What it tells you: whether pages you expect to serve are actually broken for GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. — the group of statuses that’s almost always worth fixing.
- How to pull it: the individual status rows in the report, cross-checked with
the http-status-checker (
/tools/http-status-checker) on a sample of affected URLs. - Benchmark / realistic range: zero is the honest target for URLs you intend to keep live — any nonzero count on canonical pages is worth investigating; a nonzero count on already-removed pages (expected 404s) is normal and depends on your content lifecycle.
- Cadence: check after every deploy that touches routing, auth, or templates; otherwise monthly.
Validate-fix pass rate
- What it tells you: whether your fixes are actually landing across all affected URLs, or only partially.
- How to pull it: track each Validate fix request’s outcome (Passed / Failed / N/A) in the status detail view over time.
- Benchmark / realistic range: no fixed target — a healthy pattern is most requests resolving to Passed or N/A on the first attempt. Repeated Failed results on the same status is the signal to watch, not a specific percentage.
- Cadence: per fix cycle — check back roughly two weeks after each Validate fix request.
Patrick's relevant free tools
- SEO Opportunity Finder — Choose an evidence-backed content, link, or technical SEO workflow, then open the free tool that can evaluate it without turning every warning into an opportunity.
- Google Index Checker — Check one URL’s observable indexability blockers, or reconcile sitemap, crawl, and supplied Search Console evidence across a URL set before verifying Google’s actual state in URL Inspection.
- Canonicalization Checker — Audit HTML and HTTP canonical signals, test the canonical target, and identify observable conflicts that can cause Google to choose a different URL.
Tools
The Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report tells you what Google thinks is wrong; these tools help you confirm and fix the underlying cause.
This site’s tools
- gscA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.-workbench (
/tools/gsc-workbench) — work with your own GSC data (via the Search Console APIA set of REST APIs that let you programmatically read and manage Google Search Console data for properties you've verified — Performance data, URL index status, sitemaps, and properties — authorized with OAuth 2.0.) alongside the concepts on this page. - gsc-regex-tester (
/tools/gsc-regex-tester) — build and test the regex filters GSC’s Performance and Page Indexing reportsThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. accept, useful for isolating a URL pattern behind a status. - http-status-checker (
/tools/http-status-checker) — confirm the live status code (401/403/404/5xx) behind an HTTP-error status before or after a fix. - redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't.-checker (
/tools/redirect-checker) and redirect-chain-mapper (/tools/redirect-chain-mapper) — check a single redirect or trace a full chain when you see Redirect errorA Google Search Console Page Indexing status meaning Googlebot couldn't follow a redirect — a chain too long, a loop, a URL over the max length, or a bad/empty URL in the chain. It's a broken redirect to fix, not the normal \"Page with redirect.\". - canonical-checker (
/tools/canonical-checker) — check the declared canonical tag on a URL flagged in the canonical/duplicate group. - robots-txt-tester (
/tools/robots-txt-tester) — verify whether a specific URL is actually blocked before treating Blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing. as intentional or a bug. - sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.-validator (
/tools/sitemap-validator) and xml-sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.-generator (/tools/xml-sitemap-generator) — audit or rebuild the sitemap you’ll use to scope a faster Validate-fix request. - site-audit-lite (
/tools/site-audit-lite) — a broader crawl-based sanity check when a Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. pattern (large error spike, mass duplication) looks site-wide rather than isolated to a few URLs.
Third-party tools
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. — the Page Indexing report itself, plus URL Inspection for the single-URL companion view named throughout this article.
- Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. — Site Explorer (aggregate view) and URL Inspection (single-URL view), Bing’s closest equivalents.
- Screaming Frog — crawl your own site the way GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. would to cross-check status codes, redirects, and canonical tagsA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. at scale before trusting a GSC sample.
Resources worth your time
My related writing
- How to Fix “Discovered - currently not indexed” — the deep dive on one of the most-misread statuses in this report.
- How to Remove URLs From Google Search (5 Methods) — ties to the noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. / removal-tool side of the report.
- Anonymized Queries Make Up Nearly Half of GSC Traffic — my GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. data study; a good reminder that Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. is powerful but incomplete.
- The Beginner’s Guide to Technical SEO — where indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. and GSC fit in the bigger picture.
My speaking
- How Search Works (SlideShare) — my walkthrough of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and ranking, which is the context the whole Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. sits inside. (My standing disclaimer applies: “This is my understanding of systems… not going to be 100% complete or accurate.”)
Official
From around the industry
- r/TechSEO — the community for indexing and coverage-report debugging.
- Google Explains Reasons For Crawled Not Indexed (Search Engine Journal) — covers Gary Illyes at SERP Conf 2024 explaining why high “crawled not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” numbers can hint at site-wide quality issues.
- Mueller Asked About Lag in Google Search Console Indexing Report (Search Engine Journal) — John Mueller on why the report lags and shouldn’t trigger an over-reaction.
- Google Search Console Coverage Report Renamed to Pages Report (Search Engine Roundtable) — Barry Schwartz’s first sighting of the Index Coverage → Page indexing rename at Google I/O 2022.
- Index Coverage Data Improvements (Google Search Central Blog) — the Jan 2021 update that removed “crawl anomaly,” added “indexed without contentA Google Search Console status meaning the URL is in Google's index, but Googlebot couldn't read any usable content from it — possible causes include a server/CDN block, cloaking, an empty render, or an unsupported format. Not the same as a page-level robots.txt block.,” and reworked several statuses — explains why older screenshots don’t match today’s report.
- Google Search Console Help — Page indexing report — the primary official reference: buckets, the Source column, all statuses, the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. filter, and the Validate fix flow.
Quiz
Test what you remember about reading the Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report.
Page Indexing report
The Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.
Parent concept: Indexing · Related: Crawling, Indexing, Canonicalization
Page Indexing report
The Page IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. report — formerly the Index Coverage reportThe Google Search Console report (renamed from \"Index Coverage\" in 2022) that shows which URLs Google has indexed, which it hasn't, and why. It splits your known URLs into Indexed and Not indexed, grouping the not-indexed ones by reason., and labeled “Pages” under “IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” in the Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. left nav — shows which URLs Google can find and index on a site. It groups them into an Indexed bucket (eligible to appear in Search) and a Not indexed bucket, with a “Why pages aren’t indexed” table that breaks the not-indexed URLs down by reason so you can find and fix indexing problems.
It is an aggregate, property-wide report — not a tool for checking a single page. For one URL’s live status and Google-chosen canonical, you use the URL Inspection toolA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. instead.
A key thing to internalize: “not indexed” is not necessarily bad. Many statuses (canonical/duplicate, noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., robots-blocked, intentional 404s) are correct outcomes, and Google says you should not expect every URL on your site to be indexed — only the canonical pages. If your site has fewer than 500 pages, Google says you probably don’t need this report at all.
Parent concept: Indexing · Related: Crawling, Indexing, Canonicalization
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 26, 2026.
Editorial summary and recorded change details.Summary
Removed the retired duplicate-content checker from Page Indexing diagnostics.
Change details
-
Duplicate-pattern investigation now uses canonical checks and representative crawler comparisons.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Added a compact triage order for working the report and clarified what the URL Inspection live test does and doesn't cover for canonical/duplicate statuses.
Change details
-
Added a ten-minute triage order (sitemap scope → unexpected change → Source = Website → business-critical templates → validate the pattern) to the 'How to read the report' section.
-
Clarified in 'A note on lag' that URL Inspection's live test confirms current crawlability but doesn't determine which canonical Google selects among duplicates — that's a separate indexing decision.