Indexing

How search engines store and organize pages so they can rank — content analysis, canonicalization, why crawled isn't indexed, and reading the GSC Page indexing report.

First published: Jun 23, 2026 · Last updated: Jul 18, 2026 · Advanced
demand #8 in How Search Works#29 in Technical SEO#37 on the site
1 evidence signal on this page

Indexing is stage two of search (crawl → index → serve): after a page is crawled, the engine understands it, deduplicates and canonicalizes it, and — if it qualifies — stores it in the search index. Crawled isn't indexed; Google selects what to keep, and indexing isn't guaranteed. It's not a ranking factor, but a page must be indexed before it can rank. To keep a page out, use noindex and leave it crawlable — don't block it in robots.txt. This hub explains the whole stage and routes you to the deep dives.

TL;DR — Indexing is the second of search’s three stages (crawl → index → serve): Google understands a crawled page (text, key tags, images, video; it renders JS), detects duplicates, clusters similar pages and picks the most representative one (canonicalization — rel=canonical is a hint, not a rule), computes signals, and stores the canonical in the index. Crawled ≠ indexed — “indexing isn’t guaranteed,” and the call is largely about quality/value. The Search Console Page indexing report is your cockpit. To deindex, use noindex and keep the page crawlable; never use robots.txt to remove a page, because a blocked page can still be indexed (just without a snippet).

Evidence for this claim Google must be able to crawl a page to see and apply its noindex rule. Scope: Google-supported robots meta and X-Robots-Tag directives; a robots.txt block can prevent Google from seeing the rule. Confidence: high · Verified: Google Search Central: Block Search indexing with noindex

Indexing is stage two of three

Indexing is stage two of three — the middle filter between crawling and ranking. Source: /technical-seo/how-search-works/indexing/

Three stages run left to right. Crawl discovers and downloads a URL. Index processes the page and stores eligible information. Serve or rank orders the best indexed matches for a query. The Index stage is highlighted, and a note says not every page advances through every stage.

© Patrick Stox LLC · CC BY 4.0 ·

Google is blunt about the pipeline: “Google Search works in three stages, and not all pages make it through each stage” — crawling, indexing, and serving. Indexing is the middle stage, and the doc defines it cleanly: “Indexing: Google analyzes the text, images, and video files on the page, and stores the information in the Google index, which is a large database.”

Evidence for this claim Google Search describes crawling, indexing, and serving as three distinct stages; indexing analyzes page content and stores eligible information in Google's index. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search works

A page has to be crawled before it can be indexed, and it has to be indexed before it can rank. But none of those are guarantees — each stage is a filter. Keeping the three stages separate in your head is the single most useful mental model in technical SEO, and it’s why I always ask which stage a page is failing at before changing anything. (For the stage before this one, see the crawling hub — crawl → index is the pipeline.)

What actually happens during indexing

Indexing is a sequence: understand, cluster and select a canonical, then store. Source: /technical-seo/how-search-works/indexing/

Step one analyzes a crawled page for text, title, alt text, images, and video. Step two groups duplicate URLs into a cluster and chooses the most representative page as canonical. Step three stores the canonical page and its cluster information in the Google index.

© Patrick Stox LLC · CC BY 4.0 ·

Indexing isn’t one thing; it’s a sequence:

  • Understanding the content. Google: “After a page is crawled, Google tries to understand what the page is about. This stage is called indexing.” That means “Google analyzes the textual content and key content tags and attributes, such as <title> elements and alt attributes, images, videos, and more.” JavaScript is rendered as part of this — if your content only appears after JS runs, it still has to render before it can be understood.
  • Duplicate detection & canonicalization. This is the part most explainers skip, and it’s where a lot of “why isn’t this indexed?” mysteries live. Google “determines if a page is a duplicate of another page on the internet or canonical.” The mechanic: “we first group together (also known as clustering) the pages that we found on the internet that have similar content, and then we select the one that’s most representative of the group.” That representative is the canonical — “The canonical is the page that may be shown in search results.”
  • Computing signals & storing. Finally, “The collected information about the canonical page and its cluster may be stored in the Google index, a large database hosted on thousands of computers.” Google’s named indexing system behind all this is Caffeine — the layer that ingests crawl data, renders and extracts, computes signals, and builds the index that gets served.

Canonicalization: a hint, not a command

Because canonicalization happens during indexing, it deserves its own note. “Canonicalization is the process of selecting the representative –canonical– URL of a piece of content,” and it exists because “this process helps Google show only one version of the otherwise duplicate content in its search results.”

The load-bearing detail: your rel=canonical is a suggestion. Google’s words: “indicating a canonical preference is a hint, not a rule.” Google weighs many signals — in my canonicalization guide I note that, per Google’s Allan Scott, there are roughly 40 different canonical selection signals — and it can pick a different URL than the one you flagged. That’s exactly what the “Duplicate, Google chose different canonical than user” status in Search Console is telling you.

Crawled ≠ indexed: why pages don’t get indexed

Here’s the myth-buster, straight from the docs: “Indexing isn’t guaranteed; not every page that Google processes will be indexed.” Evidence for this claim Google does not guarantee that every processed page will be indexed. Scope: Google Search indexing; the source gives examples of possible causes rather than an exhaustive decision formula. Confidence: high · Verified: Google Search Central: In-depth guide to how Google Search works Google lists common reasons it fails — “The quality of the content on page is low,” “Robots meta rules disallow indexing,” and “The design of the website might make indexing difficult.”

The reps are even more direct that this is a selection decision driven by value, not a quota you can buy past:

  • John Mueller, on how long “Discovered/Crawled – currently not indexed” can persist: “That can be forever. It’s something where we just don’t crawl and index all pages.” The fix isn’t resubmitting — it’s making the systems recognize the value, to “continue working on the website and making sure that our systems recognize that there’s value in crawling and indexing more and then over time we will crawl and index more.”
  • Mueller again: “it’s important to keep in mind that Google just doesn’t index every page on the web, even if it’s submitted directly.” And, bluntly: “Well, lots of SEOs & sites (perhaps not you/yours!) produce terrible content that’s not worth indexing.”
  • Gary Illyes, on why it’s selective: “we don’t have infinite space, so we want to index stuff that we think– well, not we– but our algorithms determine that it might be searched for…”
  • Martin Splitt frames it as a balancing act: “I usually describe it as a challenge with the balance between not overwhelming the website and also spending our resources where it matters.”

The practical takeaway: a sitemap or “request indexing” aids discovery, not selection. Submitting a page again won’t force it in. The lever is site quality and value.

Reading the Google Search Console Page indexing report

The Page indexing report is where indexing problems actually show up. Treat each status as a diagnosis. These are Google’s own verbatim descriptions:

  • Crawled – currently not indexed: “The page was crawled by Google but not indexed. It may or may not be indexed in the future; no need to resubmit this URL for crawling.” Usually a quality/value judgment — improve the page, don’t spam the resubmit button.
  • Discovered – currently not indexed: “The page was found by Google, but not crawled yet. Typically, Google wanted to crawl the URL but this was expected to overload the site; therefore Google rescheduled the crawl.” Technically a pre-crawl, capacity-driven status — but if it persists, reps tie that to value, same as the one above.
  • Duplicate without user-selected canonical: “This page is a duplicate of another page, although it doesn’t indicate a preferred canonical page. Google has chosen the other page as the canonical for this page, and so will not serve this page in Search.”
  • Duplicate, Google chose different canonical than user: “This page is marked as canonical for a set of pages, but Google thinks another URL makes a better canonical.” (The “hint, not a rule” outcome in the wild.)
  • Alternate page with proper canonical tag: “This page is marked as an alternate of another page… This page correctly points to the canonical page, which is indexed, so there is nothing you need to do.”
  • Indexed, though blocked by robots.txt: “The page was indexed despite being blocked by your website’s robots.txt file. Google always respects robots.txt, but this doesn’t necessarily prevent indexing if someone else links to your page.” This is the proof that blocking crawling does not block indexing.
  • URL blocked by robots.txt: “This page was blocked by your site’s robots.txt file.”
  • URL marked ‘noindex’: “When Google tried to index the page it encountered a ‘noindex’ directive and therefore did not index it.” (This is the deindex working as intended.)
  • Page with redirect: “This is a non-canonical URL that redirects to another page. As such, this URL will not be indexed.”
  • Soft 404: “The page request returns what we think is a soft 404 response. This means that it returns a user-friendly ‘not found’ message but not a 404 HTTP response code.”

How to control indexing the right way

To get a page indexed: make it crawlable, link to it internally, include it in your sitemap — and, above all, make it worth indexing. Discovery aids don’t override the value judgment.

To keep a page OUT — the most-botched control in SEO: use noindex, “a rule set with either a <meta> tag or HTTP response header,” and keep the page crawlable. Google’s load-bearing warning: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”

Evidence for this claim For Google to apply noindex, the crawler must be allowed to access the page or resource; a robots.txt block can prevent Google from seeing the rule. Scope: HTML and HTTP resources Confidence: high · Verified: Block Search indexing with noindex

The mistake I see constantly is adding noindex and blocking the page in robots.txt. That’s counterproductive. As I put it in How to Remove URLs From Google Search: “For these tags to be seen, a search engine needs to be able to crawl the pages—so make sure they aren’t blocked in robots.txt,” and “Crawling is not the same thing as indexing. Even if Google is blocked from crawling pages, if there are any internal or external links to a page they can still index it.” In my piece on the Indexed, though blocked by robots.txt status I say it even more plainly: “Unless Google can crawl a page, they won’t see the noindex meta tag and may still index it because it has links.” The fix: “Just add a noindex meta robots tag and make sure to allow crawling—assuming it’s canonical.”

For the fuller removal decision tree — 404/410 vs. noindex, the Removals tool’s ~6-month hold, and password protection — see how to deindex a page.

Indexing in Bing

Bing runs the same pipeline. As Microsoft describes it: “As Bingbot crawls the web, it sends information to Bing about what it finds. These pages are then added to the Bing index.” The same controls apply — a noindex directive keeps a page out, and an over-restrictive robots.txt can stop Bingbot from ever crawling it. Bing also needs at least one link pointing to your site to find it in the first place.

Not every page belongs in the index

A simple filter, not a universal rule: a page is worth indexing if it can show up for a search with a distinct, useful result. That’s the bar to check duplicates, parameter variants, private or staging URLs, and thin or repetitive inventory pages against — not a reason to noindex a whole page type by default. Run the full audit on the pages built to own it, next.

Where to go next: the indexing cluster

This hub is the overview. Two things go wrong at scale, and each gets its own deep dive:

  • Index bloat — when too many low-value, duplicate, or thin URLs end up in the index, diluting your site and wasting crawl/index resources. How to diagnose it and prune it safely.
  • Mobile-first indexing — Google indexes the mobile version of your pages, so content, links, and structured data have to reach parity between mobile and desktop. What to check and what breaks.

Both topics are nested under this hub — they’re in the sidebar too, and they’ll link back here.

The stage before this one — how bots discover and download your pages — lives in the crawling hub; crawl → index is the pipeline, and a page has to clear crawling before any of this applies. For the broader picture, see How Search Works.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin an expert quote first.