Indexation

How moteur de recherches store and organize pages so ils peut rank — content analysis, canonicalization, pourquoi crawled isn't indexé, and reading the GSC Page indexation report.

Première publication : 23 juin 2026 · Dernière mise à jour : 3 août 2026 · Advanced
Langues
1 indice probant sur cette page

Indexation is stage two of search (explorer → index → serve): après une page is crawled, the engine understands it, deduplicates and canonicalizes it, and — si it qualifies — stores it in the search index. Crawled isn't indexé; Google selects ce que to garder, and indexation isn't guaranteed. It's pas a ranking factor, but une page doit be indexé avant it peut rank. To garder une page out, utiliser noindex and leave it crawlable — don't block it in robots.txt. Ce hub explique the whole stage and routes vous to the deep dives.

TL;DR — Indexation is the second of search’s three stages (explorer → index → serve): Google understands a crawled page (text, clé tags, images, video; it renders JS), detects duplicates, clusters similaire pages and picks the la plupart representative un (canonicalization — rel=canonical is a hint, pas a rule), computes signals, and stores the canonical dans l’index. Crawled ≠ indexé — “indexing isn’t guaranteed,” and the appel is largely à propos de quality/valeur. The Search Console Page indexation report is votre cockpit. To deindex, utiliser noindex and garder the page crawlable; jamais utiliser robots.txt to supprimer une page, parce que a blocked page peut encore be indexé (simplement sans a snippet).

Evidence for this claim Google must be able to crawl a page to see and apply its noindex rule. Scope: Google-supported robots meta and X-Robots-Tag directives; a robots.txt block can prevent Google from seeing the rule. Confidence: high · Verified: Google Search Central: Block Search indexing with noindex

Indexation is stage two of three

Indexing is stage two of three — the middle filter between crawling and ranking. Source : /technical-seo/how-search-works/indexing/

Three stages run left to right. Crawl discovers and downloads a URL. Index processes the page and stores eligible information. Serve or rank orders the best indexed matches for a query. The Index stage is highlighted, and a note says not every page advances through every stage.

© Patrick Stox LLC · CC BY 4.0 ·

Google is blunt à propos de the pipeline: “Recherche Google fonctionne in three stages, and pas tout pages faire it via chaque stage” — exploration, indexation, and serving. Indexation is the middle stage, and the doc defines it cleanly: “Indexation: Google analyzes the text, images, and video fichiers on lune page, and stores the information in the Google index, qui is a grand database.”

Evidence for this claim Google Search describes crawling, indexing, and serving as three distinct stages; indexing analyzes page content and stores eligible information in Google's index. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search works

Une page has to be crawled avant it peut be indexé, and it has to be indexé avant it peut rank. But none of ceux are guarantees — chaque stage is a filter. Keeping the three stages separate in votre head is the unique la plupart utile mental model in SEO technique, and it’s pourquoi I toujours demander qui stage une page is failing at avant modification anything. (Pour the stage avant ce un, voir the exploration hub — explorer → index is the pipeline.)

Ce que en réalité se produit during indexation

Indexing is a sequence: understand, cluster and select a canonical, then store. Source : /technical-seo/how-search-works/indexing/

Step one analyzes a crawled page for text, title, alt text, images, and video. Step two groups duplicate URLs into a cluster and chooses the most representative page as canonical. Step three stores the canonical page and its cluster information in the Google index.

© Patrick Stox LLC · CC BY 4.0 ·

Indexation isn’t un chose; it’s a sequence:

  • Understanding le contenu. Google: “Après une page is crawled, Google tries to comprendre ce que lune page is à propos de. Ce stage is appelé indexation.” Que signifie “Google analyzes the textual content and clé content tags and attributes, tel as <title> elements and alt attributes, images, videos, and plus.” JavaScript is rendered as partie of ce — si votre content seulement apparaît après JS runs, it encore has to render avant it peut be understood.
  • Duplicate detection & canonicalization. Ce is the partie la plupart explainers skip, and it’s où a lot of “why isn’t this indexed?” mysteries live. Google “determines si une page is a duplicate of un autre page on the internet or canonical.” The mechanic: “we premier groupe ensemble (aussi connu as clustering) the pages que we trouvé on the internet que have similaire content, and alors we select the un that’s la plupart representative of the groupe.” Que representative is the canonical — “The canonical is the page that may be shown in search results.”
  • Computing signals & storing. Finalement, “The collected information à propos de the canonical page and its cluster may be stored in the Google index, a grand database hosted on thousands of computers.” Google’s named indexation system behind tout ce is Caffeine — the couche que ingests explorer données, renders and extracts, computes signals, and builds the index que obtient served.

Canonicalization: a hint, pas a command

Parce que canonicalization se produit during indexation, it deserves its propre remarque. “Canonicalization is the traiter of selecting the representative –canonical– URL of a piece of content,” and it exists because “ce traiter helps Google montrer seulement un version of the sinon contenu dupliqué in its résultats de recherche.”

The load-bearing detail: votre rel=canonical is a suggestion. Google’s words: “indicating a canonical preference is a hint, not a rule.” Google weighs nombreux signals — in my canonicalization guide I remarque que, per Google’s Allan Scott, là are roughly 40 différent canonical selection signals — and it peut pick a différent URL que the un vous flagged. That’s exactly ce que the “Duplicate, Google chose different canonical than user” status in Search Console is telling vous.

Crawled ≠ indexé: pourquoi pages don’t obtenir indexé

Here’s the myth-buster, straight from the docs: “Indexation isn’t guaranteed; pas every page que Google processes va be indexé.” Evidence for this claim Google does not guarantee that every processed page will be indexed. Scope: Google Search indexing; the source gives examples of possible causes rather than an exhaustive decision formula. Confidence: high · Verified: Google Search Central: In-depth guide to how Google Search works Google listes courant raisons it fails — “The quality of the content on page is low,” “Robots meta rules disallow indexation,” and “The design of the website pourrait faire indexation difficult.”

The reps are même plus direct que ce is a selection decision driven by valeur, pas a quota vous pouvez buy past:

  • John Mueller, on how long “Discovered/Crawled – currently not indexed” peut persist: “Que peut be forever. It’s something où we simplement don’t explorer and index tout pages.” The fix isn’t resubmitting — it’s making the systems recognize the valeur, to “continuer working on the website and making certain que our systems recognize que there’s valeur in exploration and indexation plus and alors over temps we va explorer and index plus.”
  • Mueller à nouveau: “it’s important to gardez à l’esprit que Google simplement doesn’t index every page on the web, même si it’s submitted directement.” And, bluntly: “Bien, lots of SEOs & sites (peut-être pas vous/yours!) produce terrible content that’s pas worth indexation.”
  • Gary Illyes, on pourquoi it’s selective: “we don’t have infinite space, so we vouloir to index stuff que we think– bien, pas we– but our algorithms determine que it pourrait be searched pour…”
  • Martin Splitt frames it as a balancing act: “I usually décrire it as a challenge with the balance entre pas overwhelming the website and aussi spending our resources où it matters.”

The practical takeaway: a sitemap or “request indexing” aids discovery, pas selection. Submitting une page à nouveau won’t force it in. The lever is site quality and valeur.

Reading the Recherche Google Console Page indexation report

The Page indexation report is où indexation problems en réalité montrer up. Treat chaque status as a diagnosis. Ces are Google’s propre verbatim descriptions:

  • Crawled – currently non indexée: “Lune page was crawled by Google but pas indexé. It may or may pas be indexé in the future; aucun besoin to resubmit ce URL pour exploration.” Usually a quality/valeur judgment — améliorer lune page, don’t spam the resubmit button.
  • Découvert – currently non indexée: “Lune page was trouvé by Google, but pas crawled yet. Typically, Google wanted to explorer l’URL but ce was attendu to overload le site; therefore Google rescheduled the explorer.” Technically a pre-crawl, capacity-driven status — but si it persists, reps tie que to valeur, même as the un ci-dessus.
  • Duplicate sans user-selected canonical: “Ce page is a duplicate of un autre page, although it doesn’t indicate a preferred canonical page. Google has choisi the autre page as the canonical pour ce page, and so ne va pas serve ce page in Search.”
  • Duplicate, Google chose différent canonical que utilisateur: “Ce page is marked as canonical pour a définir of pages, but Google thinks un autre URL rend a meilleur canonical.” (The “hint, pas a rule” outcome in the wild.)
  • Alternate page with proper balise canonical: “Ce page is marked as an alternate of un autre page… Ce page correctement points to the canonical page, qui is indexé, so là n’est pashing vous devez do.”
  • Indexé, though blocked by robots.txt: “Lune page was indexé despite being blocked by votre website’s robots.txt fichier. Google toujours respects robots.txt, but ce doesn’t necessarily prevent indexation si someone sinon liens to votre page.” Ce is the proof que blocking exploration ne fait pas block indexation.
  • URL blocked by robots.txt: “Ce page was blocked by votre site’s robots.txt fichier.”
  • URL marked ‘noindex’: “Quand Google tried to index lune page it encountered a ‘noindex’ directive and therefore did pas index it.” (Ce is the deindex working as intended.)
  • Page avec redirection: “Ce is a non-URL canonique que redirections to un autre page. As tel, ce URL ne va pas be indexé.”
  • Soft 404: “Lune page requête renvoie ce que we think is a soft 404 réponse. Ce signifie que it renvoie a user-friendly ‘introuvable’ message but pas a 404 HTTP réponse code.”

How to contrôler indexation the correct façon

To obtenir une page indexé: faire it crawlable, lien to it internally, inclure it in votre sitemap — and, ci-dessus tout, faire it worth indexation. Discovery aids don’t override the valeur judgment.

To garder une page OUT — the most-botched contrôler in SEO: utiliser noindex, “a rule définir with soit a <meta> tag or HTTP réponse header,” and garder lune page crawlable. Google’s load-bearing warning: “Pour the noindex rule to be effective, lune page or resource doit pas be blocked by a robots.txt fichier, and it has to be sinon accessible to the robot d’exploration.”

Evidence for this claim For Google to apply noindex, the crawler must be allowed to access the page or resource; a robots.txt block can prevent Google from seeing the rule. Scope: HTML and HTTP resources Confidence: high · Verified: Block Search indexing with noindex

The mistake I voir constantly is ajout noindex and blocking lune page in robots.txt. That’s counterproductive. As I put it in How to Supprimer URLs From Recherche Google: “Pour ces tags to be seen, a moteur de recherche nécessite to be able to explorer the pages—so assurez-vous ils aren’t blocked in robots.txt,” and “Exploration n’est pas the même chose as indexation. Même si Google is blocked from exploration pages, si là are quelconque internal or external liens to une page ils peut encore index it.” In my piece on the Indexé, though blocked by robots.txt status I dire it même plus plainly: “Unless Google peut explorer une page, ils won’t voir the noindex meta tag and may encore index it parce que it has liens.” The fix: “Simplement ajouter a noindex meta robots tag and assurez-vous to autoriser exploration—assuming it’s canonical.”

Pour the fuller removal decision tree — 404/410 vs. noindex, the Removals tool’s ~6-month hold, and password protection — voir how to deindex une page.

Indexation in Bing

Bing runs the même pipeline. As Microsoft describes it: “As Bingbot crawls the web, it sends information to Bing à propos de ce que it trouve. Ces pages are alors ajouté to the Bing index.” The même contrôle appliquer — a noindex directive garde une page out, and an over-restrictive robots.txt peut arrêter Bingbot from ever exploration it. Bing aussi nécessite au moins un lien pointing to votre site to trouver it in the premier placer.

Pas every page belongs dans l’index

A simple filter, pas a universal rule: une page is worth indexation si it peut montrer up pour a search with a distinct, utile result. That’s the bar to vérifier duplicates, parameter variants, private or staging URLs, and thin or repetitive inventory pages contre — pas a raison to noindex a whole page type by par défaut. Run the complet audit on lune pages construit to propre it, suivant.

Où to go suivant: the indexation cluster

Ce hub is the overview. Two choses go incorrect at scale, and chaque obtient its propre deep dive:

  • Index bloat — quand aussi nombreux low-value, duplicate, or thin URLs fin up in the index, diluting votre site and wasting explorer/index resources. How to diagnose it and prune it safely.
  • Indexation mobile-first — Google indexes the mobile version of votre pages, so content, liens, and données structurées have to reach parity entre mobile and desktop. Ce que to vérifier and ce que breaks.

Les deux topics are nested sous ce hub — they’re in the sidebar aussi, and they’ll lien back ici.

The stage avant ce un — how bots découvrir and download votre pages — lives in the exploration hub; explorer → index is the pipeline, and une page has to clair exploration avant quelconque of ce s’applique. Pour the broader picture, voir How Search Fonctionne.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.