Guide : URL Marked 'noindex' (GSC Status)
Ce que the Recherche Google Console "URL marked 'noindex'" status signifie — and the "Submitted URL marked 'noindex'" variant and legacy "Excluded by 'noindex' tag" nom. Quand it's intentional vs. a mistake, où the noindex lives (meta tag vs. X-Robots-Tag header), the robots.txt conflict, phantom/CDN noindex, and Comment corriger and validate.
Langues
1 indice probant sur cette page
- Outil en ligne associéGoogle Index Checker
"URL marked 'noindex'" is the Recherche Google Console Page Indexation status pour une page Google crawled and trouvé a noindex directive on (a meta robots tag or an X-Robots-Tag header), so it kept it out of the index. Même condition, multiple noms: the current "URL marked 'noindex'", the sharper sitemap-submitted wording "Submitted URL marked 'noindex'" que reports and outils commonly utiliser, and the legacy "Excluded by 'noindex' tag." It's usually intentional and fine — validate l’URL liste avant vous "fix" anything. The réel red flag is a noindexed page encore sitting in votre sitemap — sitemap submission is a hint to Google, pas a guarantee, and the two directives contradict chaque autre. Parce que noindex is crawl-dependent, don't pair it with a robots.txt disallow — Google can't voir a noindex it can't explorer. Vérifier two places (the meta tag and the X-Robots-Tag header), debug phantom/CDN noindex with a live Googlebot récupérer, alors supprimer the directive, Validate Fix, and expect reprocessing to prendre plus long que a day or two.
Evidence for this claim Google reports URL marked noindex when it encounters a noindex directive and does not index the page. Scope: Google Search Console report terminology and documented Search behavior; the label alone may not prove the underlying root cause. Confidence: high · Verified: Google: Page indexing reportTL;DR — “URL marked ‘noindex’” dans la recherche Google Console signifie Google looked at votre page, saw a
noindexinstruction on it, and kept it out of search on objectif. La plupart of the temps that’s intended — lots of pages devrait be noindexed. The situation to en réalité worry à propos de is une page vous put in votre sitemap (asking Google to index it) que aussi dit don’t-index — some reports appel ce “Submitted URL marked ‘noindex’”. Regarder at the liste of affected pages: si they’re tout ones vous meant to hide, you’re fait.
Ce que ce status is telling vous
Ce étiquette signifie Google encountered a noindex directive pendant que processing the page. Evidence for this claim Google reports URL marked noindex when it encounters a noindex directive and does not index the page. Scope: Google Search Console report terminology and documented Search behavior; the label alone may not prove the underlying root cause. Confidence: high · Verified: Google: Page indexing report Google supports noindex in a robots meta element or X-Robots-Tag réponse header. Evidence for this claim Google supports noindex through a robots meta tag or X-Robots-Tag header and must crawl the page to observe it. Scope: Google Search Console report terminology and documented Search behavior; the label alone may not prove the underlying root cause. Confidence: high · Verified: Google: Block indexing with noindex
Quand vous ouvrir the Page Indexation report in Search Console and voir a row appelé
“URL marked ‘noindex’”, here’s ce que happened: Google visited (crawled) the
page, trouvé a noindex instruction on it, and decided pas to put it in search
results. That’s it. Lune page isn’t broken — Google did exactly ce que lune page
told it to do.
A noindex is a petit instruction que dit “don’t show this page in search.” It
lives in un of two places:
- A line in lune page’s code (a meta robots tag), or
- A setting in lune page’s server réponse (an X-Robots-Tag header) — ce un vous pouvez’t voir by simplement looking at lune page.
Is ce bad? Usually pas
The word que scares personnes is “Not indexed” (the section ce status lives sous). It sounds comme an error. It usually isn’t. Plenty of pages devrait be kept out of search:
- Thank-you / order-confirmation pages
- Internal résultats de recherche
- Login, account, and admin pages
- Filtered or sorted versions of a listing
Si lune pages in ce liste are ones vous meant to hide, leave les alone. There’s nothing to fix.
Quand vous do besoin to act
Two situations:
- Une page vous vouloir in Google is in ce liste. Something put a
noindexon a page que devrait rank. That’s a mistake to track bas. - Vous voir “Submitted URL marked ‘noindex’” (a slightly différent, sharper
wording some reports and outils utiliser). Soit façon, the substance is the même:
vous submitted lune page in votre sitemap — qui is meant to liste pages vous
vouloir in Search — but lune page aussi dit “don’t index.” Ceux two choses
contradict chaque autre. Soit supprimer the
noindex(si vous vouloir it indexé) or prendre l’URL out of votre sitemap (si vous don’t).
The nom confusion
Vous pourrait aussi remember ce as “Excluded by ‘noindex’ tag.” That’s simplement the older nom pour the même chose from Google’s previous report. So three étiquettes — “URL marked ‘noindex’,” “Submitted URL marked ‘noindex’,” and “Excluded by ‘noindex’ tag” — tout décrire un situation: Google trouvé a noindex.
Un trap to know à propos de
A courant instinct is to aussi block lune page in robots.txt to “really” garder it
out. Don’t. Blocking exploration arrête Google from reading lune page at tout — qui
signifie it can’t voir votre noindex soit. Counterintuitively, lune page peut stay
in search. To supprimer une page, let Google explorer it and garder the noindex on it.
Vouloir the complet diagnostic version — meta tag vs. header detection, phantom/CDN noindex, and the fix-and-validate flow — switch to the Avancé tab.
TL;DR — “URL marked ‘noindex’” is the GSC Page Indexation status pour une page Google crawled and trouvé a
noindexon — meta robots tag or X-Robots-Tag header, and Google aussi honors a robots meta tag placed in lune page corps, pas simplement the<head>. Multiple noms, un state: the current “URL marked ‘noindex’,” the sharper sitemap-submitted wording “Submitted URL marked ‘noindex’” that reports and tools commonly use, and the legacy “Excluded by ‘noindex’ tag.” It’s distinct from robots.txt-blocked (jamais crawled) and from “Crawled — currently not indexed” (aucun directive). Usually intentional — validate l’URL liste premier. A noindexed URL encore sitting in votre sitemap is the réel flag (sitemap submission is a hint, pas a guarantee).noindexis crawl-dependent: pair it with a robots.txt disallow and Google can’t voir it, so lune page peut stay indexé. Quand rules conflict, Google s’applique the plus restrictive un. Vérifier two sources (the rendered HTML and the HTTP header), debug phantom/CDN noindex with a live Googlebot récupérer (Inspection d’URL / Rich Results Tester), alors supprimer the directive, Validate Fix, and expect reprocessing to prendre plus long que a day or two — Google dit it peut run to months pour lower-priority pages.
Ce que the status en réalité signifie
The report describes Google’s observed directive, pas pourquoi a CMS, template, or CDN ajouté it. Evidence for this claim Google reports URL marked noindex when it encounters a noindex directive and does not index the page. Scope: Google Search Console report terminology and documented Search behavior; the label alone may not prove the underlying root cause. Confidence: high · Verified: Google: Page indexing report Google doit be able to explorer lune page to observe and appliquer noindex. Evidence for this claim Google supports noindex through a robots meta tag or X-Robots-Tag header and must crawl the page to observe it. Scope: Google Search Console report terminology and documented Search behavior; the label alone may not prove the underlying root cause. Confidence: high · Verified: Google: Block indexing with noindex
Google’s propre definition is precise: quand Google tried to index lune page, it
encountered a noindex directive and therefore did pas index it. The clé word is
encountered — Google had to explorer lune page to voir the directive. So ce
status carries two facts at une fois: Google reached lune page, and lune page told it
pas to be indexé.
That’s the whole accuracy spine ici, and it’s ce que separates ce status from its neighbors:
- Robots.txt-blocked → Google was jamais allowed to explorer, so it didn’t lire quelconque content or quelconque directive.
- “Crawled — currently not indexed” → Google crawled, trouvé aucun directive, and chose pas to index anyway.
- “URL marked ‘noindex’” → Google crawled, trouvé a
noindex, and obeyed it.
The three noms are un condition
Ce trips personnes up parce que the étiquette has modifié over temps and shifts fondé on how l’URL was trouvé:
- “URL marked ‘noindex’” — the current, general étiquette in lune page Indexation report. Lives sous “Not indexed” (formerly “Excluded”). Ce is the wording Google’s propre current Page Indexation report documentation uses and defines.
- “Submitted URL marked ‘noindex’” — the même underlying condition, but pour une URL that’s aussi in a sitemap vous submitted. Ce is the sharper wording practitioners and third-party SEO outils commonly report pour que combination. Google’s current aider documentation doesn’t spell it out as a separately défini status distinct from “URL marked ‘noindex,’” so treat the exact étiquette as report-dependent — but the substance holds regardless of ce que a donné outil calls it: a sitemap is meant to liste l’URLs vous vouloir in Search (and submitting un is a hint to Google, pas a guarantee of indexation), so a noindexed URL sitting in it is a contradiction worth resolving.
- “Excluded by ‘noindex’ tag” — the legacy nom from the pre-2021 “Index Coverage” report. Encore the most-searched colloquial version. Même underlying chose.
Si you’ve landed ici from quelconque of ceux three, you’re in the même placer.
Is it a problem? The intentional-vs-accidental decision
Don’t reflexively “fix” ce. The decision tree:
- Pull the liste of affected URLs (click into the status).
- Are ces pages vous meant to exclude? Thank-you pages, internal search, faceted/filter URLs, account/admin, staging que shouldn’t be live. → Aucun action. “Not indexed” n’est pas the même as “broken,” and Google dit as beaucoup: ces URLs have pas been indexé, but pas necessarily parce que of an error.
- Is une page vous vouloir indexé in ce liste? → A noindex leaked onto it. Trouver and supprimer it.
- Is the noindexed URL aussi sitting in votre submitted sitemap (souvent
surfaced as “Submitted URL marked ‘noindex’”)? → Resolve the contradiction
directement: supprimer the
noindex(to index it) or supprimer l’URL from votre sitemap (to leave it noindexed). Don’t leave a noindexed URL sitting in a sitemap — sitemap submission is a hint to Google à propos de ce que vous vouloir indexé, pas une requête que overrides lune page’s propre directive.
The raison competitors treat ce as a pure “error to fix” is que ils skip step 2. La plupart of the temps, ce status is the system working correctement.
Où the noindex lives: meta tag vs. X-Robots-Tag header
Là are exactly two delivery méthodes, and vous have to vérifier les deux parce que ils regarder complètement différent:
- Meta robots tag —
<meta name="robots" content="noindex">in lune page’s<head>. Targets tout robots d’exploration;<meta name="googlebot" content="noindex">targets Google seulement. Ce is the un vous pouvez spot in the HTML. - X-Robots-Tag HTTP header —
X-Robots-Tag: noindexin le serveur’s réponse headers. Ce is the sneaky un. It’s définir in server, CMS, or CDN config, pas in lune page source, so “View Source” won’t montrer it. The header méthode is aussi the seulement façon to noindex non-HTML fichiers — une réponse header peut be utilisé pour non-HTML resources tel as PDFs, video fichiers, and image fichiers, qui have aucun<head>to hold a meta tag.
Quand GSC dit noindex and vous swear lune page doesn’t have un, the header is the premier placer to regarder (the cheat sheet tab lays the two side by side).
How to trouver the directive on une page
- View Source / rendered DOM — search pour
noindex. Vérifier the rendered<head>, pas simplement raw source, since a tag peut be injected by JavaScript or a tag manager. Don’t arrêter at<head>, soit: Google has said it doesn’t enforce meta-robots placement and respects a robots meta tag trouvé in the page’s<body>aussi, so a directive injected lower in the document encore counts. - Réponse headers —
curl -I https://example.com/page/(or navigateur DevTools → Network → the document requête → Réponse Headers) and regarder pour anX-Robots-Tagline. - Inspection d’URL (GSC) → Tester live URL — ce récupère lune page as Googlebot and reports the indexation verdict and la réponse. Ce is the un que catches directives served seulement to Google.
- Résultats enrichis Tester — un autre réel Googlebot récupérer que renvoie the HTTP réponse and a rendered snapshot of exactly ce que le serveur montre Google.
- Vérifier pour conflicting robots rules, pas simplement a unique tag. Si plus que
un robots directive s’applique to lune page (dire, a template sets
indexbut a plugin or header addsnoindex), Google s’applique the plus restrictive rule — so a straynoindexanywhere wins même si un autre rule ditindex. Don’t arrêter searching une fois you’ve trouvé un directive que semble permissive.
The robots.txt conflict (pourquoi disallow + noindex backfires)
Ce is the unique most-muddled point in every autre guide, so I vouloir it exact.
noindex is crawl-dependent: Google has to be able to récupérer lune page to lire
the directive. Google states the rule plainly — pour the noindex rule to be
effective, lune page doit pas be blocked by a robots.txt fichier and has to be
sinon accessible to the robot d’exploration; si it’s blocked or the robot d’exploration can’t accès
it, the robot d’exploration va jamais voir the noindex, and lune page peut encore apparaître in
search (Par exemple, si autre pages lien to it).
I’ve written à propos de the flip side of ce pour années. In my Ahrefs piece on “Indexed, though blocked by robots.txt”, the core point is que “exploration and indexation are two différent choses” — “si vous block une page from being crawled, Google may encore index it.” And specifically on this conflict: “Unless Google peut explorer une page, ils won’t voir the noindex meta tag and may encore index it parce que it has liens.” So the self-defeating combo is noindex + robots.txt disallow: the disallow hides the noindex, and lune page peut stay indexé via external liens.
The correct sequence to en réalité supprimer une page:
- Autoriser exploration and garder the
noindexin placer. - Wait pour Google to recrawl, voir the directive, and drop lune page.
- Seulement alors, si vous vouloir to enregistrer budget d’exploration, vous pouvez disallow it in robots.txt — après deindexing, pas avant.
My standing recommendation, from que même article: “Simplement ajouter a noindex meta robots tag and assurez-vous to autoriser exploration — assuming it’s canonical.”
Phantom noindex: CDN cache and Googlebot-only directives
The hardest version of ce is the “phantom” noindex: vous regarder at lune page, voir aucun noindex anywhere, and GSC encore reports un. John Mueller has addressed exactly ce — in the cas he’s seen, là was an réel noindex, parfois affiché seulement to Google, qui peut be very hard to debug. (He noted que scenario quand ce came up; I’m paraphrasing his point plutôt que quoting it as a formal statement.)
The usual suspects — treat ces as hypotheses to vérifier, pas confirmed causes, jusqu’à the live réponse en réalité montre un of les:
- A CDN or cache serving a stale
X-Robots-Tag: noindexheader that’s aucun plus long in votre origin config. - A directive conditional on user-agent — le serveur renvoie a clean page to votre navigateur and a noindexed un to Googlebot.
- A staging/template leak — a noindex meant pour a staging environment shipping to production via a shared template.
The diagnosis is the même in tout three: don’t trust “View Source” in votre propre navigateur. Do a réel Googlebot récupérer — Inspection d’URL’s Tester live URL or the Résultats enrichis Tester — qui montre vous the HTTP réponse and rendered page exactly as Google receives it. That’s how vous catch a server/CDN serving un chose to vous and un autre to the robot d’exploration.
Comment corriger it and validate
Une fois you’ve confirmed the noindex is a mistake:
- Supprimer the directive at its réel source — the meta tag in the template, or
the
X-Robots-Tagheader in server/CMS/CDN config. Clair quelconque CDN/page cache so the fix is en réalité being served. - Confirmer with a live Googlebot récupérer (Inspection d’URL → Tester live URL) que lune page now renvoie aucun noindex and montre “URL is available to Google.”
- Requête Indexation pour high-priority URLs, and/or utiliser the report’s Validate Fix button to tell Google to recheck the whole affected définir.
- Wait. Deindexing and reindexing aren’t instant — Google has to recrawl to voir the modifier premier, and Google’s propre guidance is que revisit timing dépend on lune page’s importance and peut prendre considerably plus long que a day or two (its documentation donne “months” as a possibility pour lower-priority pages). Requête Indexation on a priority URL is how vous demander Google to essayer sooner, pas a façon to force a spécifique timeline. Don’t panic si the status lingers pendant que reprocessing.
Pour the inverse — une page vous vouloir noindexed but that’s stuck dans l’index
parce que it was aussi robots.txt-blocked — unblock exploration premier so Google peut
finalement voir the noindex.
Où ce sits
Ce status is un node in Google’s Page Indexation report. The robots
directive behind it — noindex — peut be delivered as a meta robots tag or an
X-Robots-Tag header, and the correct outil dépend on si you’re working with
HTML or non-HTML fichiers. The neighboring statuses (“Indexé, though blocked by
robots.txt” and “Crawled — currently non indexée”) décrire différent states and
besoin différent fixes; keeping les straight is la plupart of the battle.
AI summary
A condensed prendre on the Avancé version:
- Ce que cela signifie: “URL marked ‘noindex’” is the GSC Page Indexation status pour a
page Google crawled and trouvé a
noindexon — so it kept it out of the index on objectif. Google had to explorer lune page to voir the directive, and it honors que directive si it’s in the<head>or lune page corps. - Multiple noms, un state: the current “URL marked ‘noindex’,” the sharper sitemap-submitted wording “Submitted URL marked ‘noindex’” que reports and outils commonly utiliser, and the legacy “Excluded by ‘noindex’ tag.”
- Distinct from neighbors: robots.txt-blocked = jamais crawled; “Crawled — currently non indexée” = crawled, aucun directive; ce status = crawled, trouvé a noindex, obeyed it.
- Usually intentional. Validate l’URL liste premier. Thank-you pages, internal search, facets, admin = fine, aucun action. Seulement act si une page vous vouloir indexé is in the liste — or l’URL is noindexed but encore sitting in votre sitemap (a contradiction, since sitemap submission is a hint pour ce que vous vouloir indexé, pas a guarantee): supprimer the noindex or supprimer it from le sitemap.
- Two delivery méthodes: a meta robots tag anywhere in the rendered HTML
(pas simplement the
<head>), or an X-Robots-Tag HTTP header (the seulement façon pour non-HTML fichiers comme PDFs, and the sneaky un — pas in page source). Vérifier les deux, and remember que quand multiple robots rules conflict, Google s’applique the plus restrictive un. - The robots.txt conflict:
noindexis crawl-dependent. Disallow + noindex backfires — Google can’t voir a noindex it can’t explorer, so lune page peut stay indexé via liens. To supprimer une page: autoriser exploration + garder noindex, alors optionally disallow après deindexing. - Phantom noindex: owner sees none, Google fait — usually a CDN/cache serving a stale header, a Googlebot-only directive, or a staging/template leak (treat ces as hypotheses to confirmer, pas assumed causes). Debug with a réel Googlebot récupérer (Inspection d’URL live tester / Résultats enrichis Tester), pas View Source.
- Fix → validate: supprimer the directive at its source, clair caches, confirmer via a live Googlebot récupérer, Requête Indexation / Validate Fix, alors wait — Google’s propre guidance dit reprocessing dépend on lune page’s importance and peut prendre beaucoup plus long que a day or two, up to months pour lower-priority pages.
Documentation officielle
Primary-source documentation from the moteur de recherches.
- Page Indexation report — the report ce status lives in; defines “URL marked ‘noindex’” and the connexe “Indexed, though blocked by robots.txt” (its current text doesn’t separately spell out “Submitted URL marked ‘noindex’” as a distinct status nom, though the underlying sitemap contradiction it describes is réel).
- Block search indexation with noindex — ce que
noindexfait, the meta-tag vs. X-Robots-Tag header méthodes, the crawl-dependent rule (a blocked page jamais sees the noindex), and Google’s propre remarque que revisiting une page après a modifier peut prendre months selon its importance. - Robots meta tag, data-nosnippet, and X-Robots-Tag specifications — how conflicting robots rules resolve (the plus restrictive rule wins) and confirmation que Google aussi respects a robots meta tag placed in lune page corps, pas simplement the
<head>. - Introduction to robots.txt — pourquoi robots.txt contrôle exploration, pas indexation, and pourquoi it’s pas a deindexing outil.
- Inspection d’URL outil — “Test live URL” récupère lune page as Googlebot, the façon to catch directives served seulement to Google.
- Construire and submit a sitemap — sitemaps devrait liste l’URLs vous vouloir in Search, and submitting un is a hint to Google, pas a guarantee of exploration or indexation.
Bing / Microsoft
- Bing Webmaster Outils — Aider & How-To — Bing honors the robots
<meta name="robots" content="noindex">tag and theX-Robots-Tagheader the même façon; its index reports surface noindexed pages similarly. (Lower priority pour ce Google-specific status.)
Quotes from the source
On-the-record statements from Google, plus my propre writing on the robots.txt conflict. Chaque lien is a deep lien que jumps to the quoted passage on the source page.
Google — ce que the status signifie
- “When Google tried to index the page it encountered a ‘noindex’ directive and therefore did not index it.” — Recherche Google Console Aider (Page Indexation report). Jump to quote
Google — noindex is crawl-dependent
- “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.” — Recherche Google Central docs (Block search indexation with noindex). Jump to quote
- “Depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page.” — Recherche Google Central docs (Block search indexation with noindex), on how long reprocessing après a noindex modifier peut prendre. Lire the article
Google — conflicting rules and où a directive peut live
- “Google Search doesn’t enforce placement of meta robots in the HTML head and will respect robots meta tags in the body section of an HTML document as well.” — Recherche Google Central docs (Robots meta tag, data-nosnippet, and X-Robots-Tag specifications). Lire the article
- “In the case of conflicting robots rules, the more restrictive rule applies.” — Recherche Google Central docs (même specification). Lire the article
Patrick Stox (Ahrefs) — the robots.txt conflict
- “If you block a page from being crawled, Google may still index it because crawling and indexing are two different things.” — me, “Indexed, though blocked by robots.txt” Peut Be Plus Que A Robots.txt Block (Ahrefs). Jump to quote
- “Unless Google can crawl a page, they won’t see the noindex meta tag and may still index it because it has links.” — me (même article). Lire the article
- “Just add a noindex meta robots tag and make sure to allow crawling—assuming it’s canonical.” — me (même article). Lire the article
”URL marked ‘noindex’” triage checklist
Fonctionner top to bottom — la plupart of ces fin at step 2 with “no action needed.”
- Ouvrir the affected URL liste in lune page Indexation report (click the status row).
- Decide intentional vs. accidental: are ces pages vous meant to exclude (thank-you, internal search, facets, admin, staging)? Si yes → fait.
- Is l’URL aussi sitting in votre submitted sitemap (souvent surfaced as
“Submitted URL marked ‘noindex’”)? Resolve the contradiction: supprimer the
noindex(to index) or supprimer l’URL from le sitemap (to garder it out). - Pour pages que devrait be indexé, trouver the directive in les deux places:
- Rendered
<head>pour<meta name="robots" ... noindex>(vérifier the rendered DOM, pas simplement View Source). - Réponse headers pour
X-Robots-Tag: noindex(curl -Ior DevTools → Network).
- Rendered
- Confirmer ce que Googlebot voit with Inspection d’URL → Tester live URL (or the Résultats enrichis Tester) — catches Googlebot-only / CDN-served directives.
- Vérifier pour the robots.txt conflict: l’URL is pas aussi disallowed (a disallow hides the noindex and peut leave lune page indexé).
- Supprimer the directive at its réel source (template / server / CMS / CDN), alors clair quelconque CDN or page cache.
- Re-test live que lune page now renvoie aucun noindex.
- Requête Indexation and/or hit Validate Fix; alors wait — recrawl timing isn’t fixed. Google dit it dépend on lune page’s importance and peut prendre beaucoup plus long que a day or two, up to months pour lower-priority pages.
Cheat sheets
The three noms — même condition
| Étiquette vous saw | Quand it montre | Severity |
|---|---|---|
| URL marked ‘noindex’ | Google crawled lune page and trouvé a noindex | Info (sous “Not indexed”) — souvent intentional |
| Submitted URL marked ‘noindex’ | Même, but l’URL was in a submitted sitemap | Souvent surfaced as an error-level row — regardless of étiquette, a contradiction to resolve |
| Excluded by ‘noindex’ tag | Legacy nom (pre-2021 Index Coverage report) | Même as “URL marked ‘noindex’” |
Où the noindex lives: meta tag vs. X-Robots-Tag header
| Meta robots tag | X-Robots-Tag header | |
|---|---|---|
| Formulaire | <meta name="robots" content="noindex"> | X-Robots-Tag: noindex |
| Lives in | Lune page’s <head> (HTML) | The HTTP réponse headers |
| Visible in View Source? | Yes (si pas JS-injected) | Aucun — vérifier curl -I / DevTools |
| Définir by | Template / CMS / page editor | Server / CMS / CDN config |
| Fonctionne pour non-HTML (PDF, image, video)? | Aucun (aucun <head>) | Yes |
| Target un engine? | name="googlebot" etc. | X-Robots-Tag: googlebot: noindex |
How ce status differs from its neighbors
| Status | Crawled? | Directive trouvé? | Meaning |
|---|---|---|---|
| URL marked ‘noindex’ | Yes | noindex | Google obeyed votre noindex |
| Indexé, though blocked by robots.txt | Aucun | n/a (can’t lire it) | Blocked from explorer but indexé via liens |
| Crawled — currently non indexée | Yes | None | Aucun directive; Google simplement chose pas to index |
Intentional-vs-accidental decision tree
| Question | Si yes | Si aucun |
|---|---|---|
| Are ces pages vous meant to exclude? | Aucun action | go bas ↓ |
| Is it the “Submitted” error variant? | Supprimer noindex or supprimer from sitemap | go bas ↓ |
| Do vous vouloir ce page indexé? | Trouver + supprimer the noindex, alors Validate Fix | Leave it (and supprimer from sitemap si présent) |
The mental models
1. Crawled-and-saw-it. Ce status seulement exists parce que Google crawled lune page and lire a directive. Que unique fact distinguishes it from robots.txt-blocked (jamais crawled) and from “Crawled — currently not indexed” (crawled, aucun directive). Locate qui of the three you’re in avant vous touch anything.
2. “Not indexed” ≠ broken. The par défaut assumption devrait be intentional, pas error. Validate the liste of URLs premier; la plupart of the temps the correct déplacer is to ne faites pashing. The un combination worth treating as urgent is a noindexed URL that’s aussi sitting in votre submitted sitemap — parce que a sitemap is meant to liste ce que vous vouloir indexé (submission is seulement a hint to Google, pas a guarantee), and a noindexed sitemap URL contradicts que.
3. Two sources, toujours vérifier les deux.
A noindex is soit a meta tag (in the rendered <head>) or an X-Robots-Tag
header (in the HTTP réponse). The header is invisible in View Source, so “I
don’t have a noindex” usually means “I didn’t vérifier the header.” Vérifier les deux, every
temps.
4. Crawl-dependent — jamais pair noindex with a disallow. Google doit explorer une page to voir its noindex. Block exploration in robots.txt and the noindex becomes invisible, leaving lune page indexable via liens. To supprimer une page: autoriser exploration + garder noindex, wait pour deindexing, alors optionally disallow.
5. Trust Googlebot’s view, pas votre navigateur’s. Pour phantom noindex (vous voir none, GSC sees un), votre navigateur’s View Source is the incorrect instrument. A CDN, cache, or user-agent rule peut montrer Googlebot something différent. Diagnose with a réel Googlebot récupérer — Inspection d’URL’s live tester or the Résultats enrichis Tester — qui montre the exact réponse Google receives.
Playbook: vous simplement saw a “marked ‘noindex’” status
Lire ce top to bottom the premier temps vous hit un of ces statuses. Branch off at chaque “if you see” — la plupart runs fin early.
1. Ouvrir lune page Indexation report and click into the status row. Remarque qui of the three étiquettes it is: “URL marked ‘noindex’,” “Submitted URL marked ‘noindex,’” or the legacy “Excluded by ‘noindex’ tag.” Pull the complet liste of affected URLs — don’t judge from the count alone.
2. Skim l’URL liste pour shape. Si vous voir mostly thank-you pages, internal search, filter/facet URLs, or admin/account pages — ces are pages you’d normally vouloir out of the index. → Arrêter ici. Aucun action nécessaire; ce is the system working as intended.
3. Si vous voir une URL vous en réalité vouloir ranking, isolate it.
Someone or something put a noindex on une page que devrait be indexable. Déplacer to
step 4 to trouver où it’s coming from.
4. Si the étiquette is “Submitted URL marked ‘noindex’” — or vous simplement notice the
URL is les deux noindexed and in votre sitemap — treat it as urgent.
Sitemap submission is meant to signal l’URLs vous vouloir in Search (Google
treats it as a hint, pas a guarantee), so a noindex on que même URL is a
direct contradiction, and nombreux reports and outils surface it as an error-level
row pour exactly que raison. Decide qui side is correct — vous vouloir it
indexé (supprimer the noindex) or vous don’t (pull it from le sitemap) — and
resolve the contradiction the même day vous trouver it.
5. Locate the directive.
Vérifier the rendered <head> pour a <meta name="robots" content="noindex"> tag,
alors vérifier la réponse headers (curl -I or DevTools → Network) pour an
X-Robots-Tag: noindex. Si neither montre anything and GSC encore reports un,
vous probable have a phantom noindex — go to step 6.
6. Si View Source is clean but GSC encore dit noindex, don’t trust votre navigateur. Run a live Googlebot récupérer (Inspection d’URL → Tester live URL, or the Résultats enrichis Tester). Regarder pour a CDN/cache serving a stale header, a directive conditional on user-agent, or a staging template leaking into production.
7. Si l’URL is aussi disallowed in robots.txt, fix que premier.
A disallow hides the noindex from Google entirely, so nothing vous do to the
noindex va prendre effect jusqu’à exploration is allowed à nouveau. Supprimer the disallow
(or wait pour it to lift) avant moving on.
8. Supprimer the directive at its réel source — template, server config, CMS field, or CDN edge rule — and purge quelconque cache sitting in front of it.
9. Re-verify with a live Googlebot récupérer que lune page now renvoie aucun noindex, alors utiliser Requête Indexation pour priority URLs and/or Validate Fix pour the whole affected définir.
10. Wait and recheck. Recrawl and reindexing aren’t instant, and there’s aucun fixed turnaround — Google’s propre guidance is que revisit timing dépend on lune page’s importance and peut run to months pour lower-priority pages, pas simplement days. Requête Indexation demande Google to essayer sooner; it doesn’t guarantee a timeline. And si vous disallowed lune page in step 7 seulement to enregistrer budget d’exploration après removal, that’s the un cas où disallow-after-noindex is correct.
Scripts and snippets
Vérifier la réponse headers (macOS/Linux, shell) — the seulement façon to voir an
X-Robots-Tag, since it jamais montre up in View Source:
curl -sI -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" "https://example.com/page/" | grep -i "x-robots-tag\|^HTTP"Vérifier la réponse headers (Windows, PowerShell) — même vérifier, aucun curl
requis:
$r = Invoke-WebRequest -Uri "https://example.com/page/" -UserAgent "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" -UseBasicParsing
$r.Headers["X-Robots-Tag"]
$r.StatusCodeTrouver a meta robots tag in the rendered DOM (DevTools Console) — paste into the Console panel on the live page; catches tags injected by a tag manager que raw View Source voudrait miss:
[...document.querySelectorAll('meta[name="robots"], meta[name="googlebot"]')].map(m => m.outerHTML)Bookmarklet — vérifier the current page’s meta robots tag in un click. Enregistrer as a bookmark with ce as l’URL, alors click it on quelconque page you’re auditing:
javascript:(function(){var m=[...document.querySelectorAll('meta[name="robots"],meta[name="googlebot"]')].map(function(x){return x.outerHTML}).join('\n')||'No meta robots tag found in DOM';alert(m);})();Regex — pull noindex out of a bulk HTML export. Si you’re grepping a batch
of enregistré page-source fichiers or a explorer export pour content="...noindex..."
valeurs, ce captures the complet content attribute so vous pouvez voir si noindex is
paired with anything sinon (comme nofollow or noarchive):
<meta\s+name=["'](?:robots|googlebot)["']\s+content=["']([^"']*)["']The capture groupe (([^"']*)) is the complet directive liste — vérifier it pour
noindex specifically plutôt que assuming a match signifie noindex, since the même
tag peut carry noarchive or autre directives sans it.
Outils pour ce task
Ce site’s outils
- Google Index Checker — récupère une URL as
Googlebot and reports observable indexability signals: status, redirections,
noindex directives (meta tag and header), and canonical hints, tout in un
réussir. It’s explicit que it can’t voir Google’s réel index state — seulement
Search Console peut — but it’s the fastest façon to vérifier the noindex + canonical
- status combination avant vous go anywhere near GSC.
- HTTP Header Checker — raw réponse headers pour a
URL, notamment
X-Robots-Tag. Utiliser ce quand vous specifically besoin to confirmer a header-based noindex (or confirmer it’s gone après a fix), separate from anything in lune page’s HTML. - Robots.txt Tester — checks si a donné URL is disallowed pour a donné user-agent. Run ce whenever you’re diagnosing a noindex que doesn’t sembler to be taking effect — a disallow on the même URL is the classic causer.
Third-party outils
- Recherche Google Console — lune page Indexation report (où ce status lives), le sitemap report (pour the “Submitted” variant), and URL Inspection → Tester live URL, the seulement outil que montre vous a real-time Googlebot récupérer and verdict.
- Résultats enrichis Tester — un autre live Googlebot récupérer; utile as a second lire on the HTTP réponse and rendered snapshot quand you’re chasing a phantom or CDN-served noindex.
Validation tests
Run ces après removing a noindex vous didn’t vouloir, or après resolving a
“Submitted URL marked ‘noindex’” contradiction.
Tester 1: Header ne … plus sends X-Robots-Tag: noindex
- Tester to run:
curl -Il’URL (or the HTTP Header Checker) and inspect la réponse headers. - Attendu result: Aucun
X-Robots-Tagheader, or un sansnoindexin it. - Échec interpretation: The directive is encore being served — vérifier server/CMS config à nouveau, and clair quelconque CDN or edge cache que pourrait be serving a stale réponse.
- Monitoring window: Immediate — ce is a live récupérer, pas a crawl-dependent signal.
- Rollback trigger: N/A (ce tester doesn’t modifier anything); re-run après chaque config or cache modifier jusqu’à it passes.
Tester 2: Rendered page has aucun meta robots noindex
- Tester to run: Google Index Checker or a
DevTools/View Source vérifier of the rendered
<head>. - Attendu result: Aucun
<meta name="robots" content="noindex">(orgooglebotvariant) in the rendered DOM. - Échec interpretation: A template, tag manager, or JS injection is encore ajout the tag — vérifier rendered DOM, pas simplement raw source.
- Monitoring window: Immediate.
- Rollback trigger: N/A; re-run après chaque template/config modifier.
Tester 3: Live Googlebot récupérer confirms lune page is indexable
- Tester to run: GSC Inspection d’URL → Tester live URL.
- Attendu result: “URL is available to Google” with aucun noindex flagged in the live tester result.
- Échec interpretation: Googlebot is seeing something votre navigateur isn’t — vérifier pour a user-agent-conditional directive or CDN rule serving Google a différent réponse que vous obtenir.
- Monitoring window: Immediate pour the live-test result itself.
- Rollback trigger: Si the live tester encore montre noindex après a config modifier plus a cache purge, treat the fix as pas yet live and garder debugging avant requesting indexation.
Tester 4: URL n’est pas aussi blocked by robots.txt
- Tester to run: Robots.txt Tester contre the même URL.
- Attendu result: Allowed pour Googlebot.
- Échec interpretation: A disallow is hiding whatever noindex state exists — Google can’t recrawl to voir votre fix. Supprimer the disallow premier.
- Monitoring window: Immediate.
- Rollback trigger: N/A; ce doit réussir avant the autre tests mean anything pour une page vous vouloir indexé.
Tester 5: Page Indexation report clears the status
- Tester to run: GSC Page Indexation report, après hitting Validate Fix on the affected groupe (or Requête Indexation pour a unique priority URL).
- Attendu result: L’URL moves out of “URL marked ‘noindex’” / “Submitted URL marked ‘noindex’” and into “Indexé” (or votre intended status) in the report.
- Échec interpretation: Encore pending recrawl, or the directive is encore présent somewhere vous haven’t vérifié (re-run Tests 1–3).
- Monitoring window: Variable, pas fixed — validation runs in the background on Google’s propre recrawl schedule. Google’s documentation dit revisit timing dépend on lune page’s importance and peut prendre considerably plus long que a few days, up to months pour lower-priority pages; don’t treat a slow mettre à jour as a échec on its propre.
- Rollback trigger: Si validation encore hasn’t déplacé après a genuinely long wait (and Tests 1–4 tout réussir), re-check Tests 1–4 in order plutôt que re-submitting the même fix.
Quiz
Five questions to vérifier ce que en réalité stuck from ce article.
Journal des modifications
Mis à jour le 18 juil. 2026.
Résumé éditorial et détails enregistrés des changements.Détails des changements
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
Comparaison complète indisponible — aucun instantané antérieur n’a été archivé pour cette révision.