Guide : Sitemaps

Ce que a sitemap is, si vous besoin un, the types (XML, HTML, image, video, news), how to trouver, créer, and submit un, and pourquoi sitemaps aider discovery but jamais guarantee indexation.

Première publication : 22 juin 2026 · Dernière mise à jour : 3 août 2026 · Advanced
Langues
1 indice probant sur cette page

A sitemap is a fichier que listes l’URLs (and media) on votre site so moteur de recherches peut découvrir les. It helps discovery and is a great coverage/diagnostic outil — but submitting un jamais guarantees exploration or indexation. A unique sitemap caps at 50 000 URLs or 50MB uncompressed; Google ignores <priority>/<changefreq> but reads an accurate <lastmod>. Automate it from lune pages vous en réalité have, liste seulement canonical indexable URLs, and submit it in Search Console + Bing Webmaster Outils (the old ping endpoint is dead). Ce hub routes vous to the deep dives: XML sitemap, sitemap index, image sitemap, and video sitemap.

TL;DR — A sitemap is a fichier listing votre URLs (and optionally images/videos) so engines peut découvrir les — a coverage and diagnostic outil, pas a discovery silver bullet, parce que submitting jamais guarantees exploration or indexation. A unique sitemap caps at 50 000 URLs or 50MB uncompressed (gzip is allowed; the cap is the uncompressed size); UTF-8; fully-qualified absolute URLs; entity-escape &, ', ", <, >. Google ignores <priority> and <changefreq> and seulement uses <lastmod> si it’s consistently accurate. Submit via Search Console, Bing Webmaster Outils, and the robots.txt Sitemap: line — the old ping endpoint has 404’d since Jan 2024, and IndexNow notifies Bing and others (pas Google). Construire it automatically, liste seulement canonical indexable URLs, and utiliser the submitted-vs-indexed signal in Search Console to trouver what’s manquant.

Ce que a sitemap is — and ce que it’s pour

Google’s definition is the correct anchor: “A sitemap is a fichier où vous provide information à propos de lune pages, videos, and autre fichiers on votre site, and the relationships entre les.” Evidence for this claim A sitemap supplies search engines with information about pages, videos, and other files on a site and their relationships. Scope: Google's general definition of sitemaps; supported formats and extensions have additional requirements. Confidence: high · Verified: Google Search Central: Learn about sitemaps It’s an inventory vous hand to moteur de recherches.

The la plupart important framing I peut give vous: a sitemap is a coverage and diagnostic outil, pas a discovery silver bullet. Google is explicit que “A sitemap helps moteur de recherches découvrir URLs on votre site, but it doesn’t guarantee que tout the items in votre sitemap va be crawled and indexé.” Submitting une URL doesn’t index it. The réel day-to-day payoff is the submitted-vs-indexed comparison in Search Console: split votre sitemaps by section or content type and vous pouvez voir qui partie of le site isn’t getting picked up. During a migration I’ll garder a sitemap of the old URLs autour pour a pendant que on objectif, specifically so I peut watch les drop out of the index in GSC.

Do vous besoin un? Resolving the contradiction

You’ll voir two pieces of advice que regarder comme ils conflict:

  • Google’s threshold: vous probable don’t besoin a sitemap si le site is petit (autour 500 pages) and bien internally lié; vous probable do si it’s grand, nouveau with few backlinks, or media/News-heavy.
  • Mueller’s baseline: “Making a sitemap fichier automatically seems comme a minimal baseline pour quelconque serious website, imo.”

Les deux are correct, and here’s how I reconcile les: a sitemap isn’t strictly requis pour a tiny, well-linked site, but it’s cheap insurance and it’s now the attendu baseline — so toujours generate un, automatically. Qui brings me to the un rule I won’t bend on.

Automate it, or it rots

Sitemaps devrait be generated automatically from lune pages vous en réalité have. Si someone hands me a manually-built sitemap, I déjà know it’ll fall out of date fast — somebody adds pages, nobody updates the fichier, and now it’s lying to Google. And there’s a subtler point: si you’re generating le sitemap from votre réel, crawlable pages anyway, moteur de recherches peut usually reach ceux pages on leur propre via liens. Le sitemap earns its garder by being current and complet, pas simplement by existing. Automation is ce que rend it current.

The types of sitemap

  • XML sitemap — the workhorse, written pour robots d’exploration. A liste of <loc> URLs, optionally with <lastmod>. Ce is ce que personnes usually mean by “sitemap.”
  • HTML sitemap — a human-facing page of liens. Utile pour utilisateurs and pour internal linking, but it isn’t the fichier vous submit to Search Console.
  • Image sitemap — surfaces images moteur de recherches pourrait sinon miss (pour exemple, images chargé by JavaScript or served dynamically). Peut be a standalone fichier or image tags ajouté to an existing sitemap.
  • Video sitemap — donne moteur de recherches the metadata ils besoin to comprendre and index votre videos (thumbnail, title, description, and the media or player URL).
  • News sitemap — pour sites in Google News, listing recent articles.

And XML isn’t the seulement accepted format. Google aussi takes RSS/Atom feeds and a plain-text .txt fichier (un URL per line). RSS and text formats peut seulement liste page URLs — aucun image or video metadata — so XML is the la plupart versatile.

XML vs HTML

Ces obtenir confused constantly, so garder les separate: an XML sitemap is pour robots d’exploration — machine-readable, submitted to Search Console, the chose moteur de recherches parse. An HTML sitemap is pour personnes — a regular web page linking to votre sections, qui peut aussi aider maillage interne. Si vous seulement construire un, construire the XML sitemap; that’s the un moteur de recherches en réalité utiliser.

How to trouver a sitemap (and pourquoi some sites hide theirs)

To trouver a sitemap, essayer, in order:

  • /sitemap.xml — the par défaut emplacement on la plupart platforms.
  • robots.txt — a public sitemap is usually named ici with a Sitemap: line, qui is aussi how engines auto-discover it.
  • site: operators / Search Console — pour votre propre property, the Search Console Sitemaps report is the authoritative view of what’s submitted and how it’s doing.

But here’s the partie la plupart guides skip: pas every sitemap is meant to be trouvé. Some sites deliberately omit the Sitemap: line from robots.txt and submit the fichier seulement via Search Console and Bing Webmaster Outils. A sitemap referenced in robots.txt is readable by anyone — notamment competitors who’d love a tidy liste of every URL vous publish (unlinked pages, nouveau launches, strategically pages importantes) and a façon to watch how fast vous ship. A sitemap submitted seulement via the consoles is effectively private. So si vous pouvez’t trouver a site’s sitemap at the usual paths, it may simplement be unlisted, pas absent.

The trade-off is réel, though: hiding le sitemap from robots.txt aussi signifie quelconque engine que relies on robots.txt auto-discovery won’t voir it unless you’ve submitted it in chaque console. Si vous go the private route, vous have to do the submission legwork everywhere vous care à propos de.

How to créer un

Almost every platform peut generate a sitemap pour vous — and on la plupart of les it’s automatic. I’ve put the per-platform paths (WordPress, Shopify, Wix, Squarespace, Webflow, Drupal, JS frameworks, and manual generators) in the Cheat Sheets tab so vous pouvez jump straight to yours. La version courte pour JS frameworks: search the framework nom plus “sitemap” (Par exemple, “Gatsby sitemap” or “Suivant.js sitemap”) — there’s almost toujours an existing module so vous don’t hand-roll it. Ce site runs on Astro and uses @astrojs/sitemap, configuré in astro.config.mjs.

How to submit it

The live méthodes:

  • Recherche Google Console — le sitemaps report (and the Search Console API). The principal méthode.
  • Bing Webmaster Outils — submit là aussi; Bing récupère it immédiatement, alors rechecks roughly daily.
  • robots.txt Sitemap: line — fonctionne pour quelconque engine que reads robots.txt, and is how a public sitemap obtient auto-discovered. Utiliser the complet absolute URL.

What’s dead: the old standalone ping endpoint (google.com/ping?sitemap=). It’s been deprecated since 2023 and renvoie a 404 since January 2024 — Google killed it parce que the vast majority of submissions were spam. Si you’ve got a plugin or cron job encore hitting it, strip que out. Pour fast, URL-level notification of unique changements to Bing and others, utiliser IndexNow — but remarque Google ne fait pas participate in IndexNow, so it won’t speed up Google indexation. Bing’s propre framing is que the two are complementary: sitemaps pour comprehensive coverage, IndexNow pour fast per-URL pushes.

Submission problems

Quand something goes incorrect, the Search Console Sitemaps report va tell vous — but the error étiquettes aren’t toujours self-explanatory. I’ve put a complet error → meaning → fix table in the Cheat Sheets tab. The greatest hits: “Couldn’t fetch” (incorrect URL, robots block, or simplement pas processed yet — souvent transient), “Unsupported format” (vous submitted an HTML page au lieu de a réel XML/RSS/Atom/txt fichier), and “URL pas allowed” (the classic HTTP-vs-HTTPS / www mismatch, or URLs ci-dessus le sitemap’s propre chemin). A couple of non-error gotchas worth knowing: Google may serve vous a stale mis en cache copy, so changements aren’t instant, and submitting les deux the children and le sitemap index is harmless but unnecessary.

Meilleur practices — and “exclude ≠ noindex”

The rules que matter, tout from Google’s propre spec:

  • Size: a unique sitemap holds at la plupart 50 000 URLs or 50MB uncompressed, whichever comes premier. Vous pouvez gzip it — the 50MB cap is the uncompressed size. Evidence for this claim Google limits a single sitemap to 50,000 URLs or 50 MB uncompressed. Scope: Google-supported sitemap files; larger inventories must be split across multiple sitemaps, optionally joined by an index. Confidence: high · Verified: Google Search Central: Build and submit a sitemap Over the limite, split into multiple sitemaps and tie les ensemble with a sitemap index.
  • Encoding: UTF-8.
  • URLs: fully-qualified, absolute URLs. Entity-escape &, ', ", <, and >.
  • Tags Google ignores: <priority> and <changefreq> ne faites pashing — don’t bother with les.
  • lastmod fait honestly: Google seulement trusts lastmod si it’s consistently and verifiably accurate. Définir it on significant updates (principal content, structured données, or liens) — pas a blanket copyright-year or “today” stamp on every URL. Lie à propos de it and Google arrête believing the field. Fait correct, an accurate lastmod genuinely helps re-crawling (Bing leans on it même harder que Google fait).
  • Contents: liste seulement canonical, indexable, 200-status URLs. Aucun redirections, aucun non-URL canoniques, aucun noindex’d pages. Quelconque page vous vouloir indexé devrait be in le sitemap; nothing vous don’t.

Worked audit: valid XML, polluted inventory

A sitemap peut réussir XML validation and encore send contradictory discovery signals. Ce illustrative explorer joins chaque <loc> to its live réponse and index contrôle:

Sitemap URLObserved stateGarder?Action
https://shop.example/products/trail-runner200, canonical, indexableYesGarder
https://shop.example/products/old-trail-shoe301 to the current productAucunReplace with the final URL
https://shop.example/account/login200 with noindexAucunSupprimer from le sitemap
https://shop.example/sale/spring-2025Expired campaign returning 200Usually aucunRedirection, retire, or intentionally maintain
https://staging.shop.example/products/testPublic staging hostnameAucunSupprimer and protéger the environment

Le sitemap fichier itself is bien formed. The pollution apparaît seulement après comparing its inventory with status, canonical, robots, and lifecycle evidence. Fix the generator or source requête plutôt que deleting the même rows by hand every release.

Que dernier point hides the unique la plupart courant conceptual error: exclude ≠ noindex. Leaving une URL out of votre sitemap fait pas deindex it. Le sitemap is an advertisement, pas a gate — dropping une URL simplement arrête vous advertising it; it doesn’t supprimer it from Google. Si vous vouloir une page gone, autoriser exploration and utiliser noindex. The sitemap is the incorrect outil pour que job.

Où to go suivant

Ce page is the overview. Chaque of ces is its propre deep dive nested sous the discovery topic:

  • XML sitemap — the format and anatomy: <urlset>, <loc>, <lastmod>, the ignored tags, ce que to inclure and exclude, and hreflang in sitemaps.
  • Sitemap index — the “sitemap of sitemaps” pour grand sites, quand to split, and the math on how nombreux URLs vous pouvez cover.
  • Image sitemap — surfacing images moteur de recherches pourrait miss, the current (pas deprecated) tag liste, and cross-domain rules.
  • Video sitemap — the requis tags, accepted fichier types, and getting videos understood and indexé.

Sitemaps are un half of discovery — the autre half is exploration, qui is how engines en réalité récupérer l’URLs votre sitemap points to. Pour the whole picture, voir How Search Fonctionne. Every topic ci-dessus is in the sidebar aussi.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.