XML Sitemap Generator
Free, no signup. Generate a sitemap from a small, robots-respecting crawl — with clear inclusion rules, an honest cap, and only evidence-backed last-modified dates.
Bounded and polite: at most three page requests run at once; each request is capped at 900 KB/10 seconds, and the server also rate-limits each target host. This is a sitemap seed, not a full-site crawler.
Checks run from our server; we fetch the URL you enter and don't keep the results. The start site, its robots.txt, and eligible same-host HTML pages are sent through this site's bounded fetch endpoints. Crawl decisions and XML generation run in your browser. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
Site passport Local context for this saved site
Local data
Saved targets, named lists, and recent check summaries remain only in this browser.
Sitemap output
Excluded
Unknown / not included
Rate this tool
How it works
- Fetch
robots.txtand apply its Googlebot rules before each candidate request. A temporary robots failure stops the crawl rather than treating it as permission. - Follow same-origin links only, with a maximum of three concurrent requests and the cap you selected.
- Include only successful HTML pages that are not noindex and do not canonicalize to another URL.
- Keep exclusions and unknown responses separate. A cap, timeout, failed response, or truncated HTML never becomes an implied clean page.
About priority and changefreq: Google ignores these sitemap hints, so this generator intentionally omits them. lastmod appears only when the page provides a verifiable date.
Frequently asked questions
Will this crawl my entire site?
No. You choose a hard cap up to 200 pages. The tool states when the queue is still non-empty at that cap, so a capped crawl is never presented as complete.
Which pages are excluded?
Pages with a noindex header or meta tag, pages whose canonical points elsewhere, and URLs disallowed for Googlebot in robots.txt are excluded. Failed, non-HTML, truncated, and non-2xx responses stay in the separate unknown state.
How is lastmod chosen?
Only a valid HTTP Last-Modified header, visible time datetime, or JSON-LD dateModified/datePublished value is emitted. The tool never stamps every page with today’s date.
Feature requests for Xml Sitemap Generator
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
➕ Request a feature
New requests are reviewed before they appear here.
Über das Tool
Crawle eine begrenzte Menge derselben Website, respektiere robots.txt und erzeuge daraus eine ehrliche XML-Sitemap. Noindex- und off-canonical-URLs werden ausgeschlossen; fehlgeschlagene, nicht-HTML-, abgeschnittene und nicht erfolgreiche Antworten bleiben im separaten Unbekannt-Status.
Der Crawl ist ein Sitemap-Seed, kein vollständiger Website-Crawler und wird nie als vollständig dargestellt, solange die Queue am Seitenlimit nicht leer ist.
Funktionen
- Begrenzter Same-Site-Crawl mit robots-Respekt und wählbarem Seitenlimit bis 200
- Ausschluss von Noindex-, Off-Canonical- und Googlebot-disallowed-URLs
- Separate Zustände für Fehler, Nicht-HTML, Kürzung und Unsicherheit
- lastmod nur aus sichtbaren Header-, Datetime- oder JSON-LD-Nachweisen sowie XML-Validierungslink
Funktionsweise
Gib eine Start-URL ein und wähle das Seitenlimit. Der Crawl prüft robots.txt, folgt nur derselben Website und sammelt sichtbare Canonicals, Noindex-Signale, HTTP-Status, Content-Type und Größenlimits. Für jede eingeschlossene Seite wird lastmod nur aus einem gültigen Last-Modified-Header, einem sichtbaren time-Datetime oder JSON-LD dateModified/datePublished übernommen. Erzeuge anschließend die XML-Sitemap und validiere sie im nächsten Schritt.
Einschränkungen
- Maximal 200 Seiten werden gecrawlt; eine nicht leere Queue am Limit bedeutet, dass weitere Seiten ungeprüft sind.
- Pro Request gelten 900 KB und 10 Sekunden, höchstens drei Requests laufen parallel und Zielhosts werden zusätzlich rate-limitiert.
- lastmod wird nicht pauschal auf das heutige Datum gesetzt; fehlende Evidenz bleibt ohne Datum.
Häufig gestellte Fragen
Wie wird lastmod ausgewählt?
Nur ein gültiger HTTP-Last-Modified-Header, ein sichtbares time-Datetime oder ein JSON-LD-Wert dateModified/datePublished wird ausgegeben. Das Werkzeug versieht nicht jede Seite mit dem heutigen Datum.
Crawlt der Generator meine gesamte Website?
Nein. Du setzt ein hartes Limit bis 200 Seiten. Ist die Queue an diesem Limit noch nicht leer, zeigt das Werkzeug diesen Zustand ausdrücklich an; ein begrenzter Crawl gilt nie als vollständiger Site-Crawl.
Welche Seiten werden ausgeschlossen?
Seiten mit Noindex-Header oder Meta-Tag, mit einem anderen Canonical und URLs, die Googlebot laut robots.txt nicht abrufen darf. Fehlerhafte, nicht-HTML-, abgeschnittene und nicht-2xx-Antworten bleiben separat als unbekannt sichtbar.
Ist die Sitemap danach fertig?
Sie ist ein begrenzter Seed aus den geprüften URLs. Validiere die Ausgabe anschließend und behandle nicht gecrawlte oder unsichere URLs nicht als enthalten.