XML Sitemap Generator
Free, no signup. Generate a sitemap from a small, robots-respecting crawl — with clear inclusion rules, an honest cap, and only evidence-backed last-modified dates.
Bounded and polite: at most three page requests run at once; each request is capped at 900 KB/10 seconds, and the server also rate-limits each target host. This is a sitemap seed, not a full-site crawler.
Checks run from our server; we fetch the URL you enter and don't keep the results. The start site, its robots.txt, and eligible same-host HTML pages are sent through this site's bounded fetch endpoints. Crawl decisions and XML generation run in your browser. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
Site passport Local context for this saved site
Local data
Saved targets, named lists, and recent check summaries remain only in this browser.
Sitemap output
Excluded
Unknown / not included
Rate this tool
How it works
- Fetch
robots.txtand apply its Googlebot rules before each candidate request. A temporary robots failure stops the crawl rather than treating it as permission. - Follow same-origin links only, with a maximum of three concurrent requests and the cap you selected.
- Include only successful HTML pages that are not noindex and do not canonicalize to another URL.
- Keep exclusions and unknown responses separate. A cap, timeout, failed response, or truncated HTML never becomes an implied clean page.
About priority and changefreq: Google ignores these sitemap hints, so this generator intentionally omits them. lastmod appears only when the page provides a verifiable date.
Frequently asked questions
Will this crawl my entire site?
No. You choose a hard cap up to 200 pages. The tool states when the queue is still non-empty at that cap, so a capped crawl is never presented as complete.
Which pages are excluded?
Pages with a noindex header or meta tag, pages whose canonical points elsewhere, and URLs disallowed for Googlebot in robots.txt are excluded. Failed, non-HTML, truncated, and non-2xx responses stay in the separate unknown state.
How is lastmod chosen?
Only a valid HTTP Last-Modified header, visible time datetime, or JSON-LD dateModified/datePublished value is emitted. The tool never stamps every page with today’s date.
Feature requests for Xml Sitemap Generator
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
➕ Request a feature
New requests are reviewed before they appear here.
Araç hakkında
Aynı site içindeki sınırlı bir sayfa kümesini tarayın, robots.txt kurallarına uyun, noindex ve canonical dışı URL’leri dışlayın ve güvenilir bir XML site haritası indirin.
Özellikler
- Başlangıç URL’si ve en fazla 200 sayfalık sabit bir tarama sınırı seçin.
- robots.txt, noindex ve canonical sinyallerine göre site haritasına dahil edilen ve dışlanan URL’leri ayırın.
- Başarısız, HTML olmayan, kesilmiş ve 2xx olmayan yanıtları ayrı bilinmeyen durumda tutun.
- lastmod tarihlerini yalnızca sayfanın sağladığı HTTP veya yapılandırılmış veri kanıtı olduğunda üretin.
Nasıl çalışır
Başlangıç URL’sini ve sayfa sınırını seçin, ardından taramayı başlatın. Dahil edilen, dışlanan ve bilinmeyen URL durumlarını inceleyin; queue sınırda hâlâ doluysa sonucu tamamlanmış tam site taraması olarak yorumlamayın. Oluşturulan XML’i doğrulayın ve lastmod değerlerinin kanıtını kontrol edin.
Sınırlamalar
- Bu bir tam site tarayıcısı değildir. Aynı anda en fazla üç istek çalışır; her istek 900 KB/10 saniye ile, hedef host da hız sınırıyla kısıtlanır.
- Noindex, başka bir canonical’a işaret eden veya robots.txt ile Googlebot’a yasaklanan URL’ler dışlanır. Başarısız ya da belirsiz yanıtlar site haritasına eklenmez.
Sık sorulan sorular
lastmod nasıl seçilir?
Yalnızca geçerli bir HTTP Last-Modified başlığı, görünür bir time datetime değeri veya JSON-LD dateModified/datePublished değeri yayımlanır. Araç her sayfaya bugünün tarihini eklemez.
Bu tarama tüm sitemi tarar mı?
Hayır. En fazla 200 sayfalık bir üst sınır seçersiniz. Bu sınırda kuyruk boş değilse araç bunu açıkça belirtir; sınırlı tarama tamamlanmış gibi sunulmaz.
Hangi sayfalar dışlanır?
Noindex başlığı veya meta etiketi olan, canonical’ı başka yere işaret eden ya da robots.txt içinde Googlebot’a izin verilmeyen sayfalar dışlanır. Başarısız, HTML olmayan, kesilmiş ve 2xx olmayan yanıtlar ayrı bilinmeyen durumda kalır.
Tarama sınırı neden var?
İstekleri sınırlı ve nazik tutmak için vardır. Her seferde en fazla üç istek çalışır ve her hedef host hız sınırı uygular; bu araç tam site keşfi vaat etmez.
Site haritası çıktısında bilinmeyen URL’ler ne olur?
Başarısız, HTML olmayan, kesilmiş veya 2xx olmayan yanıtlar Excluded/Unknown olarak ayrı tutulur; kesin kanıt olmadan site haritasına uygun sayılmazlar.