Historical Page Comparison
Free, no signup. Compare one bounded Common Crawl capture with the current public page. Coverage is sampled evidence—not proof of every version that existed.
Checks run from our server; we fetch the URL you enter and don't keep the results. The exact URL is sent in a POST body to this site, then used for bounded Common Crawl index/WARC lookup and a safe current-page fetch. Results may be cached; raw archived HTML is not returned. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
Site passport Local context for this saved site
Local data
Saved targets, named lists, and recent check summaries remain only in this browser.
Observed page signals
Bounded text changes
How to use it
- Enter one exact public HTTP(S) page URL.
- Optionally choose a before date to prefer an older capture.
- Run the comparison and verify the capture timestamp and crawl index.
- Review changed signals and excerpts. A difference is a review prompt, not an automatic SEO defect.
Example result
A bounded comparison can show a capture timestamp and crawl index alongside a changed title, canonical, or structured-data type. It reports only the selected capture and current raw HTML; a missing capture or a difference is evidence to review, not a complete version history or an automatic SEO defect.
How it works
The service selects one date-relevant Common Crawl index record, reads one bounded WARC byte range, and fetches the current public page through the guarded endpoint. It extracts comparable raw-HTML facts and capped text excerpts; it does not replay archived JavaScript or subresources.
What it checks
The result compares the title, description, H1, word count, canonical URLHow search engines pick one canonical URL among duplicates and consolidate signals onto it., robots directivesThe robots meta tag is an HTML element in a page's head — <meta name="robots" content="noindex"> — that tells search engines how to index and serve that page. It's crawl-then-obey: a page blocked in robots.txt is never fetched, so the tag is never seen., and JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. types. Similarity and excerpts use normalized extracted text from bounded raw HTML; they are not a rendered-page replay.
Common Crawl attribution and limitations
Historical capture data comes from Common Crawl; this site is not affiliated with it. Common Crawl samples the web, so captures can be absent, incomplete, truncated, or represented by revisit records without a payload. The tool queries only a few date-relevant crawl indexes and never promises lifetime coverage.
The comparison shows derived observations and small bounded excerpts—not a license to republish a third party’s archived page. Review the Common Crawl Terms of Use and any rights or terms that apply to the originating page before commercial reuse.
Frequently asked questions
What does the tool compare?
It compares one exact URL’s selected Common Crawl capture with the current public page: title, H1, meta description, canonical, robots directives, word count, structured-data types, bounded text similarity, and capped added/removed excerpts. Capture and current HTTP statuses are shown separately as source observations, not as a status-history comparison.
Does a missing capture mean the page never existed?
No. Common Crawl samples the web and does not capture every URL on every crawl. Missing, incomplete, revisit-only, and truncated records are reported as coverage limits rather than converted into a pass or failure.
Can I view or download the complete archived page?
No. The service retrieves one exact WARC byte range and returns only derived facts and bounded text excerpts. It does not provide a full-page archive viewer, execute archived JavaScript, or load archived subresources.
Can I use the result commercially?
This tool reports derived observations and does not grant rights in third-party archived content. Review Common Crawl’s current terms and the originating site’s rights before relying on archived material in a commercial product.
Rate this tool
Feature requests for Historical Page Comparison
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
➕ Request a feature
New requests are reviewed before they appear here.
Araç hakkında
Tek bir Common Crawl yakalamasını güncel sayfanın sınırlı ham HTML SEO sinyalleri ve çıkarılmış metniyle karşılaştırın; kapsam örneklenmiş kanıttır, var olmuş her sürümün kanıtı değildir.
Özellikler
- Başlık, H1, meta açıklaması, canonical, robots yönergeleri ve yapılandırılmış veri türlerindeki farkları karşılaştırın.
- Yakalama zamanı, crawl dizini ve güncel sayfayı aynı sonuçta ayrı gözlemler olarak görün.
- Kelime sayısı, sınırlı metin benzerliği ve eklenen/çıkarılan metin parçalarını inceleyin.
- Common Crawl atıf ve kullanım sınırlarını, yeniden yayımlama haklarından ayrı şekilde değerlendirin.
Nasıl çalışır
Tam sayfa URL’sini girin ve isteğe bağlı tarih sınırı seçin. Araç tarih açısından uygun bir Common Crawl kaydıyla güncel ham HTML’yi karşılaştırır; başlık, canonical, robots, yapılandırılmış veri, kelime sayısı ve metin farklarını kanıt olarak inceleyin. Eksik yakalama veya farkı otomatik SEO hatası değil, gözden geçirilecek kanıt olarak değerlendirin.
Sınırlamalar
- Common Crawl web’i örnekler; eksik, yarım, kesilmiş veya yalnızca revisit kaydı olan yakalamalar kapsam sınırı olarak raporlanır, başarı ya da başarısızlığa çevrilmez.
- Araç tek bir sınırlı WARC byte aralığı okur ve güncel herkese açık sayfayı alır; arşivlenmiş JavaScript’i veya alt kaynakları çalıştırmaz ve tam sürüm geçmişi sunmaz.
Sık sorulan sorular
Araç neyi karşılaştırır?
Bir URL’nin seçilen Common Crawl yakalamasını güncel herkese açık sayfayla; başlık, H1, meta açıklaması, canonical, robots yönergeleri, kelime sayısı, yapılandırılmış veri türleri, sınırlı metin benzerliği ve eklenen/çıkarılan parçalarla karşılaştırır.
Eksik yakalama sayfanın hiç var olmadığını mı gösterir?
Hayır. Common Crawl her taramada her URL’yi yakalamaz. Eksik, yarım, revisit-only veya kesilmiş kayıtlar kapsam sınırı olarak raporlanır.
Sonucu ticari olarak kullanabilir miyim?
Sonuç türetilmiş gözlemler sağlar ve üçüncü taraf arşiv içeriğini kullanma hakkı vermez. Ticari üründe kullanmadan önce Common Crawl kullanım şartlarını ve kaynak sitenin haklarını inceleyin.
Tam arşivlenmiş sayfayı görüntüleyebilir veya indirebilir miyim?
Hayır. Hizmet tek bir WARC byte aralığından türetilmiş gerçekleri ve sınırlı metin parçalarını döndürür; tam sayfa görüntüleyici sağlamaz, arşivlenmiş komut dosyalarını çalıştırmaz veya alt kaynakları yüklemez.
Sonuç bir SEO hatasını otomatik olarak kanıtlar mı?
Hayır. Fark veya eksik yakalama incelenmesi gereken kanıttır. Araç yalnızca seçilen yakalama ile güncel ham HTML’yi raporlar; tam geçmiş veya otomatik SEO kusuru iddiasında bulunmaz.