Historical Page Comparison

Free, no signup. Compare one bounded Common Crawl capture with the current public page. Coverage is sampled evidence—not proof of every version that existed.

Checks run from our server; we fetch the URL you enter and don't keep the results. The exact URL is sent in a POST body to this site, then used for bounded Common Crawl index/WARC lookup and a safe current-page fetch. Results may be cached; raw archived HTML is not returned. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.

Feedback
Report a bug

Found something broken in Historical Page Comparison? Let us know what happened — this goes straight to a private triage queue, not a public list.

What will be sent
 No tool inputs, uploads, pasted source, complete results, query parameters, or URL fragments are attached automatically. You can edit or remove the selected passage above. Browser and anti-abuse metadata is processed for spam prevention. 
Local data

Saved targets, named lists, and recent check summaries remain only in this browser.

How to use it

  1. Enter one exact public HTTP(S) page URL.
  2. Optionally choose a before date to prefer an older capture.
  3. Run the comparison and verify the capture timestamp and crawl index.
  4. Review changed signals and excerpts. A difference is a review prompt, not an automatic SEO defect.

Example result

A bounded comparison can show a capture timestamp and crawl index alongside a changed title, canonical, or structured-data type. It reports only the selected capture and current raw HTML; a missing capture or a difference is evidence to review, not a complete version history or an automatic SEO defect.

How it works

The service selects one date-relevant Common Crawl index record, reads one bounded WARC byte range, and fetches the current public page through the guarded endpoint. It extracts comparable raw-HTML facts and capped text excerpts; it does not replay archived JavaScript or subresources.

What it checks

The result compares the title, description, H1, word count, canonical URLHow search engines pick one canonical URL among duplicates and consolidate signals onto it., robots directivesThe robots meta tag is an HTML element in a page's head — <meta name="robots" content="noindex"> — that tells search engines how to index and serve that page. It's crawl-then-obey: a page blocked in robots.txt is never fetched, so the tag is never seen., and JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. types. Similarity and excerpts use normalized extracted text from bounded raw HTML; they are not a rendered-page replay.

Common Crawl attribution and limitations

Historical capture data comes from Common Crawl; this site is not affiliated with it. Common Crawl samples the web, so captures can be absent, incomplete, truncated, or represented by revisit records without a payload. The tool queries only a few date-relevant crawl indexes and never promises lifetime coverage.

The comparison shows derived observations and small bounded excerpts—not a license to republish a third party’s archived page. Review the Common Crawl Terms of Use and any rights or terms that apply to the originating page before commercial reuse.

Frequently asked questions

What does the tool compare?

It compares one exact URL’s selected Common Crawl capture with the current public page: title, H1, meta description, canonical, robots directives, word count, structured-data types, bounded text similarity, and capped added/removed excerpts. Capture and current HTTP statuses are shown separately as source observations, not as a status-history comparison.

Does a missing capture mean the page never existed?

No. Common Crawl samples the web and does not capture every URL on every crawl. Missing, incomplete, revisit-only, and truncated records are reported as coverage limits rather than converted into a pass or failure.

Can I view or download the complete archived page?

No. The service retrieves one exact WARC byte range and returns only derived facts and bounded text excerpts. It does not provide a full-page archive viewer, execute archived JavaScript, or load archived subresources.

Can I use the result commercially?

This tool reports derived observations and does not grant rights in third-party archived content. Review Common Crawl’s current terms and the originating site’s rights before relying on archived material in a commercial product.

Next stepXML Sitemap Generator — generate the corrected version.

Feature requests for Historical Page Comparison

Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.

Loading…

➕ Request a feature

New requests are reviewed before they appear here.

Where this tool helps

Common use cases

Compare an archived page with today

Load one bounded Common Crawl capture and compare it with the current public page.

Investigate metadata drift

Review changes to titles, canonicals, robots directives, and structured-data types across the two observations.

Measure content change

Compare extracted text and word-count differences to understand the scope of a historical rewrite.

Document evidence for an SEO incident

Keep archive and current HTTP statuses separate while creating a dated record for further investigation.

Watch the full workflow

Historical Page Comparison walkthrough

Read the transcript

Historical Page Comparison

Common Crawl can provide a bounded historical checkpoint when you need to investigate how an exact page changed. I’ll show you how to select the U-R-L and optional date, verify the chosen capture, read signal and text differences, distinguish source observations from history claims, understand sampled coverage and rights limits, and combine the evidence with current diagnostics.

Step 1

Use this tool for migration audits, suspected metadata changes, content refresh analysis, and historical competitor research. It compares one exact Common Crawl capture with the current public page’s bounded raw-H-T-M-L signals and extracted text.

Step 2

Enter the full H-T-T-P-S page U-R-L—not a domain pattern. Optionally choose a before date to prefer an older capture. The request uses the exact U-R-L for bounded index lookup, one WARC byte range, and a guarded current-page fetch.

Step 3

The optional date asks the service to select a date-relevant capture on or before that boundary. It does not promise the nearest historical version, because Common Crawl samples the web and a usable record may not exist for every crawl.

Step 4

This walkthrough uses fictional example dot com observations so no live capture is implied. In a real run, select Compare historical versus current, complete the anti-abuse check when shown, and wait while the service looks up one record and fetches the current page.

Step 5

Start with the capture timestamp, Common Crawl index, capture H-T-T-P status, current final U-R-L and current status. Those statuses describe two separate source observations; they are not a redirect chain or complete historical status timeline.

Step 6

Compare title, H-one, description, canonical, robots directives, word count, and J-S-O-N-L-D types. Highlighting means the bounded extracted values differ. It does not mean the historical or current value is automatically better or defective.

Step 7

Extracted-text similarity summarizes normalized bounded text. Added and removed excerpts are capped samples that help locate likely edits. Navigation, templates, incomplete records, and extraction differences can affect them, so verify important passages elsewhere.

Step 8

A missing capture does not prove the page never existed. Records can be absent, incomplete, truncated, or revisit-only without a usable payload. Partial states and coverage notes remain limitations rather than being converted into a pass, failure, or complete history.

Step 9

Use a changed canonical, title, robots directive, schema type, or content excerpt to narrow an investigation. Then compare the capture window with C-M-S revisions, releases, Search Console, analytics, server logs, and stakeholder records before proposing cause or remediation.

Step 10

The service reads one bounded WARC byte range and current raw H-T-M-L. It does not execute archived JavaScript, load old styles, images, fonts, or other subresources, or provide a full archive viewer. Rendered experiences may have differed.

Step 11

Historical data comes from Common Crawl, and this site is not affiliated. The output contains derived observations and small excerpts, not permission to republish someone else’s archived content. Review Common Crawl’s current terms and the originating site’s rights before commercial reuse.

Step 12

Record the exact U-R-L, before date, selected timestamp, crawl index, status observations, changed fields, and coverage limits. Verify the current page separately, gather independent timing evidence, and form a narrow hypothesis. Make changes only when today’s evidence supports them.

Use one capture as a checkpoint—not a complete history.

Save the capture timestamp and crawl index, verify the changed field against current goals, and corroborate timing with deployment records, analytics, Search Console, logs, and other archives. Document uncertainty and rights constraints, then fix only current problems you can support with present-day evidence.