Crawl Export Comparator

Free, no signup. Compare two crawl exports without forcing unlike vendor issue taxonomies into fake equivalence. Shared page evidence is normalized; unmapped columns and incomplete crawl coverage remain visible.

Runs entirely in your browser — nothing you paste is uploaded or stored. Both files stay in your browser. Files over 15 MiB and rows beyond the 100,000-row comparison cap are rejected or reported as truncated. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.

Feedback
Report a bug

Found something broken in Crawl Export Comparator? Let us know what happened — this goes straight to a private triage queue, not a public list.

What will be sent
 No tool inputs, uploads, pasted source, complete results, query parameters, or URL fragments are attached automatically. You can edit or remove the selected passage above. Browser and anti-abuse metadata is processed for spam prevention. 

Sample report Static output-shape example

5 changed URLs · 2 added · 2 removed · 1 field-changed

/products/:number — 1 added, 1 removed; verify whether product IDs changed or the crawl scopes differed.

https://example.com/ — meta description and word count changed.

This illustrates the report structure. It is not a live observation or a claim that every difference is a defect.

How to use it

  1. Export page-level CSV or TSV files from the before and after crawls.
  2. Use comparable crawl scope, robots, rendering, authentication, and limits.
  3. Load both files and review the detected vendor, mapped fields, ignored columns, invalid URLs, duplicates, and caps.
  4. Triage inventory changes by path template, then inspect field changes per URL.
  5. Export the normalized change list and verify important regressions in the source crawlers or on the live pages.
Local data

Saved targets, named lists, and recent check summaries remain only in this browser.

How it works

The browser parses quoted CSV or TSV, detects known vendor headers, normalizes public HTTP(S) URLs, and keeps the last row when an export contains duplicate URLs. Only fields present in both exports are compared. Numeric evidence is compared numerically; canonicals and content types receive bounded normalization. The path-template rollup replaces only numeric IDs, UUIDs, and long hexadecimal segments.

What the results mean

  • Added — the normalized URL appears only in the after export.
  • Removed — it appears only in the before export.
  • Changed — at least one mutually mapped evidence field differs.
  • Ignored column — the source data is preserved in the original file but is not given cross-vendor meaning here.

Limitations

Different crawl configurations are the largest confounder. This tool does not reproduce rendering modes, authentication, robots behavior, custom extraction, vendor scoring, crawl scheduling, or site completeness. It does not infer that two differently named vendor checks mean the same thing. Template grouping is path-shape grouping, not a learned CMS template classification. The report is triage evidence, not proof of a regression.

Frequently asked questions

Which crawl exports can I compare?

The tool recognizes common page-export headers from Screaming Frog, Sitebulb, Scout Site Audit Free, and generic CSV or TSV files. It needs a URL column and compares only fields recognized in both files.

Are vendor issue names treated as equivalent?

No. Vendor-specific issue taxonomies can have different definitions and thresholds. The comparator normalizes shared page evidence such as status, indexability, title, canonicalHow search engines pick one canonical URL among duplicates and consolidate signals onto it., H1, word count, depth, and inlinks. Issue labels are compared as literal sorted text only when both files provide them.

Does an added or removed URL prove a crawlability change?

No. A URL can appear or disappear because of crawl scope, limits, authentication, robots rulesA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere., settings, or a real site change. Confirm comparable crawl configurationsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. before treating inventory differences as regressions.

Are the exports uploaded?

No. Parsing, normalization, comparison, filtering, and CSV export happen in your browser. The selected files are not sent to this site.

How are path templates created?

The grouping replaces numeric IDs, UUIDs, and long hexadecimal tokens in path segments. It is a deterministic triage aid, not a CMS-aware template detector.

Next stepXML Sitemap Generator — generate the corrected version.

Feature requests for Crawl Export Comparator

Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.

Loading…

➕ Request a feature

New requests are reviewed before they appear here.

Where this tool helps

Common use cases

Verify a migration or release with page-level evidence

Compare before-and-after crawl inventories and shared fields to find URLs and evidence that deserve verification after a site change.

Compare exports from different crawl tools cautiously

Normalize shared page evidence from Screaming Frog, Sitebulb, Scout, or generic exports without pretending vendor-specific issue taxonomies are equivalent.

Triage inventory changes by path shape

Group added and removed numeric-ID, UUID, and long-hash URLs into deterministic path templates so repeated patterns are easier to investigate.

Audit important on-page and crawl-field changes

Review status, indexability, titles, descriptions, canonicals, H1s, word counts, content types, depth, inlinks, and literal issue text when both files provide them.

Create a normalized review handoff

Filter the retained change evidence and export CSV for validation in the source crawlers or on live pages while preserving crawl-scope caveats.

Watch the full workflow

Crawl Export Comparator walkthrough

Read the transcript

Crawl Export Comparator

This beginner walkthrough defines the basic file and crawl terms, compares two fictional exports, explains every kind of difference, filters the review list, and downloads a reproducible C-S-V handoff.

Step 1

Crawl Export Comparator lines up two website-crawl files by U-R-L and reports what appears to have been added, removed, or changed. It recognizes common exports from Screaming Frog, Sitebulb, Scout Site Audit Free, and generic files. Everything is processed locally in your browser, without uploading the files to this site.

Step 2

A crawl is a tool’s attempt to visit and record website pages. A crawl export is the spreadsheet it saves. C-S-V means comma-separated values, and T-S-V uses tabs instead of commas. A field is one column, such as status code, title, canonical, word count, depth, or inlinks. Before and after simply mean the two files you choose; they do not have to come from the same vendor.

Step 3

Use the comparator after a migration or release, when comparing repeated crawls, when two teams use different crawl tools, when you want to review inventory by path pattern, or when you need a normalized handoff for another person. A reported difference is a question to verify, not automatic proof of a defect.

Step 4

Choose the earlier file under Before crawl and the comparison file under After crawl. This safe example selects a fictional three-row Screaming Frog C-S-V and a fictional three-row Sitebulb T-S-V. The files contain no real customer or company data. Files over fifteen mebibytes are rejected, and the comparison has a disclosed one-hundred-thousand-row cap.

Step 5

Both files include U-R-L, status, indexability, title, description, canonical, H one, word count, content type, crawl depth, and inlinks. The before file also has Custom Segment, while the after file has Project Label and Hints. The tool maps recognized evidence but never invents a relationship between columns it cannot compare honestly.

Step 6

Select Compare exports. An isolated browser worker reads both local files, detects the delimiter and likely vendor, normalizes public H-T-T-P or H-T-T-P-S U-R-Ls, keeps the last row if a file repeats a U-R-L, compares mutually mapped fields, and sends only the structured result back to the page.

Step 7

The example reports five changed U-R-Ls: two added, two removed, and one field-changed. Each source has three valid U-R-Ls. The comparable-evidence line names the fields present in both files. This summary describes the retained spreadsheet evidence; it does not tell us why a difference happened.

Step 8

The left card identifies Screaming Frog and maps Address to U-R-L, Title one to title, and Inlinks to inlinks. The right card identifies Sitebulb and maps U-R-L, Title, and Internal Inlinks to the same normalized fields. Both have three valid U-R-Ls, zero duplicate U-R-Ls, and zero invalid rows. Always confirm this mapping before trusting a cross-vendor comparison.

Step 9

Open Ignored columns on each card. Custom Segment remains ignored in the before file, and Project Label remains ignored in the after file. The original data is still in the source exports. Ignored means this report gives the column no cross-vendor meaning, which is safer than silently treating different project labels or vendor checks as equivalent.

Step 10

The first group replaces the numeric product identifiers with slash products slash colon number. It contains one added and one removed U-R-L. That can reveal an inventory pattern worth checking. The grouping only replaces numeric I-Ds, U-U-I-Ds, and long hexadecimal path segments. It does not understand your C-M-S or page templates.

Step 11

Choose field changes and one shared homepage remains. Its description changes from Old home to New home, word count from five hundred to five hundred fifty, and inlinks from twenty to twenty-two. These are observed normalized differences. The comparator cannot decide whether the newer values are correct, intentional, complete, or caused by different crawl settings.

Step 12

Choose added U-R-Ls. The list now shows slash added and product four-five-six. Inventory only means each U-R-L appears in the after file but not the before file, so there is no same-U-R-L field pair to compare. It does not prove that the page launched, became crawlable, or belongs in the index.

Step 13

Choose removed U-R-Ls and the list shows product one-two-three and slash removed. A U-R-L can disappear because of crawl scope, robots rules, authentication, rendering, limits, settings, or a real website change. Confirm that the two crawls were comparable before calling either row a loss or regression.

Step 14

Type slash products slash into the filter. The table keeps the removed product one-two-three and added product four-five-six rows. Search checks the full U-R-L, normalized path template, and changed field names. That makes it easy to create a focused review set without losing the distinct source U-R-Ls.

Step 15

With the product filter active, choose Download change C-S-V. The dated file contains one header and exactly the two visible product rows. Its columns are state, U-R-L, path template, field, before value, and after value. Preserve both original crawl exports and their settings with this handoff so another reviewer can reproduce the comparison.

Step 16

Added means the normalized U-R-L appears only after. Removed means it appears only before. Changed means a U-R-L exists in both and at least one mutually mapped field differs. Ignored means the source column is preserved but not interpreted across vendors. None of these labels means good, bad, correct, broken, indexable, or complete by itself.

Step 17

First, use comparable scope, robots behavior, rendering, authentication, and limits. Second, confirm the detected vendors, mappings, ignored columns, invalid rows, duplicates, and caps. Third, review repeated path patterns, then important U-R-Ls. Finally, verify high-impact differences in the original crawl tools or on the live pages before filing a bug or changing the site.

Step 18

The main features are local C-S-V and T-S-V parsing, common vendor-header recognition, public-U-R-L normalization, duplicate and invalid-row disclosure, mutually mapped field comparison, ignored-column visibility, added, removed, and changed states, deterministic path grouping, state and text filters, visible counts, and a complete filtered C-S-V export.

Step 19

Different crawl configurations are the largest confounder. This tool does not reproduce JavaScript rendering, authentication, robots handling, custom extraction, vendor scoring, schedules, issue definitions, or site completeness. It does not crawl the live website. Path grouping is shape-based, not learned C-M-S classification. The report is bounded triage evidence, not proof of a regression or its cause.

Treat every difference as a lead to verify.

Save both source exports and their crawl settings, investigate the most important U-R-Ls and path patterns, and confirm meaningful differences in the original crawler or on the live site. The free comparator narrows the work; it does not make the final decision.