Enterprise SEO Audit
How to audit a large enterprise site — segment before you crawl, start with indexing, think in templates, and ship 5-10 prioritized fixes people will actually implement.
An enterprise SEO audit isn't a bigger checklist — it's a different job. You can't audit everything, so you scope ruthlessly: interview stakeholders for real pain points, segment the site by CMS/region/template/team, then start with indexing (GSC Page Indexing report) before content or links. Think in templates, not pages — one bad canonical hits hundreds of thousands of URLs. Prioritize by business impact and feasibility, ship 5-10 fixes (not a 300-slide deck) in developer-ticket format, and wrap governance around it so the fixes don't regress. The hardest part is organizational, not technical.
TL;DR — An enterprise SEO audit is a health check for a very large website — think millions of pages across multiple systems and teams. The trick isn’t to check more things; it’s to not try to check everything. You figure out what actually hurts, look at the parts that matter, and hand back a short list of the biggest, most fixable problems — not a giant report nobody reads.
What an enterprise SEO audit is
An SEO audit is a review of a website to find what’s stopping it from showing up well in search. An enterprise audit is the same idea on a much bigger, messier site: a company with millions of URLs, several different content systems, teams in different countries, and a lot of moving parts.
That scale changes the job. On a small site you can look at every page. On an enterprise site you can’t — there are too many pages, and most of them follow the same handful of templates anyway. So instead of checking pages one at a time, you check the templates, because fixing one template fixes thousands of pages at once.
The order that matters
When everything looks broken, people freeze. Here’s the order I work in, simplest first:
- Can search engines reach the pages at all? This is the foundation. Start in Google Search Console’s Page Indexing report — it tells you which pages Google has indexed, and which it’s skipping (and often why).
- Once pages are findable, are they good? Content quality, duplicates, the right pages set to show vs. hide.
- Then links — internal links that help important pages, and recovering links pointing at broken pages.
Indexing first, content second, links third. Do it out of order and you’ll polish pages search engines can’t even see.
Why these audits usually fail
Not because someone missed a setting. They fail because the report is too big and nobody does anything with it. I’ve seen 300-slide audit decks gathering dust. A useful enterprise audit ends with 5 to 10 prioritized fixes written so an engineer can pick them up — plus a plain-English reason each one matters to the business (“this recovers links worth real traffic”), not just to SEO.
The part people skip
The hardest part of enterprise SEO isn’t technical — it’s getting different teams (engineering, content, legal) to agree on what to fix and actually do it. The best audit in the world is worthless if it sits in a folder. When a company and its people finally get behind SEO, they can dominate an industry.
Want the full methodology — how to scope, segment, prioritize, and keep an audit from going stale? Switch to the Advanced tab.
Google’s minimum technical requirements for indexing eligibility are Googlebot
access, an HTTP 200 response, and indexable content; eligibility still does not
guarantee indexing. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements
Google reserves its large-site crawl-budget guide mainly for sites above one million
unique pages or above 10,000 pages that change daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
TL;DR — Enterprise audits are different beasts because of scale and the org behind it. Don’t audit everything: interview stakeholders for real pain points, segment the site (CMS / region / template / team), then scope deliberately. Start with indexing — the GSC Page Indexing report — before content or links. Think in templates, not pages: one bad canonical hits hundreds of thousands of URLs. Crawl waste and hreflang are the uniquely-enterprise problems, and hreflang is the single most expensive issue you’ll find. Prioritize by impact × feasibility, frame findings as quantified business value in developer-ticket format, ship 5-10 items, and wrap governance around it so nothing regresses. Don’t trust tool scores — Google’s own people tell you not to.
Why “enterprise” changes the whole job
Enterprise SEO audits are entirely different beasts to “regular” audits because of all the complications that come from large sites — multiple systems, diverse teams, and data at a scale that breaks a normal spreadsheet. A standard audit applies a checklist to one site. An enterprise audit has to work across millions of URLs, many CMSs and CDNs, international footprints with separate regional teams, and JavaScript-heavy stacks — inside approval workflows where the person who has to ship the fix doesn’t report to you.
And the stakes are asymmetric. On a small site a mistake costs you a page. On an enterprise site, one mistake can keep millions of pages out of the index or remove an entire site from search results. A single URL-parameter setting once broke pagination indexing across an entire site I worked on. That’s the world you’re auditing in.
So the mindset shift is this: stop thinking in pages, start thinking in templates. One misconfigured canonical on a category-page template affects hundreds of thousands of URLs simultaneously. The wins (and the disasters) come from patterns, not individual pages. One mistake many SEOs make in an enterprise environment is getting caught up in busy work — fixing one page at a time, racking up micro gains while the macro problems sit untouched. Enterprise is all about scale.
The methodology: scope before you crawl
My enterprise audit process is four steps, and the first three all happen before you run a single crawler.
1. Find the pain points. Start with stakeholder interviews. What’s actually wrong — a traffic drop, a new market launch, a redesign, a compliance concern, a platform migration? The audit should answer specific questions, not produce a generic findings dump. This is also where you do stakeholder mapping: who owns what? Engineering owns the templates and the CDN. Content owns the copy. Legal can veto outreach and dictate disclaimer text. Knowing the owners up front is what makes the findings actionable later.
2. Segment the website. Break the site into manageable sections — by page type, region, language, or technical platform — using site-structure reports and custom filters. This is non-negotiable at scale: an unsegmented full-site crawl of a 40M-URL site produces data nobody can use, and a crawl that big can take 48-72 hours before you even start. Segmenting also lets you assign fixes to the right team later, since multiple websites, CMSs, and CDNs create infrastructure complexity that maps to different owners.
3. Determine scope. If you try to audit everything, an enterprise SEO audit is going to be expensive and time-consuming — and a lot of that time is likely to be wasted. Scope to what matters: the segments with traffic, revenue, or the known pain points. A focused technical audit of one segment might be ~10 hours; a full enterprise audit with a 12-month roadmap is 50-70.
4. Build deliverables that get implemented — covered in its own section below, because it’s where most audits die.
Start with indexability — always
Before content quality, before links, confirm the foundation: can Googlebot reach the
pages, do they return 200, and are they actually getting indexed? Google’s
technical requirements
are blunt — Googlebot must have access, the page must return an HTTP 200 status, and
it needs indexable content — and even then, eligibility isn’t a guarantee. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements Indexing
isn’t ranking.
The first diagnostic layer is the Google Search Console Page Indexing (Coverage) report. It shows what’s indexed and what’s being skipped, bucketed by reason — “Discovered – currently not indexed,” “Crawled – currently not indexed,” duplicates, soft 404s. On a large site those buckets are your map. Then you bring in crawl data (Botify, Lumar, Sitebulb, Screaming Frog) and, ideally, log files to see what bots actually fetched. GSC first, crawler second, logs for ground truth.
A warning while you read the Page Indexing report: don’t panic at a 404 spike. As Martin Splitt put it, a high number of 404s is expected if you removed a lot of content recently. The concern is an unexplained spike, not an expected one.
The uniquely-enterprise problems
Most audit findings exist on small sites too. A few are genuinely enterprise-only:
Crawl waste. At millions of URLs, parameter variants, faceted navigation, session
IDs, and thin internal-search results can eat the majority of your crawl budget while
burying high-value pages. This is the crawl-budget problem Google says matters only at
scale — its guide is aimed mainly at 1M+ unique pages or 10K+ pages changing daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
Below that, as Gary Illyes says, crawl budget is not something to worry about. The fix
is rarely “make Google crawl more” — it’s removing waste: consolidate duplicates,
robots.txt-block low-value spaces, return 404/410 for dead pages, and keep
sitemaps’ lastmod honest.
Duplicate content at scale. Templates generate near-identical pages by the
thousand. Analyze the pattern — don’t treat each as a random URL to canonicalize.
As John Mueller advises, instead of canonicalizing URLs one by one, find the pattern
and apply a tailored fix (e.g., block a parameter variant in robots.txt).
JavaScript rendering at scale. During the crawl Google renders the page and runs its JavaScript in a recent Chrome — but rendering is deferred and separate from fetching the HTML. Content that only exists after JS may be indexed late or not at all. Enterprise audits have to cross-reference the raw HTML response against the rendered output, at the template level.
Hreflang at scale — the most expensive issue you’ll find. On a large international site, misconfigured hreflang means Google serves the wrong country or language version across thousands of pages, quietly destroying conversion rates. I treat it as the most expensive issue in the audit because the loss compounds daily across every affected template. (Worth noting what isn’t worth your time here: minor locale variations like underscore-vs-dash rarely matter — don’t burn budget on them.)
Don’t trust the tool score
This deserves its own flag. Audit tools love to hand you a number — “SEO Health 82/100.” Ignore it. Martin Splitt, November 2025: “Please, please don’t follow your tools blindly. Make sure your findings are meaningful for the website in question.” Tool scores aren’t ranking factors and they lack site-specific context. A “low” score caused by correctly-noindexed faceted pages is fine. A 404 spike from a content purge is fine. Context determines severity — the tool can’t supply it. As Splitt frames it, finding technical issues is only half of an audit; understand the site’s technology before you run diagnostics, and validate findings with the people who know how the site is built.
This is also why I separate findings from recommendations. A findings dump is “here are 4,000 issues.” A recommendation is “here are the 8 that matter, in this order, and here’s why.” Executives need the second thing.
Prioritize: impact × feasibility, in dollars
No major site is technically perfect. To get there would be a waste of money. So prioritize on an impact/effort matrix and translate everything into business value. Crawl/index issues come first (highest immediate impact — they gate everything downstream), then on-page improvements at scale, then link work.
The framing is everything. You can say “redirect these 100 pages,” but it won’t go over as well as “redirect these 100 pages, which will recover 2,500 links.” I’ve gotten redirect projects funded by assigning a value of $400 per referring domain recovered — suddenly it’s a revenue conversation, not a tech chore. Translate SEO metrics into the language the budget-holder uses.
Some things SEOs over-prioritize, from auditing a lot of sites: short redirect chains (Google follows up to ~10 hops — only sweat it past 5), double slashes in URLs, multiple H1s (fine in HTML5), and the locale-format nitpicks above. Most common ≠ most important.
Deliverables that actually get implemented
This is where audits live or die. Nobody is going to read a 300-slide deck. The deliverable is a prioritized list of 5-10 main issues or opportunities, each with:
- a plain-language problem description,
- quantified business impact,
- implementation steps in developer-ticket format — detailed problem, acceptance criteria, reproducible steps, and the business impact attached,
- an effort estimate.
Keep the executive summary separate from the technical detail. Executives want the prioritized fix list and the expected outcome; engineers want the ticket. Don’t make either group read the other’s document.
Governance: stop the audit from going stale
Enterprise sites change every week — new templates, new teams, new CMS plugins — and issues you fix come back within months without guardrails. An audit should feed a continuous program, not be a one-off:
- SEO standards and SOPs, and pre-launch checklists (even pre-launch unit tests on staging) so regressions are caught before production.
- Crawl strategy by velocity: monthly/biweekly full crawls, daily sampling crawls across key templates for faster detection, and pre-launch staging audits. The emerging option is always-on crawling wired to IndexNow for real-time issue alerts.
- Cadence: full audits quarterly to biannually; continuous monitoring (Page Indexing report, Core Web Vitals, sampling crawls) always-on; trigger-based audits (migration, platform change, traffic drop) immediately.
The real lever
After all the technical detail, the thing that decides whether an enterprise SEO program wins is organizational, not technical. The key to enterprise SEO is doing the basics better than anyone else — and most major brands rank only for branded searches, leaving enormous non-branded visibility on the table. The blocker is almost never “we don’t know what’s wrong.” It’s getting the company aligned to fix it. When a company and its people finally get behind SEO, they can dominate an industry.
Fund an enterprise SEO audit to produce a short, owned fix plan—not a bigger checklist: scope the business problem, inspect patterns at template scale, and prioritize implementation.
- One template error can affect hundreds of thousands or millions of URLs, so page-by-page busywork misses the largest risks and wins.
- Indexing comes before content and links; the GSC Page Indexing report is the first diagnostic layer.
- Stakeholder mapping, developer-ticket deliverables, and ongoing guardrails determine whether findings reach production and stay fixed.
A segmented, impact-versus-feasibility audit concentrates engineering time on the site sections and template problems tied to traffic, revenue, or a known business concern.
Risk if ignored: A generic findings dump or tool score can consume budget without changing the site, while template-level indexation, rendering, international, and crawl issues continue to compound.
Ask your team: Which 5–10 fixes have the greatest business impact, who owns each ticket, and what pre-launch or monitoring control will prevent the issue from returning?
Google’s indexing-eligibility baseline is crawl access, an HTTP 200, and indexable
content. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements Its dedicated
crawl-budget guidance is aimed mainly at sites with more than one million unique pages
or more than 10,000 pages changing daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
AI summary
A condensed take on the Advanced version:
- Enterprise audits are a different job, not a bigger checklist — driven by scale (millions of URLs, multiple CMSs/CDNs, international teams, JS stacks) and by the org that has to ship the fixes.
- Scope before you crawl. My four-step process: (1) interview stakeholders for real pain points + map who owns what; (2) segment by CMS/region/template/team; (3) scope deliberately — auditing everything wastes time and money; (4) build deliverables that get implemented.
- Indexing first. Start in the GSC Page Indexing report, then crawl data, then
logs as ground truth. Google’s bar: accessible to Googlebot,
200status, indexable content — and eligibility still isn’t a guarantee. - Think in templates, not pages. One bad canonical hits hundreds of thousands of URLs. Micro-gains on single pages while macro problems sit is the classic mistake.
- Uniquely-enterprise problems: crawl waste (only matters past ~1M pages/week or 10K/day), duplicate content at scale, JS rendering at scale, and hreflang — the most expensive issue on large international sites.
- Don’t trust the tool score (Splitt: “don’t follow your tools blindly”); context decides severity. A 404 spike from a content purge is expected, not a fire.
- Prioritize impact × effort, in dollars (e.g. $400/referring domain recovered): crawl/index → on-page at scale → links. Separate findings from recommendations.
- Ship 5-10 fixes, in developer-ticket format with acceptance criteria + business impact — not a 300-slide deck. Separate the exec summary from the technical detail.
- Wrap governance around it (standards, pre-launch checks, sampling/always-on crawls) so fixes don’t regress. The real lever is organizational alignment.
Official documentation
Primary-source documentation that should ground an enterprise audit’s recommendations.
- Search Essentials — Technical requirements — the three minimums for indexing eligibility: Googlebot access, HTTP
200, indexable content. - Optimize your crawl budget — crawl capacity + demand, who actually needs it (1M+ pages/week, 10K+/day), and the URL-inventory recommendations.
- Core Web Vitals — LCP, INP (replaced FID in March 2024), and CLS, measured at the 75th percentile of CrUX field data.
- JavaScript SEO basics — rendering is deferred and separate; cross-reference HTML against rendered output.
- Do I need an SEO? — Google’s own description of what a proper audit should deliver (realistic estimates, expected outcomes — and never a rankings guarantee).
- Page Indexing report (Search Console Help) — the first diagnostic layer for indexability at scale.
Bing / Microsoft
- Bing Webmaster Guidelines — content quality and technical requirements.
- Keeping content discoverable with sitemaps in AI-powered search (July 2025) — enterprise sitemap limits (50,000 URLs/file; index files referencing billions of URLs) and why accurate
lastmodmatters. - IndexNow — real-time URL submission for high-velocity enterprise sites; the basis for always-on crawl alerting.
Quotes from the source
On-the-record statements from Google reps relevant to auditing large sites.
Martin Splitt, Google — what an audit is actually for
- “A technical audit, in my opinion, should make sure no technical issues prevent or interfere with crawling or indexing.” — Martin Splitt, Search Central, November 2025. Coverage
- “Please, please don’t follow your tools blindly. Make sure your findings are meaningful for the website in question.” — Martin Splitt, November 2025. Coverage
- “A high number of 404s, for instance, is expected if you removed a lot of content recently.” — Martin Splitt, November 2025. Coverage
Gary Illyes, Google — crawl budget and quality
- “For most sites, crawl budget is not something to worry about. For really large sites, it becomes something to consider looking at.” — Gary Illyes. Crawl budget docs
John Mueller, Google — indexing large sites
- “I strongly recommend not relying on trying to force indexing” for large sites — manual indexing requests don’t scale. — John Mueller, Google Search Central.
- “Internal links help Google understand site structure, and if all pages are linked to all other pages, there’s no real structure.” — John Mueller, Google Search Central.
Enterprise audit checklist
Run it in this order — scope first, then foundation, then everything else.
Before any crawler runs
- Stakeholder interviews done — known pain points captured (traffic drop, launch, migration, redesign, compliance).
- Stakeholder map drawn — who owns templates, content, CDN, legal sign-off.
- Site segmented by CMS / region / page type / template / team.
- Scope agreed — which segments are in, which are explicitly out, and why.
Indexability (the foundation)
- GSC Page Indexing report reviewed; “Discovered/Crawled – currently not indexed” buckets analyzed by pattern.
- Key templates return
200; dead pages return404/410. - Crawl waste identified (parameters, facets, session IDs, internal search) and a
robots.txt/consolidation plan drafted. - Sitemaps list only canonical, indexable URLs with honest
lastmod. - JS-rendered content cross-referenced: raw HTML vs. rendered output at the template level.
Content & international
- Duplicate/thin content assessed by template pattern, not page-by-page.
- hreflang validated at scale (return tags, correct region/language) on every international template — highest-stakes check.
Links
- 404s with backlinks flagged for reclamation (with $ value attached).
- Orphaned and deeply-buried important pages (>3 clicks) identified.
Deliverable & governance
- Findings distilled to 5-10 prioritized recommendations.
- Each fix written as a developer ticket (problem, acceptance criteria, steps, business impact).
- Executive summary separated from technical detail.
- Monitoring cadence + pre-launch checklist agreed so fixes don’t regress.
The mental models
1. Scope before you crawl — the four steps. Pain points → segment → scope → deliverables. The first three happen before any tool runs. Skipping them is how you end up with a 40M-row crawl nobody can use.
2. Templates, not pages. On an enterprise site, every meaningful problem and fix is a pattern. Ask “which template generates this?” before “which page is broken?” One canonical fix can move hundreds of thousands of URLs; one page fix moves one.
3. The indexability ladder — climb in order.
Accessible to Googlebot → returns 200 → indexable content → actually indexed →
ranks. Find the rung a segment is failing on (start in the Page Indexing report)
before changing anything above it.
4. Impact × effort, priced in dollars. Plot findings on an impact/effort matrix, then convert impact to money ($/referring-domain, $/traffic recovered). Order of operations: crawl/index issues → on-page at scale → links. Findings ≠ recommendations — executives get the short, prioritized recommendation list.
5. Context beats the score. A tool’s number has no idea what your site is supposed to do. A 404 spike after a purge, a “low score” from correctly-noindexed facets — both fine. Validate every finding against how the site is actually built.
6. Audit → governance loop. An audit is a snapshot; an enterprise site is a movie. Feed findings into standards, pre-launch checks, and continuous/sampling crawls so the same issues don’t return next quarter.
Enterprise audit cheat sheet
Order of operations
| Phase | What you do | Primary source |
|---|---|---|
| 0. Scope | Pain points → segment → scope | Stakeholder interviews, site-structure reports |
| 1. Index | Can engines reach + index it? | GSC Page Indexing report → crawl data → logs |
| 2. Content | Duplicates/thin at template level | Crawl + content reports |
| 3. International | hreflang validation at scale | Crawl + GSC international targeting |
| 4. Links | 404 reclamation, orphans, depth | Backlink + internal-link reports |
| 5. Deliver | 5-10 fixes as dev tickets, $-framed | — |
| 6. Govern | Standards, pre-launch checks, monitoring | Sampling/always-on crawls, IndexNow |
Who needs to worry about crawl budget (Google’s threshold)
- 1M+ unique pages updating weekly, or 10K+ pages updating daily, or lots of “Discovered – currently not indexed.” Below that: don’t worry about it.
Core Web Vitals targets
- LCP < 2.5s · INP < 200ms (replaced FID, March 2024) · CLS < 0.1 — at the 75th percentile of CrUX field data.
Over-prioritized (lower than SEOs think)
- Redirect chains (Google follows ~10 hops; worry past 5) · double slashes · multiple H1s (fine in HTML5) · locale format nitpicks (underscore vs. dash).
Highest-stakes
- hreflang on international templates — the most expensive issue · canonical/template errors that ripple across hundreds of thousands of URLs · anything that can deindex at scale.
Deliverable rules
- 5-10 items, not 300 slides · dev-ticket format (problem + acceptance criteria + steps + business impact) · exec summary separate from technical detail · findings ≠ recommendations.
Patrick's relevant free tools
- SEO Incident Simulator — Practice thirty deterministic technical SEO incident investigations — indexability, crawl controls, redirects, sitemaps, markup, caching, DNS, bot verification, rendering, hreflang, and faceted navigation — with clearly labeled fixture evidence and Find → Fix → Verify handoffs.
- hreflang Generator + Linter — Enter your URL × locale matrix and get bidirectional hreflang markup as head tags, sitemap XML (auto-split past 50,000 URLs), and Link headers — linted live for wrong region codes, duplicates, and missing fallbacks. Runs entirely in your browser.
- returntag - hreflang checker — Enter one URL or an XML sitemap and the returntag - hreflang checker crawls the whole hreflang cluster — missing return tags, broken targets, self-reference and x-default checks, language-code validation, and head vs Link header vs sitemap disagreements — on an interactive cluster map with CSV export. The check Search Console's International Targeting report used to run.
Tools for enterprise SEO audits
No single tool covers an enterprise audit. The stack:
- Google Search Console (free, official) — start here. Page Indexing report, Core Web Vitals, Crawl Stats, URL Inspection. The first diagnostic layer.
- Bing Webmaster Tools (free, official) — Site Scan (on-demand full-site audit), Top Insights (pages missing from sitemaps, crawl errors), IndexNow Insights, Crawl Control, and 16-month Search Performance.
- Enterprise crawlers — Botify, Lumar (DeepCrawl), Sitebulb, and Screaming Frog for crawling at scale, segmented by template. Crawls of millions of URLs can take 48-72 hours, which is exactly why you segment first.
- Log file analysis — the ground truth for what bots actually fetched and where crawl budget is being wasted (Screaming Frog Log File Analyser, or pipe logs into BigQuery / a log platform).
- Backlink analysis — Ahrefs for the link audit and 404 reclamation (attaching $ value to recovered referring domains).
- Enterprise SEO platforms — BrightEdge, Conductor, seoClarity, Ahrefs Enterprise for ongoing monitoring across large sites.
- Performance — PageSpeed Insights + CrUX for field Core Web Vitals.
A reminder that overrides all of the above: don’t let any tool’s overall “health score” set your priorities. It’s not a ranking factor and it doesn’t know your site.
Quarterly enterprise SEO audit SOP
Use this as a recurring audit cycle, not a one-time findings dump.
- Confirm scope and owners. Re-interview the owners of each CMS, region, and template. Record major launches, migrations, traffic changes, and any systems that moved in or out of scope. Done means every included segment has a named owner and a business question the audit must answer.
- Compare expected and indexed inventory. Export sitemap totals and review the GSC Page Indexing report by segment. Investigate material gaps by reason before running a full crawl. Done means each gap is classified as intentional or queued for diagnosis.
- Sample every important template. Check status, canonical, robots directives, rendered content, internal links, and hreflang where applicable. Done means every revenue-driving template has a current pass/fail record.
- Crawl only the segments that need evidence. Apply saved include/exclude rules so parameters and low-value spaces do not bury the findings. Done means every crawl maps back to an owner, template, and audit question.
- Prioritize the short list. Score confirmed findings by business impact and feasibility, then reduce the deliverable to 5–10 recommendations. Done means each recommendation has an affected pattern, owner, impact case, and effort estimate.
- Write implementation-ready tickets. Include reproduction steps, acceptance criteria, and the check that proves the fix worked. Done means engineering can estimate the ticket without reopening the audit deck.
- Close the loop. Add shipped fixes to monitoring and pre-launch checks, and carry unresolved items into the next cycle with a reason. Done means the next audit starts from a change log rather than from zero.
Enterprise audit mistakes that waste the most time
Crawling the whole site before defining the question
Why it fails: a 40-million-URL export mixes unrelated systems, templates, and owners into one pile. The volume looks impressive but makes patterns harder to see.
Do instead: interview stakeholders, segment the site, and decide which business questions each crawl must answer before it runs.
Treating every URL as a separate problem
Why it fails: fixing individual pages leaves the template that generated the error untouched, so the issue returns across thousands of URLs.
Do instead: identify the template, CMS rule, or parameter pattern behind each finding and write the recommendation at that level.
Ranking findings by a tool’s health score
Why it fails: the score cannot know that a noindexed facet is intentional or a 404 spike followed a planned content purge.
Do instead: validate findings against site behavior and business impact, then rank them by impact and feasibility.
Delivering hundreds of unowned findings
Why it fails: a large deck shifts the sorting work to stakeholders and leaves no clear first move.
Do instead: deliver 5–10 recommendations with an owner, acceptance criteria, effort, and a plain-language impact case.
Ending the audit when the report is sent
Why it fails: templates and platforms keep changing, so fixed issues quietly return.
Do instead: turn each shipped fix into a regression check, monitoring rule, or pre-launch requirement.
Prompts for enterprise audit work
Turn segmented crawl data into template-level findings
Paste a CSV excerpt with URL, template, status, canonical, robots, depth, and index state. Remove sensitive fields first. Expect a grouped analysis, not page-by-page advice.
You are helping triage an enterprise SEO crawl. Group the rows below by template and
failure pattern. For each pattern, report: affected segment, observable evidence,
likely system-level cause, pages affected in this sample, business risk, owner to
involve, and the next check needed to confirm the diagnosis. Do not infer revenue or
claim causation from correlation. Separate intentional states from probable defects.
[PASTE SEGMENTED CRAWL ROWS]Convert a confirmed finding into a developer ticket
Paste one verified finding plus the relevant template behavior. Expect a ticket an engineer can estimate without reading the full audit.
Convert this confirmed enterprise SEO finding into a developer-ready ticket. Return:
problem statement, affected template or rule, reproduction steps, expected behavior,
acceptance criteria, validation test, rollback trigger, dependencies, owner, and a
plain-language business impact statement. Preserve unknowns as questions. Do not
invent traffic, revenue, or implementation estimates.
[PASTE VERIFIED FINDING AND EVIDENCE] Resources worth your time
My related writing
- What is an Enterprise SEO Audit & How To Do One — the full 4-step process and audit types.
- Enterprise Sites Are Where Technical SEO Shines — crawl strategy options, the impact/effort matrix, and high-priority projects.
- Enterprise SEO Challenges & Mistakes You Need To Overcome — the organizational layer: buy-in, incentives, and micro-vs-macro.
- Enterprise SEO Strategies For Maximum Growth — the broader program around the audit.
My speaking
- Enterprise SEO Chaos (SMX Seattle 2016) — what enterprise scale actually looks like, from 4 years in-house at IBM (378,000 employees, 170+ countries): 24 versions of one URL, 14-hop redirect chains, and “Everything Has To Work Together.”
- What I Learned from Auditing Over 1,000,000 Websites — why most common ≠ most important, and the prioritization thresholds that actually matter.
From around the industry
- Google — Search Central technical-audit guidance — Martin Splitt on tool scores and audit methodology (Nov 2025); Search Engine Journal coverage.
- Martin Splitt — “Why we need to talk about audits” (YouTube, Nov 2025) — primary source video on audit purpose and tool-score misconceptions, from Google Search Central.
- Martin Splitt — “How to perform a technical SEO audit” (YouTube, Nov 2025) — companion video covering the three-step audit framework from Google.
- Screaming Frog: How to Do an Enterprise SEO Audit the Right Way — tool-centric, four-stage framework from one of the most-used enterprise crawlers.
- Sitebulb: Key Considerations in Enterprise SEO Auditing — crawl-first methodology guide covering segmentation and prioritization at scale.
- Search Engine Land: What your enterprise SEO audit may be missing — covers canonical, indexing gaps, and holistic audit blind spots.
- Search Engine Land: 6 reasons why SEO audits seem like a waste and how to fix them — practical fixes for audit deliverables that don’t get implemented.
- r/TechSEO — the community for large-site crawl/index debugging.
Test yourself: Enterprise SEO audits
Five questions on scoping, diagnosis, and delivery. Pick an answer for each, then check.
Stats worth citing
- Crawl-budget threshold: Google’s own bar for when crawl budget matters — 1M+ pages updating weekly or 10K+ pages updating daily. Below that, it’s not worth worrying about. Source
- Enterprise sitemap scale (Bing): a single sitemap index file can reference up to
2.5 billion URLs (50,000 child sitemaps × 50,000 URLs) — and multiple index files
push that far higher. The scale is the point: at this size, accurate
lastmodis what keeps discovery efficient. Source - Buy-in by the numbers: I’ve funded redirect projects by assigning $400 per referring domain recovered — turning a tech chore into a revenue case is what gets enterprise fixes shipped. Source
Enterprise SEO Audit
An enterprise SEO audit is a systematic evaluation of a large-scale website — usually millions of URLs across multiple CMS platforms, teams, and regions — to find the technical, indexing, content, and link issues that limit organic visibility. Unlike a standard audit, it must account for organizational complexity and ruthless prioritization, because technical perfection at scale is neither possible nor worth paying for.
Related: Crawl Budget, Hreflang, Indexing, Core Web Vitals, JavaScript SEO
Enterprise SEO Audit
An enterprise SEO audit evaluates a large enterprise website to find the issues and opportunities that will improve its rankings and visibility in search. What makes it “enterprise” isn’t a fancier checklist — it’s scale and the organization behind it. You’re dealing with millions of URLs, multiple content management systems and CDNs, international footprints across regional teams, and JavaScript-heavy stacks, all owned by different groups with their own roadmaps and incentives. As I’ve said before, enterprise SEO audits are entirely different beasts to “regular” audits because of all the complications that come from large sites.
That changes the work in two big ways. First, you can’t audit everything — an unsegmented full-site crawl of a 40-million-URL site produces unusable data. So enterprise audits start by finding stakeholder pain points, segmenting the site (by CMS, region, page type, and team ownership), and scoping deliberately before any tool runs. Second, you think in templates, not pages: one misconfigured canonical on a category template can affect hundreds of thousands of URLs at once, so the fixes that matter are the ones that ripple across a pattern.
The other half of the job is organizational. The #1 reason enterprise audits fail isn’t a missed finding — it’s that nobody acts on the deck. The deliverable can’t be a 300-slide dump; it has to be a prioritized list of the 5–10 highest-impact, feasible fixes, framed in business terms (“recover 2,500 links worth $X”) and handed over in developer-ticket format with acceptance criteria. And because enterprise sites change constantly, an audit feeds into ongoing governance — standards, pre-launch checks, and monitoring — so the issues you fix don’t quietly come back.
Related: Crawl Budget, Hreflang, Indexing, Core Web Vitals, JavaScript SEO
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 25, 2026.
Editorial summary and recorded change details.Summary
Aligned legacy tool display names with returntag and Scout Site Audit Free.
Change details
-
Updated the linked tool names to match their current public labels.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 19, 2026.
Editorial summary and recorded change details.Summary
Removed an unsourced Core Web Vitals origin pass-rate statistic ("Jan 2026 CrUX": ~68% LCP / ~87% INP / ~81% CLS / ~56% pass-all-three) from the Stats lens — it carried no first-party citation, and the aggregate-dashboard tool that kind of figure is normally pulled from (Google's CrUX Dashboard in Looker Studio) is now deprecated. Reviewed the article against a Wave 94 structured-research packet (screened synthetic/templated, used for topic-coverage only); verified the crawl-budget doc citation still resolves and its 1M-pages-weekly/10K-pages-daily threshold text is unchanged, and confirmed the quoted Martin Splitt lines against Search Engine Journal's and ppc.land's coverage.
Change details
-
Removed the unsourced ~68%/~87%/~81%/~56% Core Web Vitals pass-rate stat from the Stats lens; no defensible dated first-party source for it.
Full comparison unavailable — no prior snapshot was archived for this revision.