Enterprise SEO Audit
How to audit a large enterprise site — segment before you crawl, start with indexing, think in templates, and ship 5-10 prioritized fixes people will actually implement.
An enterprise SEO audit isn't a bigger checklist — it's a different job. You can't audit everything, so you scope ruthlessly: interview stakeholders for real pain points, segment the site by CMS/region/template/team, then start with indexing (GSC Page Indexing report) before content or links. Think in templates, not pages — one bad canonical hits hundreds of thousands of URLs. Prioritize by business impact and feasibility, ship 5-10 fixes (not a 300-slide deck) in developer-ticket format, and wrap governance around it so the fixes don't regress. The hardest part is organizational, not technical.
TL;DR — An enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. audit is a health check for a very large website — think millions of pages across multiple systems and teams. The trick isn’t to check more things; it’s to not try to check everything. You figure out what actually hurts, look at the parts that matter, and hand back a short list of the biggest, most fixable problems — not a giant report nobody reads.
What an enterprise SEO audit is
An SEO auditAn SEO audit checklist is the structured set of things you review across a site's technical health, on-page elements, content quality, and off-page authority — used to produce a short, prioritized action plan, not an exhaustive 100–200 item inventory. is a review of a website to find what’s stopping it from showing up well in search. An enterprise audit is the same idea on a much bigger, messier site: a company with millions of URLs, several different content systems, teams in different countries, and a lot of moving parts.
That scale changes the job. On a small site you can look at every page. On an enterprise site you can’t — there are too many pages, and most of them follow the same handful of templates anyway. So instead of checking pages one at a time, you check the templates, because fixing one template fixes thousands of pages at once.
The order that matters
When everything looks broken, people freeze. Here’s the order I work in, simplest first:
- Can search engines reach the pages at all? This is the foundation. Start in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.’s Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. — it tells you which pages Google has indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and which it’s skipping (and often why).
- Once pages are findable, are they good? Content quality, duplicates, the right pages set to show vs. hide.
- Then links — internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. that help important pages, and recovering links pointing at broken pages.
IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. first, content second, links third. Do it out of order and you’ll polish pages search engines can’t even see.
Why these audits usually fail
Not because someone missed a setting. They fail because the report is too big and nobody does anything with it. I’ve seen 300-slide audit decks gathering dust. A useful enterprise audit ends with 5 to 10 prioritized fixes written so an engineer can pick them up — plus a plain-English reason each one matters to the business (“this recovers links worth real traffic”), not just to SEO.
The part people skip
The hardest part of enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. isn’t technical — it’s getting different teams (engineering, content, legal) to agree on what to fix and actually do it. The best audit in the world is worthless if it sits in a folder. When a company and its people finally get behind SEO, they can dominate an industry.
Want the full methodology — how to scope, segment, prioritize, and keep an audit from going stale? Switch to the Advanced tab.
Google’s minimum technical requirements for indexing eligibility are GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer.
access, an HTTP 200 response, and indexable content; eligibility still does not
guarantee indexing. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements
Google reserves its large-site crawl-budget guide mainly for sites above one million
unique pages or above 10,000 pages that change daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
TL;DR — Enterprise audits are different beasts because of scale and the org behind it. Don’t audit everything: interview stakeholders for real pain points, segment the site (CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. / region / template / team), then scope deliberately. Start with indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — the GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. — before content or links. Think in templates, not pages: one bad canonical hits hundreds of thousands of URLs. Crawl waste and hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. are the uniquely-enterprise problems, and hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. is the single most expensive issue you’ll find. Prioritize by impact × feasibility, frame findings as quantified business value in developer-ticket format, ship 5-10 items, and wrap governance around it so nothing regresses. Don’t trust tool scores — Google’s own people tell you not to.
Why “enterprise” changes the whole job
Enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. audits are entirely different beasts to “regular” audits because of all the complications that come from large sites — multiple systems, diverse teams, and data at a scale that breaks a normal spreadsheet. A standard audit applies a checklist to one site. An enterprise audit has to work across millions of URLs, many CMSs and CDNs, international footprints with separate regional teams, and JavaScript-heavy stacks — inside approval workflows where the person who has to ship the fix doesn’t report to you.
And the stakes are asymmetric. On a small site a mistake costs you a page. On an enterprise site, one mistake can keep millions of pages out of the index or remove an entire site from search results. A single URL-parameter setting once broke paginationPagination splits a large set of content — product listings, blog archives, search results — across multiple sequentially numbered URLs. For SEO, each paginated page should be crawlable, indexable, and self-canonical; Google no longer uses rel=prev/next, but Bing still does. indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. across an entire site I worked on. That’s the world you’re auditing in.
So the mindset shift is this: stop thinking in pages, start thinking in templates. One misconfigured canonical on a category-page template affects hundreds of thousands of URLs simultaneously. The wins (and the disasters) come from patterns, not individual pages. One mistake many SEOs make in an enterprise environment is getting caught up in busy work — fixing one page at a time, racking up micro gains while the macro problems sit untouched. Enterprise is all about scale.
The methodology: scope before you crawl
My enterprise audit process is four steps, and the first three all happen before you run a single crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
1. Find the pain points. Start with stakeholder interviews. What’s actually wrong — a traffic drop, a new market launch, a redesign, a compliance concern, a platform migration? The audit should answer specific questions, not produce a generic findings dump. This is also where you do stakeholder mapping: who owns what? Engineering owns the templates and the CDN. Content owns the copy. Legal can veto outreach and dictate disclaimer text. Knowing the owners up front is what makes the findings actionable later.
2. Segment the website. Break the site into manageable sections — by page type, region, language, or technical platform — using site-structure reports and custom filters. This is non-negotiable at scale: an unsegmented full-site crawl of a 40M-URL site produces data nobody can use, and a crawl that big can take 48-72 hours before you even start. Segmenting also lets you assign fixes to the right team later, since multiple websites, CMSs, and CDNs create infrastructure complexity that maps to different owners.
3. Determine scope. If you try to audit everything, an enterprise SEO auditAn SEO audit checklist is the structured set of things you review across a site's technical health, on-page elements, content quality, and off-page authority — used to produce a short, prioritized action plan, not an exhaustive 100–200 item inventory. is going to be expensive and time-consuming — and a lot of that time is likely to be wasted. Scope to what matters: the segments with traffic, revenue, or the known pain points. A focused technical audit of one segment might be ~10 hours; a full enterprise audit with a 12-month roadmap is 50-70.
4. Build deliverables that get implemented — covered in its own section below, because it’s where most audits die.
A bounded raw-HTML crawl is one audit input, not the audit score. Rendered JavaScript, logs, Search Console, business impact, and governance stay separate.
Run a scoped technical baseline with my free Scout Site Audit Free Free
- Define representative templates, hosts, markets, and crawl caps before running the sample.
- Record discovered versus crawled coverage and group repeatable findings by template owner.
- Reconcile priority patterns with logs, Search Console, rendered pages, and business impact before recommending work.
The sample report scores 82 out of 100, crawls 24 of 31 discovered URLs, labels the crawl partial, reports two missing-title warnings across 24 checked pages, and says rendered JavaScript and AI checks were not evaluated.
Start with indexability — always
Before content quality, before links, confirm the foundation: can GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. reach the
pages, do they return 200, and are they actually getting indexed? Google’s
technical requirements
are blunt — Googlebot must have access, the page must return an HTTP 200 status, and
it needs indexable content — and even then, eligibility isn’t a guarantee. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements Indexing
isn’t ranking.
The first diagnostic layer is the Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing (Coverage) report. It shows what’s indexed and what’s being skipped, bucketed by reason — “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.,” “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.,” duplicates, soft 404sA soft 404 is a URL that returns a success status code (usually 200 OK) even though the page is empty, missing, or shows a 'not found' message. It isn't a status code a server sends — it's a label search engines apply after comparing the response code against the rendered content, and they treat the page like a 404 for indexing.. On a large site those buckets are your map. Then you bring in crawl data (Botify, Lumar, Sitebulb, Screaming Frog) and, ideally, log files to see what bots actually fetched. GSC first, crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. second, logs for ground truth.
A warning while you read the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.: don’t panic at a 404 spike. As Martin Splitt put it, a high number of 404s is expected if you removed a lot of content recently. The concern is an unexplained spike, not an expected one.
The uniquely-enterprise problems
Most audit findings exist on small sites too. A few are genuinely enterprise-only:
Crawl waste. At millions of URLs, parameter variants, faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals., session
IDs, and thin internal-search results can eat the majority of your crawl budgetThe number of URLs an engine will crawl in a timeframe. while
burying high-value pages. This is the crawl-budget problem Google says matters only at
scale — its guide is aimed mainly at 1M+ unique pages or 10K+ pages changing daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
Below that, as Gary Illyes says, crawl budget is not something to worry about. The fix
is rarely “make Google crawl more” — it’s removing waste: consolidate duplicates,
robots.txt-block low-value spaces, return 404/410 for dead pages, and keep
sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.’ lastmod honest.
Duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. at scale. Templates generate near-identical pages by the
thousand. Analyze the pattern — don’t treat each as a random URL to canonicalize.
As John Mueller advises, instead of canonicalizing URLs one by one, find the pattern
and apply a tailored fix (e.g., block a parameter variant in robots.txt).
JavaScript renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. at scale. During the crawl Google renders the page and runs its JavaScript in a recent Chrome — but rendering is deferred and separate from fetching the HTML. Content that only exists after JS may be indexed late or not at all. Enterprise audits have to cross-reference the raw HTML response against the rendered output, at the template level.
Hreflang at scale — the most expensive issue you’ll find. On a large international site, misconfigured hreflang means Google serves the wrong country or language version across thousands of pages, quietly destroying conversion rates. I treat it as the most expensive issue in the audit because the loss compounds daily across every affected template. (Worth noting what isn’t worth your time here: minor locale variations like underscore-vs-dash rarely matter — don’t burn budget on them.)
Don’t trust the tool score
This deserves its own flag. Audit tools love to hand you a number — “SEO Health 82/100.” Ignore it. Martin Splitt, November 2025: “Please, please don’t follow your tools blindly. Make sure your findings are meaningful for the website in question.” Tool scores aren’t ranking factors and they lack site-specific context. A “low” score caused by correctly-noindexed faceted pages is fine. A 404 spike from a content purge is fine. Context determines severity — the tool can’t supply it. As Splitt frames it, finding technical issues is only half of an audit; understand the site’s technology before you run diagnostics, and validate findings with the people who know how the site is built.
This is also why I separate findings from recommendations. A findings dump is “here are 4,000 issues.” A recommendation is “here are the 8 that matter, in this order, and here’s why.” Executives need the second thing.
Prioritize: impact × feasibility, in dollars
No major site is technically perfect. To get there would be a waste of money. So prioritize on an impact/effort matrix and translate everything into business value. Crawl/index issues come first (highest immediate impact — they gate everything downstream), then on-page improvements at scale, then link work.
The framing is everything. You can say “redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. these 100 pages,” but it won’t go over as well as “redirect these 100 pages, which will recover 2,500 links.” I’ve gotten redirect projects funded by assigning a value of $400 per referring domain recovered — suddenly it’s a revenue conversation, not a tech chore. Translate SEO metrics into the language the budget-holder uses.
Some things SEOs over-prioritize, from auditing a lot of sites: short redirect chains (Google follows up to ~10 hops — only sweat it past 5), double slashes in URLs, multiple H1s (fine in HTML5), and the locale-format nitpicks above. Most common ≠ most important.
Deliverables that actually get implemented
This is where audits live or die. Nobody is going to read a 300-slide deck. The deliverable is a prioritized list of 5-10 main issues or opportunities, each with:
- a plain-language problem description,
- quantified business impact,
- implementation steps in developer-ticket format — detailed problem, acceptance criteria, reproducible steps, and the business impact attached,
- an effort estimate.
Keep the executive summary separate from the technical detail. Executives want the prioritized fix list and the expected outcome; engineers want the ticket. Don’t make either group read the other’s document.
Governance: stop the audit from going stale
Enterprise sites change every week — new templates, new teams, new CMS plugins — and issues you fix come back within months without guardrails. An audit should feed a continuous program, not be a one-off:
- SEO standards and SOPs, and pre-launch checklists (even pre-launch unit tests on staging) so regressions are caught before production.
- Crawl strategy by velocity: monthly/biweekly full crawls, daily sampling crawls across key templates for faster detection, and pre-launch staging audits. The emerging option is always-on crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. wired to IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. for real-time issue alerts.
- Cadence: full audits quarterly to biannually; continuous monitoring (Page Indexing report, Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data., sampling crawls) always-on; trigger-based audits (migration, platform change, traffic drop) immediately.
The real lever
After all the technical detail, the thing that decides whether an enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. program wins is organizational, not technical. The key to enterprise SEO is doing the basics better than anyone else — and most major brands rank only for branded searches, leaving enormous non-branded visibility on the table. The blocker is almost never “we don’t know what’s wrong.” It’s getting the company aligned to fix it. When a company and its people finally get behind SEO, they can dominate an industry.
Fund an enterprise SEO audit to produce a short, owned fix plan—not a bigger checklist: scope the business problem, inspect patterns at template scale, and prioritize implementation.
- One template error can affect hundreds of thousands or millions of URLs, so page-by-page busywork misses the largest risks and wins.
- Indexing comes before content and links; the GSC Page Indexing report is the first diagnostic layer.
- Stakeholder mapping, developer-ticket deliverables, and ongoing guardrails determine whether findings reach production and stay fixed.
A segmented, impact-versus-feasibility audit concentrates engineering time on the site sections and template problems tied to traffic, revenue, or a known business concern.
Risk if ignored: A generic findings dump or tool score can consume budget without changing the site, while template-level indexation, rendering, international, and crawl issues continue to compound.
Ask your team: Which 5–10 fixes have the greatest business impact, who owns each ticket, and what pre-launch or monitoring control will prevent the issue from returning?
Google’s indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.-eligibility baseline is crawl access, an HTTP 200, and indexable
content. Evidence for this claim Google lists three minimum technical requirements for indexing eligibility: Googlebot must not be blocked, the page must return HTTP 200, and the page must contain indexable content. Scope: Minimum eligibility requirements for Google Search; meeting them does not guarantee crawling, indexing, serving, or ranking. Confidence: high · Verified: Google Search Central: Technical requirements Its dedicated
crawl-budget guidance is aimed mainly at sites with more than one million unique pages
or more than 10,000 pages changing daily. Evidence for this claim Google directs its crawl-budget guidance mainly to sites with more than one million unique pages or more than 10,000 pages that change daily. Scope: Google's examples for deciding whether its large-site crawl-budget guide is relevant; these are not crawl guarantees or definitions of an enterprise company. Confidence: high · Verified: Google Search Central: Large site's guide to managing crawl budget
AI summary
A condensed take on the Advanced version:
- Enterprise audits are a different job, not a bigger checklist — driven by scale (millions of URLs, multiple CMSsA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms./CDNs, international teams, JS stacks) and by the org that has to ship the fixes.
- Scope before you crawl. My four-step process: (1) interview stakeholders for real pain points + map who owns what; (2) segment by CMS/region/template/team; (3) scope deliberately — auditing everything wastes time and money; (4) build deliverables that get implemented.
- IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. first. Start in the GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason., then crawl data, then
logs as ground truth. Google’s bar: accessible to GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer.,
200status, indexable content — and eligibility still isn’t a guarantee. - Think in templates, not pages. One bad canonical hits hundreds of thousands of URLs. Micro-gains on single pages while macro problems sit is the classic mistake.
- Uniquely-enterprise problems: crawl waste (only matters past ~1M pages/week or 10K/day), duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. at scale, JS renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. at scale, and hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. — the most expensive issue on large international sites.
- Don’t trust the tool score (Splitt: “don’t follow your tools blindly”); context decides severity. A 404 spike from a content purge is expected, not a fire.
- Prioritize impact × effort, in dollars (e.g. $400/referring domain recovered): crawl/index → on-page at scale → links. Separate findings from recommendations.
- Ship 5-10 fixes, in developer-ticket format with acceptance criteria + business impact — not a 300-slide deck. Separate the exec summary from the technical detail.
- Wrap governance around it (standards, pre-launch checks, sampling/always-on crawls) so fixes don’t regress. The real lever is organizational alignment.
Official documentation
Primary-source documentation that should ground an enterprise audit’s recommendations.
- Search Essentials — Technical requirements — the three minimums for indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. eligibility: GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. access, HTTP
200, indexable content. - Optimize your crawl budget — crawl capacityThe number of URLs an engine will crawl in a timeframe. + demand, who actually needs it (1M+ pages/week, 10K+/day), and the URL-inventory recommendations.
- Core Web Vitals — LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good., INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good. (replaced FIDFirst Input Delay — a retired Core Web Vital that measured the delay before the browser could begin processing a page's first interaction. Good was ≤100 ms. Replaced by INP in March 2024. in March 2024), and CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good., measured at the 75th percentile of CrUXChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program. field dataPerformance metrics captured from real users, not lab tests..
- JavaScript SEO basics — renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. is deferred and separate; cross-reference HTML against rendered output.
- Do I need an SEO? — Google’s own description of what a proper audit should deliver (realistic estimates, expected outcomes — and never a rankings guarantee).
- Page Indexing report (Search Console Help) — the first diagnostic layer for indexability at scale.
Bing / Microsoft
- Bing Webmaster Guidelines — content quality and technical requirements.
- Keeping content discoverable with sitemaps in AI-powered search (July 2025) — enterprise sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. limits (50,000 URLs/file; index files referencing billions of URLs) and why accurate
lastmodmatters. - IndexNow — real-time URL submission for high-velocity enterprise sites; the basis for always-on crawl alerting.
Quotes from the source
On-the-record statements from Google reps relevant to auditing large sites.
Martin Splitt, Google — what an audit is actually for
- “A technical audit, in my opinion, should make sure no technical issues prevent or interfere with crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. or indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” — Martin Splitt, Search Central, November 2025. Coverage
- “Please, please don’t follow your tools blindly. Make sure your findings are meaningful for the website in question.” — Martin Splitt, November 2025. Coverage
- “A high number of 404s, for instance, is expected if you removed a lot of content recently.” — Martin Splitt, November 2025. Coverage
Gary Illyes, Google — crawl budgetThe number of URLs an engine will crawl in a timeframe. and quality
- “For most sites, crawl budgetThe number of URLs an engine will crawl in a timeframe. is not something to worry about. For really large sites, it becomes something to consider looking at.” — Gary Illyes. Crawl budget docs
John Mueller, Google — indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. large sites
- “I strongly recommend not relying on trying to force indexing” for large sites — manual indexing requests don’t scale. — John Mueller, Google Search Central.
- “Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. help Google understand site structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders., and if all pages are linked to all other pages, there’s no real structure.” — John Mueller, Google Search Central.
Enterprise audit checklist
Run it in this order — scope first, then foundation, then everything else.
Before any crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. runs
- Stakeholder interviews done — known pain points captured (traffic drop, launch, migration, redesign, compliance).
- Stakeholder map drawn — who owns templates, content, CDN, legal sign-off.
- Site segmented by CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. / region / page type / template / team.
- Scope agreed — which segments are in, which are explicitly out, and why.
Indexability (the foundation)
- GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. reviewed; “Discovered/Crawled – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” buckets analyzed by pattern.
- Key templates return
200; dead pages return404/410. - Crawl waste identified (parameters, facets, session IDs, internal search) and a
robots.txt/consolidation plan drafted. - SitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. list only canonical, indexable URLs with honest
lastmod. - JS-rendered content cross-referenced: raw HTML vs. rendered output at the template level.
Content & international
- Duplicate/thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count. assessed by template pattern, not page-by-page.
- hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. validated at scale (return tags, correct region/language) on every international template — highest-stakes check.
Links
- 404s with backlinks flagged for reclamation (with $ value attached).
- Orphaned and deeply-buried important pages (>3 clicks) identified.
Deliverable & governance
- Findings distilled to 5-10 prioritized recommendations.
- Each fix written as a developer ticket (problem, acceptance criteria, steps, business impact).
- Executive summary separated from technical detail.
- Monitoring cadence + pre-launch checklist agreed so fixes don’t regress.
The mental models
1. Scope before you crawl — the four steps. Pain points → segment → scope → deliverables. The first three happen before any tool runs. Skipping them is how you end up with a 40M-row crawl nobody can use.
2. Templates, not pages. On an enterprise site, every meaningful problem and fix is a pattern. Ask “which template generates this?” before “which page is broken?” One canonical fix can move hundreds of thousands of URLs; one page fix moves one.
3. The indexability ladder — climb in order.
Accessible to GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. → returns 200 → indexable content → actually indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. →
ranks. Find the rung a segment is failing on (start in the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.)
before changing anything above it.
4. Impact × effort, priced in dollars. Plot findings on an impact/effort matrix, then convert impact to money ($/referring-domain, $/traffic recovered). Order of operations: crawl/index issues → on-page at scale → links. Findings ≠ recommendations — executives get the short, prioritized recommendation list.
5. Context beats the score. A tool’s number has no idea what your site is supposed to do. A 404 spike after a purge, a “low score” from correctly-noindexed facets — both fine. Validate every finding against how the site is actually built.
6. Audit → governance loop. An audit is a snapshot; an enterprise site is a movie. Feed findings into standards, pre-launch checks, and continuous/sampling crawls so the same issues don’t return next quarter.
Enterprise audit cheat sheet
Order of operations
| Phase | What you do | Primary source |
|---|---|---|
| 0. Scope | Pain points → segment → scope | Stakeholder interviews, site-structure reports |
| 1. IndexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. | Can engines reach + index it? | GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. → crawl data → logs |
| 2. Content | Duplicates/thin at template level | Crawl + content reports |
| 3. International | hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. validation at scale | Crawl + GSC international targeting |
| 4. Links | 404 reclamation, orphans, depth | Backlink + internal-link reports |
| 5. Deliver | 5-10 fixes as dev tickets, $-framed | — |
| 6. Govern | Standards, pre-launch checks, monitoring | Sampling/always-on crawls, IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. |
Who needs to worry about crawl budgetThe number of URLs an engine will crawl in a timeframe. (Google’s threshold)
- 1M+ unique pages updating weekly, or 10K+ pages updating daily, or lots of “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal..” Below that: don’t worry about it.
Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data. targets
- LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good. < 2.5s · INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good. < 200ms (replaced FID, March 2024) · CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good. < 0.1 — at the 75th percentile of CrUXChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program. field dataPerformance metrics captured from real users, not lab tests..
Over-prioritized (lower than SEOs think)
- Redirect chainsA → B → C instead of A → C. Each hop loses link equity and adds latency. (Google follows ~10 hops; worry past 5) · double slashes · multiple H1s (fine in HTML5) · locale format nitpicks (underscore vs. dash).
Highest-stakes
- hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. on international templates — the most expensive issue · canonical/template errors that ripple across hundreds of thousands of URLs · anything that can deindex at scale.
Deliverable rules
- 5-10 items, not 300 slides · dev-ticket format (problem + acceptance criteria + steps + business impact) · exec summary separate from technical detail · findings ≠ recommendations.
Patrick's relevant free tools
- SEO Incident Simulator — Practice thirty deterministic technical SEO incident investigations — indexability, crawl controls, redirects, sitemaps, markup, caching, DNS, bot verification, rendering, hreflang, and faceted navigation — with clearly labeled fixture evidence and Find → Fix → Verify handoffs.
- hreflang Generator + Linter — Enter your URL × locale matrix and get bidirectional hreflang markup as head tags, sitemap XML (auto-split past 50,000 URLs), and Link headers — linted live for wrong region codes, duplicates, and missing fallbacks. Runs entirely in your browser.
- returntag - hreflang checker — Enter one URL or an XML sitemap and the returntag - hreflang checker crawls the whole hreflang cluster — missing return tags, broken targets, self-reference and x-default checks, language-code validation, and head vs Link header vs sitemap disagreements — on an interactive cluster map with CSV export. The check Search Console's International Targeting report used to run.
Tools for enterprise SEO audits
No single tool covers an enterprise audit. The stack:
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. (free, official) — start here. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason., Core Web VitalsWeb Vitals is Google's initiative (launched May 2020) for unified page-experience quality signals. Core Web Vitals — LCP, INP, and CLS — are the subset used in ranking; the rest (TTFB, FCP, TBT, Speed Index) are diagnostic, not ranking factors., Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root)., URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version.. The first diagnostic layer.
- Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. (free, official) — Site Scan (on-demand full-site audit), Top Insights (pages missing from sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., crawl errors), IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. Insights, Crawl Control, and 16-month Search Performance.
- Enterprise crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — Botify, Lumar (DeepCrawl), Sitebulb, and Screaming Frog for crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. at scale, segmented by template. Crawls of millions of URLs can take 48-72 hours, which is exactly why you segment first.
- Log file analysisLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. — the ground truth for what bots actually fetched and where crawl budgetThe number of URLs an engine will crawl in a timeframe. is being wasted (Screaming Frog Log File Analyser, or pipe logs into BigQuery / a log platform).
- Backlink analysis — Ahrefs for the link audit and 404 reclamation (attaching $ value to recovered referring domains).
- Enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. platforms — BrightEdge, Conductor, seoClarity, Ahrefs Enterprise for ongoing monitoring across large sites.
- Performance — PageSpeed InsightsPageSpeed Insights (PSI) is a free Google tool at pagespeed.web.dev that reports two kinds of data for a URL: real-user field data from the Chrome UX Report and a single Lighthouse lab run with the 0–100 Performance score. Only the field Core Web Vitals are what Google uses for ranking. + CrUXChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program. for field Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data..
A reminder that overrides all of the above: don’t let any tool’s overall “health score” set your priorities. It’s not a ranking factor and it doesn’t know your site.
Quarterly enterprise SEO audit SOP
Use this as a recurring audit cycle, not a one-time findings dump.
- Confirm scope and owners. Re-interview the owners of each CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms., region, and template. Record major launches, migrations, traffic changes, and any systems that moved in or out of scope. Done means every included segment has a named owner and a business question the audit must answer.
- Compare expected and indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. inventory. Export sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. totals and review the GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. by segment. Investigate material gaps by reason before running a full crawl. Done means each gap is classified as intentional or queued for diagnosis.
- Sample every important template. Check status, canonical, robots directives, rendered content, internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., and hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. where applicable. Done means every revenue-driving template has a current pass/fail record.
- Crawl only the segments that need evidence. Apply saved include/exclude rules so parameters and low-value spaces do not bury the findings. Done means every crawl maps back to an owner, template, and audit question.
- Prioritize the short list. Score confirmed findings by business impact and feasibility, then reduce the deliverable to 5–10 recommendations. Done means each recommendation has an affected pattern, owner, impact case, and effort estimate.
- Write implementation-ready tickets. Include reproduction steps, acceptance criteria, and the check that proves the fix worked. Done means engineering can estimate the ticket without reopening the audit deck.
- Close the loop. Add shipped fixes to monitoring and pre-launch checks, and carry unresolved items into the next cycle with a reason. Done means the next audit starts from a change log rather than from zero.
Enterprise audit mistakes that waste the most time
Crawling the whole site before defining the question
Why it fails: a 40-million-URL export mixes unrelated systems, templates, and owners into one pile. The volume looks impressive but makes patterns harder to see.
Do instead: interview stakeholders, segment the site, and decide which business questions each crawl must answer before it runs.
Treating every URL as a separate problem
Why it fails: fixing individual pages leaves the template that generated the error untouched, so the issue returns across thousands of URLs.
Do instead: identify the template, CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. rule, or parameter pattern behind each finding and write the recommendation at that level.
Ranking findings by a tool’s health score
Why it fails: the score cannot know that a noindexed facet is intentional or a 404 spike followed a planned content purge.
Do instead: validate findings against site behavior and business impact, then rank them by impact and feasibility.
Delivering hundreds of unowned findings
Why it fails: a large deck shifts the sorting work to stakeholders and leaves no clear first move.
Do instead: deliver 5–10 recommendations with an owner, acceptance criteria, effort, and a plain-language impact case.
Ending the audit when the report is sent
Why it fails: templates and platforms keep changing, so fixed issues quietly return.
Do instead: turn each shipped fix into a regression check, monitoring rule, or pre-launch requirement.
Prompts for enterprise audit work
Turn segmented crawl data into template-level findings
Paste a CSV excerpt with URL, template, status, canonical, robots, depth, and indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. state. Remove sensitive fields first. Expect a grouped analysis, not page-by-page advice.
You are helping triage an enterprise SEO crawl. Group the rows below by template and
failure pattern. For each pattern, report: affected segment, observable evidence,
likely system-level cause, pages affected in this sample, business risk, owner to
involve, and the next check needed to confirm the diagnosis. Do not infer revenue or
claim causation from correlation. Separate intentional states from probable defects.
[PASTE SEGMENTED CRAWL ROWS]Convert a confirmed finding into a developer ticket
Paste one verified finding plus the relevant template behavior. Expect a ticket an engineer can estimate without reading the full audit.
Convert this confirmed enterprise SEO finding into a developer-ready ticket. Return:
problem statement, affected template or rule, reproduction steps, expected behavior,
acceptance criteria, validation test, rollback trigger, dependencies, owner, and a
plain-language business impact statement. Preserve unknowns as questions. Do not
invent traffic, revenue, or implementation estimates.
[PASTE VERIFIED FINDING AND EVIDENCE] Resources worth your time
My related writing
- What is an Enterprise SEO Audit & How To Do One — the full 4-step process and audit types.
- Enterprise Sites Are Where Technical SEO Shines — crawl strategy options, the impact/effort matrix, and high-priority projects.
- Enterprise SEO Challenges & Mistakes You Need To Overcome — the organizational layer: buy-in, incentives, and micro-vs-macro.
- Enterprise SEO Strategies For Maximum Growth — the broader program around the audit.
My speaking
- Enterprise SEO Chaos (SMX Seattle 2016) — what enterprise scale actually looks like, from 4 years in-house at IBM (378,000 employees, 170+ countries): 24 versions of one URL, 14-hop redirect chainsA → B → C instead of A → C. Each hop loses link equity and adds latency., and “Everything Has To Work Together.”
- What I Learned from Auditing Over 1,000,000 Websites — why most common ≠ most important, and the prioritization thresholds that actually matter.
From around the industry
- Google — Search Central technical-audit guidance — Martin Splitt on tool scores and audit methodology (Nov 2025); Search Engine Journal coverage.
- Martin Splitt — “Why we need to talk about audits” (YouTube, Nov 2025) — primary source video on audit purpose and tool-score misconceptions, from Google Search Central.
- Martin Splitt — “How to perform a technical SEO audit” (YouTube, Nov 2025) — companion video covering the three-step audit framework from Google.
- Screaming Frog: How to Do an Enterprise SEO Audit the Right Way — tool-centric, four-stage framework from one of the most-used enterprise crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
- Sitebulb: Key Considerations in Enterprise SEO Auditing — crawl-first methodology guide covering segmentation and prioritization at scale.
- Search Engine Land: What your enterprise SEO audit may be missing — covers canonical, indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. gaps, and holistic audit blind spots.
- Search Engine Land: 6 reasons why SEO audits seem like a waste and how to fix them — practical fixes for audit deliverables that don’t get implemented.
- r/TechSEO — the community for large-site crawl/index debugging.
Test yourself: Enterprise SEO audits
Five questions on scoping, diagnosis, and delivery. Pick an answer for each, then check.
Stats worth citing
- Crawl-budget threshold: Google’s own bar for when crawl budgetThe number of URLs an engine will crawl in a timeframe. matters — 1M+ pages updating weekly or 10K+ pages updating daily. Below that, it’s not worth worrying about. Source
- Enterprise sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. scale (Bing): a single sitemap index fileA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file. can reference up to
2.5 billion URLs (50,000 child sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. × 50,000 URLs) — and multiple indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. files
push that far higher. The scale is the point: at this size, accurate
lastmodis what keeps discovery efficient. Source - Buy-in by the numbers: I’ve funded redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. projects by assigning $400 per referring domain recovered — turning a tech chore into a revenue case is what gets enterprise fixes shipped. Source
Enterprise SEO Audit
An enterprise SEO audit is a systematic evaluation of a large-scale website — usually millions of URLs across multiple CMS platforms, teams, and regions — to find the technical, indexing, content, and link issues that limit organic visibility. Unlike a standard audit, it must account for organizational complexity and ruthless prioritization, because technical perfection at scale is neither possible nor worth paying for.
Related: Crawl Budget, Hreflang, Indexing, Core Web Vitals, JavaScript SEO
Enterprise SEO Audit
An enterprise SEOEnterprise SEO is the practice of doing SEO at scale — for large, complex sites (often tens of thousands to millions of pages) across multiple teams, CMSs, and stakeholders. It uses the same ranking factors as any site; what changes is the scale, the technical debt, and the organizational coordination. audit evaluates a large enterprise website to find the issues and opportunities that will improve its rankings and visibility in search. What makes it “enterprise” isn’t a fancier checklist — it’s scale and the organization behind it. You’re dealing with millions of URLs, multiple content management systems and CDNs, international footprints across regional teams, and JavaScript-heavy stacks, all owned by different groups with their own roadmaps and incentives. As I’ve said before, enterprise SEO auditsAn SEO audit checklist is the structured set of things you review across a site's technical health, on-page elements, content quality, and off-page authority — used to produce a short, prioritized action plan, not an exhaustive 100–200 item inventory. are entirely different beasts to “regular” audits because of all the complications that come from large sites.
That changes the work in two big ways. First, you can’t audit everything — an unsegmented full-site crawl of a 40-million-URL site produces unusable data. So enterprise audits start by finding stakeholder pain points, segmenting the site (by CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms., region, page type, and team ownership), and scoping deliberately before any tool runs. Second, you think in templates, not pages: one misconfigured canonical on a category template can affect hundreds of thousands of URLs at once, so the fixes that matter are the ones that ripple across a pattern.
The other half of the job is organizational. The #1 reason enterprise audits fail isn’t a missed finding — it’s that nobody acts on the deck. The deliverable can’t be a 300-slide dump; it has to be a prioritized list of the 5–10 highest-impact, feasible fixes, framed in business terms (“recover 2,500 links worth $X”) and handed over in developer-ticket format with acceptance criteria. And because enterprise sites change constantly, an audit feeds into ongoing governance — standards, pre-launch checks, and monitoring — so the issues you fix don’t quietly come back.
Related: Crawl Budget, Hreflang, Indexing, Core Web Vitals, JavaScript SEO
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 25, 2026.
Editorial summary and recorded change details.Summary
Aligned legacy tool display names with returntag - hreflang checker and Scout Site Audit Free.
Change details
-
Updated the linked tool names to match their current public labels.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 19, 2026.
Editorial summary and recorded change details.Summary
Removed an unsourced Core Web Vitals origin pass-rate statistic ("Jan 2026 CrUX": ~68% LCP / ~87% INP / ~81% CLS / ~56% pass-all-three) from the Stats lens — it carried no first-party citation, and the aggregate-dashboard tool that kind of figure is normally pulled from (Google's CrUX Dashboard in Looker Studio) is now deprecated. Reviewed the article against a Wave 94 structured-research packet (screened synthetic/templated, used for topic-coverage only); verified the crawl-budget doc citation still resolves and its 1M-pages-weekly/10K-pages-daily threshold text is unchanged, and confirmed the quoted Martin Splitt lines against Search Engine Journal's and ppc.land's coverage.
Change details
-
Removed the unsourced ~68%/~87%/~81%/~56% Core Web Vitals pass-rate stat from the Stats lens; no defensible dated first-party source for it.
Full comparison unavailable — no prior snapshot was archived for this revision.