Ecommerce SEO Audit
How I audit an online store — and why I start with Google Search Console and crawl data, not on-page tweaks. A prioritized, scale-first ecommerce SEO audit framework: crawlability, indexation, duplicate content, on-page at scale, technical, structured data, internal links, out-of-stock handling, and off-page — sorted by impact and effort.
An ecommerce SEO audit isn't a 500-point checklist — it's the work of finding the small set of issues holding a store's rankings and revenue back, then prioritizing them by impact and effort. Because ecommerce sites hit problems at scale, I start the audit in Google Search Console (the Page Indexing report) and crawl data, not on-page tweaks: the biggest wins are usually crawling and indexation issues affecting thousands of URLs at once. Duplicate content from facets, collections, and variants is the #1 ecommerce-specific issue audits uncover; thin category pages are #2. The goal is to fix the stuff that matters most and skip the busywork.
Evidence for this claim Search Console's Page indexing report identifies indexed and non-indexed URLs and groups reasons pages are not indexed. Scope: Google Search Console audit data. Confidence: high · Verified: Search Console Help: Page indexing report Evidence for this claim Google's Rich Results Test and product structured-data requirements can be used to validate product markup eligibility. Scope: Google product structured-data validation. Confidence: high · Verified: Google Search Central: Product structured dataTL;DR — An ecommerce SEOEcommerce SEO is the practice of optimizing an online store so its product and category pages rank in organic search and attract purchase-intent visitors. It uses the same Google algorithm as any other site, but compounds the usual SEO work with commerce-specific challenges like faceted navigation, product variants, and platform-imposed URLs. audit is a health check for your online store — finding the things that stop Google from showing your product and category pages, and fixing the ones that matter most. The trick: don’t start by tweaking titles and adding keywords. Start in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and find out which pages Google can even find and indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., because on a store the biggest problems hit thousands of pages at once.
What an ecommerce SEO audit is
An audit is just a structured look at your store to answer one question: what’s holding it back in search? It’s the same idea as a checkup at the doctor — you’re looking for problems before they cost you, and you fix the serious ones first.
The reason ecommerce stores need their own kind of audit is scale. A blog has a few hundred pages. A store can have tens of thousands — one page per product, plus a category page for every way to group them, plus a near-identical copy of each page for every color, size, and sort order. That’s a lot of pages, and a lot of them look almost the same to Google.
Where to start (and where not to)
Most people start an audit by rewriting title tagsThe title tag is the HTML title element in a page's head that specifies the document's title. It's the primary source for the SERP title link and a confirmed light ranking factor — but since August 2021 Google doesn't always show it verbatim. and stuffing in keywords. That’s backwards. If Google can’t crawl or index a page, no amount of keyword work will help — the page isn’t in the race at all.
So I start in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., which is Google’s free dashboard for site owners. The most useful screen is the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.: it tells you which of your pages Google has actually indexed, and gives you a reason for every page it hasn’t. On a store, that report usually surfaces the big money problems right away — thousands of duplicate pagesThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling., pages Google foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. but never bothered to crawl, or (worst case) pages accidentally marked “do not index.”
The two problems almost every store has
- Duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.. This is the big one. Your filters (color, size, price) create a brand-new web address every time someone clicks one, and most of those addresses show nearly the same products. Variants (the same shirt in five colors) often each get their own page too. Google ends up wading through thousands of near-copies.
- Thin category pages. A category page that’s just a grid of products with no description is hard for Google to rank — there’s not enough on it to tell Google what it’s about.
What you’ll need
- Google Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. — free, and the single most important tool here.
- A site crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — something that visits all your pages like Google does and lists the problems. I use Ahrefs Site Audit (this is my site, so that’s the honest answer), but Ahrefs Webmaster Tools gives you a free crawl of sites you verify, and Screaming Frog is a popular desktop option.
- PageSpeed InsightsPageSpeed Insights (PSI) is a free Google tool at pagespeed.web.dev that reports two kinds of data for a URL: real-user field data from the Chrome UX Report and a single Lighthouse lab run with the 0–100 Performance score. Only the field Core Web Vitals are what Google uses for ranking. and the Rich ResultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. Test — both free from Google — for checking speed and your product markup.
Want the full, prioritized checklist — every area to audit, in order, with the ecommerce gotchas and how to rank the fixes by impact? Switch to the Advanced tab.
Evidence for this claim Search Console's Page indexing report identifies indexed and non-indexed URLs and groups reasons pages are not indexed. Scope: Google Search Console audit data. Confidence: high · Verified: Search Console Help: Page indexing report Evidence for this claim Google's Rich Results Test and product structured-data requirements can be used to validate product markup eligibility. Scope: Google product structured-data validation. Confidence: high · Verified: Google Search Central: Product structured dataTL;DR — I audit a store in roughly this order: crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index. → indexation → duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. → on-page at scale → technical (CWVGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data., HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', canonicals) → structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. → internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. → out-of-stock handling → off-page. I start in the GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. and a fresh crawl, not on-page, because on a store the highest-leverage problems hit thousands of URLs at once: faceted-nav duplication, accidental sitewide
noindex, “Discovered – currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” at scale, sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. full of redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't.. The deliverable is a prioritized impact/effort list, not a 500-point report. Most stores never need to worry about crawl budgetThe number of URLs an engine will crawl in a timeframe. — but the ones that do really do.
Audit philosophy: client-first, not issue-first
A crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. will hand you 170+ issue types. If you dump all of them into a report, you’ve produced a document, not an audit. The job is to find the handful that move rankings and revenue and ignore the rest.
The starting point isn’t the tool — it’s the pain. As I’ve put it before: “If clients are coming to you asking for an audit, they already have a pain point. Talk to them. Solve that one thing and they’ll be happy with the audit.” On a store the pain is usually concrete: traffic dropped, a category stopped ranking, a migration went sideways, new products aren’t getting indexed. Anchor the audit to that, then expand.
And know when to stop. From our study of over a million domains: “Sometimes, the best course of action is to do nothing because the costs outweigh the benefits.” Not every flagged issue is worth a developer ticket.
Why I start with Search Console and crawl data, not on-page
A crawl is an evidence source, not the audit itself. Record how many discovered URLs were actually crawled and what the tool did not evaluate before turning any warning into a ticket.
Use my lightweight crawler for a fast raw-HTML baseline before you move into Search Console samples and rendered-page checks. Scout Site Audit Free Free
- Run the store or a representative section and note the discovered-versus-crawled scope.
- Group findings by template and business impact instead of copying every warning into the report.
- Verify high-priority patterns in Search Console and representative rendered pages.
The instinct to open the audit with title tagsThe title tag is the HTML title element in a page's head that specifies the document's title. It's the primary source for the SERP title link and a confirmed light ranking factor — but since August 2021 Google doesn't always show it verbatim. and headings is the most common mistake I see. On an ecommerce site the math is against it: tweaking a title helps one page; fixing an indexation pattern helps thousands. The biggest wins on a store are almost always crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexation problems at scale, and the stakes are highest exactly here — “one mistake can keep millions of pages out of the index or remove an entire site from search results.”
So the audit runs in priority order, top to bottom. Here’s the whole thing.
Step 1 — Crawlability
Can bots reach what matters, and are they wasting their time on what doesn’t?
The ecommerce-specific traps:
- Faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals. URL explosion. Filters (color, size, price range) combine into thousands of unique URLs. Google is candid about the cost: “the crawlers will typically access a very large number of faceted navigation URLs before the crawlers’ processes determine the URLs are in fact useless.”
- Parameter variants — session IDs, tracking codes (
?utm_source=), sort orders — multiply the same content across URLs. - Crawl traps — internal search results, infinite calendars, unbounded filter stacking.
robots.txtover- or under-blocking — accidentally blocking CSS/JS needed to render, or a valid product path; or failing to block a junk parameter space.- JS-gated navigation. Google wants real links: “use
<a href>tags when creating links to other content. Don’t use JavaScript events on other HTML DOM elements for navigation.”
Audit steps:
- Fetch and read
robots.txt; confirm nothing important is disallowed and that low-value parameter spaces are. - In the crawler, pull “Blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing.” — any important pages caught?
- GSC Crawl Stats: look for spikes in crawled URLs against a flat indexed count — that gap is crawl waste.
- Export all crawled URLs and cluster by path/parameter pattern to size the faceted/parameter problem.
On crawl budget — de-escalate first. Most stores don’t have a crawl-budget problem, and I’ll say so in the audit rather than invent one: “Most sites don’t need to worry about crawl budget, but there are few cases where you may want to take a look.” It becomes real around Google’s thresholds — “Large sites (1 million+ unique pages)…” changing weekly, or 10,000+ pages changing daily — or when “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” is large. When it is real, you fix it by removing waste, not by asking Google to crawl more. (Deep dive: crawl budgetThe number of URLs an engine will crawl in a timeframe..)
Step 2 — Indexation
Crawlable isn’t indexed. The GSC Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. is the center of gravity for the whole audit — “see which pages Google can find and index on your site, and learn about any indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. problems encountered.”
The statuses that matter most on a store:
- Duplicate without user-selected canonicalA Google Search Console Page Indexing status: Google found this page to be a duplicate, you didn't declare a canonical, so Google chose a different page as the canonical — and this URL isn't indexed. / Google chose a different canonical — the faceted/variant duplication problem, made visible.
- Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error. — usually thin product or near-duplicate filter pages.
- Discovered – currently not indexed — Google knows the URL but hasn’t crawled it; on large catalogs this is a crawl-priority signal.
- URL marked ‘noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.’ — verify it’s intentional. The accidental sitewide version is the “one mistake” scenario.
- Soft 404A soft 404 is a URL that returns a success status code (usually 200 OK) even though the page is empty, missing, or shows a 'not found' message. It isn't a status code a server sends — it's a label search engines apply after comparing the response code against the rendered content, and they treat the page like a 404 for indexing. — a 200 response on a page that’s effectively empty (out-of-stock, zero-result search). Google warns these “will continue to be crawled, and waste your budget.”
Audit steps:
- Export “Not indexed” URLs by reason; quantify each bucket.
- Compare GSC indexed count against your known catalog size — a large gap means Google isn’t finding pages.
- Run URL Inspection on a sample from each reason bucket.
- Audit the XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.: it should list only canonical, indexable, 200-status URLs. Redirected, noindexed, or canonicalized-away URLs in a sitemap send conflicting signals — pull them.
Step 3 — Duplicate content (the #1 ecommerce issue)
This is the single most common thing an ecommerce audit uncovers, and it’s worth its own step. Google’s Gary Illyes has estimated that roughly 60% of the internet is duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. — and stores are overrepresented.
Where it comes from:
- Faceted navigation —
?color=blue,?size=S&color=blue,?size=S&color=blue&sort=priceall serve near-identical content. - Parameter-order permutations —
?color=blue&size=Sand?size=S&color=blueare the same page twice. - Product variants as separate URLs —
?variant=…without a self-referencing canonical or a canonical to the parent. - A product living under multiple category paths, both indexable.
- HTTP/HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', www/non-www, trailing-slash inconsistency that doesn’t 301 to one canonical.
The Shopify gotcha worth calling out by name. Shopify does canonicalize
/collections/{collection}/products/{product} to the clean /products/{product}
URL — so people assume it’s handled. It isn’t fully: the internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. from
collection pages still point at the non-canonical /collections/... version, so
the canonical product URL receives no internal link equity through normal
navigation, and it can show up as an orphan. Fixes: add an “All products” page that
links the canonical /products/ URLs, or edit collection templates to link the
canonical URL directly.
Audit steps:
- GSC: pull both “Duplicate” buckets from Page Indexing.
- Crawler Content Quality / duplicate-cluster report: find clusters with no canonical specified.
- Confirm faceted parameter URLs are handled — Google’s recommended prevention
is
robots.txtblocking of filter parameters, or fragment-based (#) filtering, which “will have no impact on crawling.” Don’t reach fornoindexhere (see Myths). - Spot-check that Shopify
/collections/*/products/*URLs canonicalize to/products/*— and that something links the canonical version.
(Background: duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. and canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it..)
Step 4 — On-page at scale
On a store, on-page is a templating problem, not a copywriting one — you’re auditing patterns, not pages.
The recurring issues:
- Thin category pages (the #2 ecommerce issue). John Mueller is direct about it: “When the ecommerce category pages don’t have any other content at all, other than links to the products, then it’s really hard for us to rank those pages.” But don’t overcorrect into a wall of text — “maybe 90%, 95% of that text is unnecessary,” and “our algorithms sometimes get confused when they have a list of products on top and essentially a giant article on the bottom.” The target is a short, useful block (sourcing, materials, sizing, popularity) near the top — not filler that buries the grid.
- Templated title collisions —
Buy {Product} | Storepatterns that go near-identical across products, or two products sharing a name. - Missing/empty H1s and titles at the template level.
- Meta descriptionsThe meta description is an HTML head tag — `<meta name=\"description\" content=\"…\">` — that suggests a short summary of the page for the search snippet. It's not a Google ranking factor, and Google rewrites it the majority of the time, but a good one can still lift click-through. — worth doing on your top category and product pages, less worth it on the long tail (Google rewrites them most of the time anyway).
Audit steps:
- Crawler On-Page report: filter for missing title, missing H1, missing meta description, duplicate titles, duplicate H1s.
- Flag category pages with little to no unique body copy (thin-content candidates) and cross-reference against “Crawled – currently not indexed.”
- Export titles and check for template collisions.
A note on what not to spend audit time on: multiple H1s. They’re valid HTML5 and near-irrelevant to ranking — 51.3% of sites have them somewhere, per my own million-domain study. Skip it.
Step 5 — Technical (CWV, HTTPS, canonicals, mobile)
Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data.. Google’s passing thresholds: LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good. < 2.5s, INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good. < 200ms, CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good. < 0.1. Ecommerce failure modes are predictable — unoptimized hero images and third-party scripts (chat, A/B, pixels) hurt LCP; add-to-cart and checkout JS hurt INP; images without explicit dimensions and injected promo bars hurt CLS. Audit by URL group in the GSC CWV report so you can see whether it’s product, category, or checkout templates failing, then confirm on representative templates in PageSpeed InsightsPageSpeed Insights (PSI) is a free Google tool at pagespeed.web.dev that reports two kinds of data for a URL: real-user field data from the Chrome UX Report and a single Lighthouse lab run with the 0–100 Performance score. Only the field Core Web Vitals are what Google uses for ranking.. (See Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data..)
HTTPS. Everything — product, cart, checkout — on HTTPS, no mixed contentMixed content is when a page served over HTTPS loads a sub-resource — a script, stylesheet, image, iframe, or similar — over insecure HTTP. Browsers' current taxonomy is upgradable versus blockable; active mixed content (scripts, styles, iframes) is blocked, and passive mixed content (images, audio, video) is warned about or increasingly auto-upgraded, with exceptions like CORS-enabled images and srcset/picture candidates that are blockable, not upgradable. (HTTP images/resources on HTTPS pages), and canonicals/redirects pointing at the HTTPS version.
Canonicals. The common ecommerce errors: canonicals pointing at 4XX pages;
non-canonical URLs in the sitemap; paginated pages canonicalized to page one
(don’t); and — the quiet one — internal links pointing at non-canonical versions so
the canonical never accumulates link equity. Multiple rel=canonical tags on one
page get ignored.
Mobile-first. Google indexes the mobile version. Make sure the mobile product page isn’t a stripped-down one missing the description, the structured data, or full-size images.
Step 6 — Structured data
For a store, the priority types are Product / ProductGroup, BreadcrumbList, Organization (with return policy), and Review / AggregateRating. Adding more valid properties widens eligibility — Google: “adding the more properties you can add, the more enhancements your page can be eligible for.” (The one myth to keep in your head: schema makes you eligible for rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show.; it doesn’t make you rank.)
Audit steps:
- Run the Rich Results Test on a representative product and category page.
- GSC Enhancements: check Product Snippets and Merchant Listings for errors/warnings.
- Check required vs. recommended separately per feature — don’t collapse them.
For merchant listings, Google’s required set is
name,image, andoffers(anOffer, withprice,priceCurrency, andavailability). For product snippets,nameis required and you need at least one ofreview,aggregateRating, oroffersto be eligible — Google listsaggregateRating,offers, andreviewas recommended, not required, on that feature. - Common faults: missing
image, missingoffers, missingavailability, schema prices that don’t match the visible price (a policy violation), breadcrumb schema that doesn’t match the visible trail.
Step 7 — Internal linking
The issues: orphan product pages (classic Shopify symptom, above); important category/product pages buried 5+ clicks deep; internal links pointing at redirects or 404s; best sellers not linked from high-authority hubs. Google: “add links from menus to category pages, from category pages to sub-category pages, and finally from sub-category pages to all product pages.”
Audit steps:
- Crawler Links report: pull orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. and crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops.; confirm key pages sit within ~3 clicks of the homepage.
- Pull links to redirects and broken links.
- Check that your highest-revenue products are linked from the main nav and editorial/hub pages.
Step 8 — Out-of-stock and discontinued products
A genuinely ecommerce-only section. My honest answer here is the SEO cliché for a reason: “though it’s a joke in the SEO community, ‘it depends’ is really the answer when dealing with out-of-stock products.” And “ultimately, there’s no perfect solution.” The decision turns on whether it’s temporary vs. permanent and whether the page has traffic or links:
| Scenario | What I do | Why |
|---|---|---|
| Temporarily OOS | Keep the page live | Add restock dates, waitlist, notify-me; don’t throw away ranking |
| Permanently discontinued, close replacement | 301 to the similar product | Preserves link equity if they’re genuinely similar |
| Discontinued, no match, has links/traffic | Keep live with “related products” | Retains ranking potential; route users onward |
| Discontinued, no links/traffic | 404 or 410 | Clean it up; fix internal links pointing at it |
Watch for soft 404s: out-of-stock pages, zero-result searches, or empty cart/account pages returning 200. Pull the Soft 404 bucket in GSC and resolve by pattern. (See out-of-stock productsAn out-of-stock product page is a product URL whose item can't currently be bought. The right SEO treatment depends on whether the stockout is temporary, indefinite, or permanent — keep temporary stockouts live at 200 with OutOfStock schema; 301 permanently discontinued products to a relevant replacement. for the full framework.)
Step 9 — Off-page
Lighter for most stores, but worth a pass:
- GSC Performance: branded vs. non-branded click share — heavy branded reliance means weak non-branded discovery.
- Backlink profile by page type — are category/product pages earning links, or only the homepage and blog?
- Find 404 pages that have backlinks and 301 them to the closest live page (reclaim the equity).
- Competitor link gap on category pages.
How to prioritize the findings
Everything above produces a list. The list is not the deliverable — the sorted list is. I score each finding on an impact/effort matrix: “anything high-impact and low-effort is a quick win, so those tasks should be tackled first.” For a store:
- High impact / low effort — do first: a stray sitewide
noindex; a sitemap full of redirect/noindex URLs; missing canonicals on variant URLs; internal links pointing at redirects; out-of-stock soft 404s. - High impact / high effort — plan and schedule: faceted-nav architecture; Core Web VitalsWeb Vitals is Google's initiative (launched May 2020) for unified page-experience quality signals. Core Web Vitals — LCP, INP, and CLS — are the subset used in ranking; the rest (TTFB, FCP, TBT, Speed Index) are diagnostic, not ranking factors. work; rolling Product schemaProduct schema (schema.org/Product) is structured data that tells search engines a page's product name, price, availability, and reviews so it can appear in Shopping-style rich results. It's separate from a Google Merchant Center feed, though Google reconciles the two. across templates; crawl-depth / architecture fixes.
- Low impact / low effort — when time allows: meta descriptions on long-tail products; minor title-template tidy-ups.
- Low impact / high effort — skip: redirect-chain cleanup on zero-traffic pages; Open GraphOpen Graph (OG) tags are `<meta>` elements in a page's head, defined by the Open Graph protocol (ogp.me, created by Facebook), that describe a page as a shareable object — its title, description, image, URL, and type. They control how a link preview card looks when the page is shared on Facebook, LinkedIn, Slack, Discord, WhatsApp, and iMessage. They are not a direct Google ranking factor, though Google reads og:title, og:image, and og:site_name as inputs to how a result appears. tags (social, not ranking).
The whole point of the matrix is permission to not do things.
Common myths to clear up in the report
- “More indexed pages = better SEO.” No — Google’s own advice is to “eliminate duplicate content to focus crawling on unique content rather than unique URLs.” Inflating the index with thin variant pages usually hurts.
- “Use
noindexto save crawl budget on facets.” No: “don’t use noindex, as Google will still request, but then drop the page… wasting crawling time.” Block the fetch (robots.txt) or use fragment filtering instead. - “Shopify handles all canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it..” It sets the canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content., but it doesn’t fix the internal-linking-to-non-canonical problem. See Step 3.
- “Crawl budget affects every store.” It doesn’t — most stores never need to think about it.
AI summary
A condensed take on the Advanced version:
- An ecommerce SEOEcommerce SEO is the practice of optimizing an online store so its product and category pages rank in organic search and attract purchase-intent visitors. It uses the same Google algorithm as any other site, but compounds the usual SEO work with commerce-specific challenges like faceted navigation, product variants, and platform-imposed URLs. audit finds the few issues holding a store back — it’s a prioritized list, not a 500-point checklist. The philosophy is client-first (anchor to the actual pain) and “sometimes the best course of action is to do nothing.”
- Start in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and crawl data, not on-page. On a store, the biggest wins are crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor./indexation problems at scale; “one mistake can keep millions of pages out of the indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..”
- Audit order: crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index. → indexation (GSC Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason.) → duplicate content → on-page at scale → technical (CWVGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data. LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good.<2.5s / INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good.<200ms / CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good.<0.1, HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', canonicals, mobile) → structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. → internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. → out-of-stock → off-page.
- #1 ecommerce issue: duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. from faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals., parameter
permutations, and variants. Fix with robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. blocking or fragment filtering —
not noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.. The Shopify catch: it canonicalizes
/collections/.../products/...but internal links still point at the non-canonical URL (orphans). - #2 issue: thin category pages. Mueller: hard to rank category pages that are just product grids — but a short useful block beats a giant article at the bottom.
- Structured data (Product, Breadcrumb, Org, Review) makes you eligible for rich resultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show.; it doesn’t make you rank.
- Prioritize on impact/effort: quick wins (stray noindex, dirty sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., variant canonicals) first; architecture (facets, CWV, schema rollout) planned; busywork skipped.
- Crawl budgetThe number of URLs an engine will crawl in a timeframe. is a non-issue for most stores — it matters at ~1M+ pages changing weekly or 10k+ daily, or with a big “Discovered – not indexed” bucket.
Official documentation
The primary sources behind each audit step.
Google — crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. & indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.
- Optimize your crawl budget — who needs it, soft-404 waste, why not to use
noindexto save budget. - Managing crawling of faceted navigation URLs — the URL-explosion problem and the robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. / fragment fixes.
- Page indexing report (Search Console Help) — every “not indexed” status, including the duplicate-parameter note.
- URL Inspection tool (Search Console Help) — checking a single URL’s indexed state and structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding..
Google — ecommerce specialty
- Ecommerce URL structure best practices — minimizing alternate URLs, descriptive paths,
&separators, canonical parameter handling. - Help Google understand your ecommerce site structure — the menu → category → sub-category → product hierarchy and
<a href>requirement. - Structured data for ecommerce sites — the recommended schema typesSchema markup is code that uses the schema.org vocabulary to label what your content means so search engines can understand it and show rich results. It's most often written in JSON-LD, and it's not a direct ranking factor..
- Product snippet structured data — required and recommended Product fields.
Google — technical
- Understanding Core Web Vitals and Google Search results — the LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good./INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good./CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good. thresholds and the CWVGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data. report.
- Rich Results Test — validate Product, Breadcrumb, and other markup.
Bing / Microsoft
- bingbot Series: Maximizing Crawl Efficiency — Bing’s “crawl efficiency” framing; relevant to large-catalog stores.
Quotes from the source
On-the-record statements behind the audit. The deep links jump to the quoted passage on the source page.
Me — audit methodology (from my Ahrefs writing)
- “If clients are coming to you asking for an audit, they already have a pain point. Talk to them. Solve that one thing and they’ll be happy with the audit.” — Patrick Stox, Free SEO Audit Template.
- “Anything high-impact and low-effort is a quick win, so those tasks should be tackled first.” — Patrick Stox, Enterprise Technical SEO.
- “One mistake can keep millions of pages out of the indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. or remove an entire site from search results.” — Patrick Stox, Enterprise Technical SEO.
- “Sometimes, the best course of action is to do nothing because the costs outweigh the benefits.” — Patrick Stox, We Studied Over 1 Million Domains.
- “Most sites don’t need to worry about crawl budgetThe number of URLs an engine will crawl in a timeframe., but there are few cases where you may want to take a look.” — Patrick Stox, When Should You Worry About Crawl Budget?
- “Though it’s a joke in the SEO community, ‘it depends’ is really the answer when dealing with out-of-stock products on e-commerce websites.” … “Ultimately, there’s no perfect solution.” — Patrick Stox, How Should You Handle Out-of-Stock Products?.
The four Ahrefs-blog quotes above are from my own published articles; their
#:~:text= deep links resolve on the live pages but a couple were flagged for
browser confirmation during research — worth a spot-check before treating the
fragments as final.
John Mueller, Google Search Advocate — thin category pages (relayed)
- “When the ecommerce category pages don’t have any other content at all, other than links to the products, then it’s really hard for us to rank those pages.”
- “Maybe 90%, 95% of that text is unnecessary. But some amount of text is useful to have on a page so that we can understand what this page is about.”
- “Our algorithms sometimes get confused when they have a list of products on top and essentially a giant article on the bottom.”
Mueller’s quotes are relayed via Ahrefs’ 11 Ways to Improve E-commerce Category Pages, which sourced them from his Search Central office-hours; confirm against the original hangout before quoting as primary.
Google — official docs
- “Eliminate duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. to focus crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. on unique content rather than unique URLs.” Jump to quote
- “Don’t use noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., as Google will still request, but then drop the page when it sees a noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. meta tag or header in the HTTP response, wasting crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. time.” Jump to quote
- “The crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. will typically access a very large number of faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals. URLs before the crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.’ processes determine the URLs are in fact useless.” — Managing crawling of faceted navigation URLs.
- “Use
<a href>tags when creating links to other content. Don’t use JavaScript events on other HTML DOM elements for navigation.” — Help Google understand your ecommerce site structure.
The ecommerce SEO audit checklist
Run top to bottom — the order is the prioritization. Don’t start at on-page.
1. CrawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index.
- Read
robots.txt; nothing important disallowed, low-value parameter spaces blocked. - “Blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing.” report reviewed for false positives.
- GSC Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root). checked for crawled-vs-indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. gaps (crawl waste).
- All URLs exported and clustered by parameter/path to size the faceted problem.
- Crawl-budget concern confirmed real (1M+ weekly / 10k+ daily / big “Discovered”) before acting.
2. Indexation
- GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page IndexingThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. “Not indexed” exported by reason and quantified.
- Indexed count compared against known catalog size.
- URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. run on a sample from each reason bucket.
- XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. contains only canonical, indexable, 200-status URLs.
- No accidental sitewide
noindex.
3. Duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.
- Both GSC “Duplicate” buckets reviewed.
- Duplicate clusters with no canonical identified in the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
- Faceted parameters handled via robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. or
#fragments (not noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.). - Shopify
/collections/*/products/*canonicalized — and the canonical is internally linked. - HTTP/HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', www/non-www, trailing-slash all 301 to one canonical.
4. On-page at scale
- Missing/duplicate titles, H1s, and meta descriptionsThe meta description is an HTML head tag — `<meta name=\"description\" content=\"…\">` — that suggests a short summary of the page for the search snippet. It's not a Google ranking factor, and Google rewrites it the majority of the time, but a good one can still lift click-through. filtered in the crawler.
- Thin category pages flagged; short useful copy block, not a wall of text.
- Title templates checked for collisions.
5. Technical
- CWV by URL group: LCPLargest Contentful Paint — render time of the largest visible image or text block, relative to when the page started loading. ≤2.5 s (at the 75th percentile) is good. < 2.5s, INPInteraction to Next Paint — the input-to-paint latency at the 75th percentile of a page's interactions. ≤200 ms is good. < 200ms, CLSCumulative Layout Shift — a unitless score for unexpected visual movement, taken from the largest burst (session window) of layout shifts, not the lifetime sum. ≤0.1 is good. < 0.1.
- HTTPS everywhere; no mixed contentMixed content is when a page served over HTTPS loads a sub-resource — a script, stylesheet, image, iframe, or similar — over insecure HTTP. Browsers' current taxonomy is upgradable versus blockable; active mixed content (scripts, styles, iframes) is blocked, and passive mixed content (images, audio, video) is warned about or increasingly auto-upgraded, with exceptions like CORS-enabled images and srcset/picture candidates that are blockable, not upgradable..
- Canonicals: no 4XX targets, no non-canonical sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. URLs, no paginated-to-page-one, single tag per page.
- Mobile version isn’t stripped of description/schema/images.
6. Structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding.
- Rich ResultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. Test passes on product + category templates.
- GSC Product Snippets / Merchant Listings reportsThe Merchant listings report is a Google Search Console report (under Shopping) that validates the Product/Offer structured data on your own pages for free merchant listing experiences — Popular Products, Shopping Knowledge Panels, and shopping surfaces in Google Images and Lens. It classifies items as valid or invalid based on critical issues, with non-critical issues flagged as enhancement opportunities, and is distinct from both the Product snippets report and Google Merchant Center's feed diagnostics. clean.
- Merchant listing required set present:
name,image,offers(price/currency/availability). - Product snippet eligibility present:
nameplus at least one ofreview,aggregateRating, oroffers. - Schema prices match visible prices; breadcrumb schema matches visible trail.
7. Internal linkingLinks between pages on the same site.
- Orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. and crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. reviewed; key pages within ~3 clicks.
- Links to redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. and broken links pulled.
- Best sellers linked from nav and hub/editorial pages.
8. Out-of-stock / discontinued
- Soft 404A soft 404 is a URL that returns a success status code (usually 200 OK) even though the page is empty, missing, or shows a 'not found' message. It isn't a status code a server sends — it's a label search engines apply after comparing the response code against the rendered content, and they treat the page like a 404 for indexing. bucket pulled and resolved by pattern.
- Temporary vs. permanent decisions applied (keep / 301 / 404-410).
- Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to deleted products removed or updated.
9. Off-page
- Branded vs. non-branded click share reviewed.
- Links by page type assessed.
- 404s with backlinks 301’d to live pages.
10. Prioritize & deliver
- Every finding scored on impact/effort.
- Quick wins first; architecture planned; busywork skipped.
- Report focused on the few issues that matter, quantified in business impact.
The frameworks behind the audit
1. Scale-first ordering. The audit runs in priority order because leverage differs by area: an on-page tweak fixes one page; an indexation pattern fixes thousands. So: crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index. → indexation → duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. → on-page → technical → structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. → internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. → out-of-stock → off-page. Resist the urge to open with title tagsThe title tag is the HTML title element in a page's head that specifies the document's title. It's the primary source for the SERP title link and a confirmed light ranking factor — but since August 2021 Google doesn't always show it verbatim..
2. The impact/effort matrix. Score every finding on two axes and act by quadrant:
| Low effort | High effort | |
|---|---|---|
| High impact | Quick wins — do first (stray noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., dirty sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., variant canonicals, soft-404 OOS) | Plan & schedule (faceted-nav architecture, CWVGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data., schema rollout, crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops.) |
| Low impact | When time allows (long-tail meta descriptionsThe meta description is an HTML head tag — `<meta name=\"description\" content=\"…\">` — that suggests a short summary of the page for the search snippet. It's not a Google ranking factor, and Google rewrites it the majority of the time, but a good one can still lift click-through., title tidy-ups) | Skip (zero-traffic redirect chainsA → B → C instead of A → C. Each hop loses link equity and adds latency., Open GraphOpen Graph (OG) tags are `<meta>` elements in a page's head, defined by the Open Graph protocol (ogp.me, created by Facebook), that describe a page as a shareable object — its title, description, image, URL, and type. They control how a link preview card looks when the page is shared on Facebook, LinkedIn, Slack, Discord, WhatsApp, and iMessage. They are not a direct Google ranking factor, though Google reads og:title, og:image, and og:site_name as inputs to how a result appears.) |
The matrix’s real value is permission to skip — “sometimes the best course of action is to do nothing.”
3. Client-first scoping. Don’t audit everything. Start from the pain the store actually has (traffic drop, a category that stopped ranking, a bad migration), solve that, then widen. A focused report of 5–10 quantified issues beats a 200-row crawl export every time.
4. The three “not equals” (inherited from technical SEOTechnical SEO is the practice of making a site easy for search engines to crawl, render, index, and (now) be eligible for AI answers. It's the foundation that lets your content and links rank — not a ranking trick of its own.). CrawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. ≠ indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. (a blocked page can still be indexed), crawling ≠ ranking (more crawl ≠ higher positions), crawling ≠ renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. (JS runs separately). Most ecommerce confusion — “why is my product not showing up?” — resolves once you locate which stage it’s failing at.
5. Out-of-stock decision tree. Temporary → keep live (restock/waitlist). Permanent + close match → 301. Permanent + links/traffic, no match → keep with related products. Permanent + nothing → 404/410. “It depends” is the honest default; the variables are permanence and equity.
Patrick's relevant free tools
- Scout Site Audit Free — Run a bounded same-site raw-HTML crawl, group technical SEO findings by detector, and export an honest coverage report.
- SEO Incident Simulator — Practice thirty deterministic technical SEO incident investigations — indexability, crawl controls, redirects, sitemaps, markup, caching, DNS, bot verification, rendering, hreflang, and faceted navigation — with clearly labeled fixture evidence and Find → Fix → Verify handoffs.
- Faceted Navigation Auditor — Classify supplied parameter URLs, surface crawl traps, and advise on facet controls.
Tools for an ecommerce SEO audit
Mentioned in this audit
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. — the backbone: Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (every “not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” reason), Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root)., Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data., SitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version., and Enhancements (Product Snippets, Merchant Listings). Free, and where the audit starts.
- Ahrefs Site Audit — my primary crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.: crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index., indexability, duplicate clusters, canonicals, orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site., internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., performance, and structured-data checks across 170+ issue types.
- Ahrefs Site Explorer — backlink profile by page type, broken pages with links (redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't.-reclamation targets), organic keywords by page, competitor link gaps.
- Google Rich ResultsRich results (formerly 'rich snippets') are enhanced search listings — stars, images, prices, breadcrumbs, video thumbnails, and more — that Google and Bing build from structured data. They're a display feature, not a ranking factor, and eligibility never guarantees they'll show. Test — validate Product, BreadcrumbList, and other schema on a live URL.
- PageSpeed InsightsPageSpeed Insights (PSI) is a free Google tool at pagespeed.web.dev that reports two kinds of data for a URL: real-user field data from the Chrome UX Report and a single Lighthouse lab run with the 0–100 Performance score. Only the field Core Web Vitals are what Google uses for ranking. — field (CrUXChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program.) + lab CWV data on representative templates.
Free options
- Ahrefs Webmaster Tools — free crawl + Site Audit for sites you verify; the no-cost way to get most of the above.
- Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. — crawl info, Site Scan, and IndexNowIndexNow is an open push protocol that lets you instantly tell participating search engines (Bing, Yandex, Naver, Seznam, and Yep) which URLs you've added, changed, or removed via a simple HTTP request — and one submission is shared across all of them. Google does not use it. (push price/stock changes instead of waiting for a recrawl).
Other crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.
- Screaming Frog SEO Spider — desktop bulk export of URLs, status codes, canonicals, meta data, hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.; handy for very large custom crawls.
- Chrome DevTools — confirm whether product content is in the initial HTML or injected by JavaScript (“View source” vs. “Inspect”).
Playbook: organic traffic drops across an ecommerce site
- Confirm the scope. Split Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. data by page type, country, device, and query class. If only one template or market moved, keep the investigation there; if all segments moved, continue sitewide.
- Align the timing. Mark deployments, migrations, feed changes, inventory events, seasonality, and known search changes on the same timeline. If the drop begins with a release, inspect that release before compiling a generic issue list.
- Check access and response changes. Compare current and prior robots rules, status codes, canonicals, sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., and rendered internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. If important URLs are blocked, redirected, noindexed, or orphaned, contain that failure first.
- Reconcile indexation. Join intended URLs from the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing./catalog with crawl results and Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. Page IndexingThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. reasons. If the loss clusters in duplicate or canonical states, inspect URLA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. signals; if it clusters in crawled-not-indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., examine content and page value.
- Test representative templates. Validate one known-good and one affected category, product, facet, and editorial page for renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., schema, internal linkingLinks between pages on the same site., and Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data.. If a shared component fails, widen the sample before changing individual pages.
- Prioritize by affected value and confidence. Ship the smallest reversible fix for the highest-value confirmed cause. If evidence is inconclusive, gather a targeted sample rather than bundling speculative changes.
- Verify and annotate. Record the exact change, test its technical output, and monitor the affected segment against its own baseline. If the intended signal did not change, roll back or reopen the diagnosis.
Audit practices that waste the team’s attention
Export every crawler warning and call it an audit
Why it fails: severity labels do not know the site’s templates, traffic, revenue, or intended URL policy. Do instead: connect each finding to affected URLs, evidence, business impact, and a specific fix owner.
Start with title-tag rewrites during an indexation loss
Why it fails: on-page polish cannot fix blocked, redirected, canonicalized, or unrendered pages. Do instead: establish crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index. and indexation scope before moving down the stack.
Treat every excluded URL as a problem
Why it fails: alternate variants, filtered URLs, and noncanonical duplicates may be intentionally excluded. Do instead: compare the observed state with the documented indexation policy for each URL class.
Change several systems before measuring anything
Why it fails: combined template, canonical, content, and linking changes destroy causal clarity and complicate rollback. Do instead: group related fixes, state the expected signal, and validate each release.
Diagnose audit evidence that does not line up
Crawl totals are far above the catalog size
Likely cause: faceted, sorting, tracking, search, or session parameters are generating URL combinations. Fix: classify the parameter patterns, inspect internal discovery and canonical behavior, and crawl a bounded sample before recommending controls.
Search Console and the crawler disagree on indexability
Likely cause: the crawl sees today’s response while Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. reflects an earlier crawl, or renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. changes directives after raw HTML. Fix: compare timestamps, raw and rendered output, canonical signals, and representative URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. results.
Structured-data errors appear on only some products
Likely cause: optional catalog fields, variant logic, or out-of-stock states enter a different template branch. Fix: segment errors by template and data condition, reproduce one affected record, and correct the shared mapping rather than hand-editing URLs.
Recommendations keep growing but nothing ships
Likely cause: findings lack impact, ownership, dependencies, or acceptance tests. Fix: convert each confirmed issue into a ticket with affected scope, evidence, proposed change, expected signal, owner, and validation step.
Prompts for organizing audit evidence
Cluster findings without inventing severity
Paste a sanitized crawl/issues export after this prompt.
Group these ecommerce SEO findings by root cause and affected template. Preserve the
original evidence and URL counts. For each group, return: observed signal, likely
system owner, evidence still needed, affected page type, reversible first test, and
validation method. Do not assign business impact or severity unless the input
contains traffic, revenue, or indexation evidence supporting it.
[PASTE AUDIT EXPORT]Turn confirmed findings into implementation tickets
Convert only the confirmed findings below into engineering-ready tickets. Each ticket
must include current behavior, intended behavior, affected URL pattern, reproduction
steps, proposed acceptance tests, monitoring window, and rollback condition. Separate
facts from hypotheses and place unresolved questions in a final section.
[PASTE CONFIRMED FINDINGS AND EVIDENCE] Ecommerce audit signal map
| Observed signal | First evidence to inspect | Avoid assuming | Useful next cut |
|---|---|---|---|
| Important pages not discovered | Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. membership, rendered navigation | The sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. alone supplies enough context | Template, depth, orphan status |
| Duplicate/canonical exclusions grow | Canonicals, redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., sitemap URLs, internal-link targets | Every excluded variant should be indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. | URL pattern and product family |
| Indexed pages lose impressions | Query/page cohorts, content changes, inventory, competitors | A crawl warning caused the loss | Category, query intentSemantic search is meaning-based retrieval — matching what a user means, not just the words they typed. Search engines detect entities, expand synonyms, infer intent, and rank by conceptual relevance, which is why keyword stuffing lost its power and topical depth gained it., stock state |
| Product enhancement errors | Visible facts, raw/rendered Product markupProduct schema (schema.org/Product) is structured data that tells search engines a page's product name, price, availability, and reviews so it can appear in Shopping-style rich results. It's separate from a Google Merchant Center feed, though Google reconciles the two., catalog fields | One valid sample proves the template | Error type and data condition |
| Crawl volume explodes | Parameter patterns, facets, calendar/search URLs, logs | More crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. means more indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. | Parameter and botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index./user-agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. |
| Core Web VitalsGoogle's three real-user UX metrics — LCP (loading), INP (responsiveness), and CLS (visual stability) — used by Google's ranking systems, with no official weight attached, measured on field data. regress | CrUXChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program. page groups, template releases, lab traces | One lab score represents the field | Template, device, metric |
| Revenue falls without click loss | Landing-page/checkout behavior, price, stock, analytics | SEO visibility is the cause | Product group and conversion path |
Measure whether the audit program improves the site
Confirmed high-impact findings resolved
Metric: count and share of evidence-backed priority findings shipped and validated, not merely closed. What it tells you: whether audit work reaches production and produces the expected technical signal. How to pull it: join the audit ledger with issue-tracker status and acceptance-test evidence. Benchmark / realistic range: establish a baseline by team capacity and dependency class; do not reward closure of low-value findings to inflate the percentage. Cadence: every sprint and quarterly by root cause.
Intended index coverage by page type
Metric: intended canonical URLsHow search engines pick one canonical URL among duplicates and consolidate signals onto it. represented in the expected indexation state, segmented by product, category, editorial, and approved facet pages. What it tells you: whether the searchable inventory matches policy. How to pull it: reconcile sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing./catalog URLs, crawl states, and Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page IndexingThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. exports. Benchmark / realistic range: the target depends on explicit URL policy; excluded variants should not be counted as failures. Cadence: monthly and after platform releases.
Organic performance of affected cohorts
Metric: clicks, impressions, and qualified organic outcomes for the exact URL/query cohorts tied to shipped fixes. What it tells you: whether technically validated work corresponds with durable search and business improvement. How to pull it: save pre-change Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. cohorts and join them to analytics or commerce outcomes where governance allows. Benchmark / realistic range: compare each cohort with its own seasonal baseline and an unaffected comparison group when possible. Cadence: annotate at release, then review after sufficient recrawling and monthly thereafter.
Recurrence rate
Metric: validated issues that reappear on the same template or URL class after remediation. What it tells you: whether the root cause was fixed or only the current symptoms were patched. How to pull it: compare scheduled crawl detectors and audit ledger fingerprints across runs. Benchmark / realistic range: use the first two comparable audits to set a baseline and treat repeated systemic defects as prevention work. Cadence: each audit cycle.
Test yourself: ecommerce SEO audits
Five questions on audit sequencing, evidence, and prioritization.
Ecommerce SEO Audit
An ecommerce SEO audit is a systematic review of an online store's crawlability, indexation, duplicate content, on-page, technical, and link health — designed to surface the small set of issues that actually hold rankings and revenue back, not to produce a 500-point checklist.
Related: Ecommerce SEO, Faceted Navigation, Crawl Budget, Canonicalization
Ecommerce SEO Audit
An ecommerce SEOEcommerce SEO is the practice of optimizing an online store so its product and category pages rank in organic search and attract purchase-intent visitors. It uses the same Google algorithm as any other site, but compounds the usual SEO work with commerce-specific challenges like faceted navigation, product variants, and platform-imposed URLs. audit is a structured evaluation of an online store’s technical, on-page, content, and off-page health, run to find the issues that keep its product and category pages from being crawled, indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., ranked, and clicked. It uses the same toolkit as any SEO auditAn SEO audit checklist is the structured set of things you review across a site's technical health, on-page elements, content quality, and off-page authority — used to produce a short, prioritized action plan, not an exhaustive 100–200 item inventory. — but ecommerce sites hit those problems at scale: large catalogs, faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals. generating thousands of near-duplicateThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. URLs, product-variant canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it., thin category pages, out-of-stock handling, and Product structured dataProduct schema (schema.org/Product) is structured data that tells search engines a page's product name, price, availability, and reviews so it can appear in Shopping-style rich results. It's separate from a Google Merchant Center feed, though Google reconciles the two..
The output is not a 500-point checklist. The point is to find and fix the stuff that matters most. The biggest wins on a store are usually crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexation problems affecting thousands of URLs at once — which is why an ecommerce audit starts with the Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. and crawl data, not with on-page tweaks. Duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. (from facets, collections, and variants) is the single most common ecommerce-specific issue an audit uncovers; thin category pages are a close second.
Findings get prioritized on an impact/effort basis: high-impact, low-effort fixes (a stray sitewide noindex, a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. full of redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't.) get done first; architectural overhauls get planned; low-value busywork gets skipped.
Related: Ecommerce SEO, Faceted Navigation, Crawl Budget, Canonicalization
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 25, 2026.
Editorial summary and recorded change details.Summary
Aligned legacy tool display names with returntag - hreflang checker and Scout Site Audit Free.
Change details
-
Updated the linked tool names to match their current public labels.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 19, 2026.
Editorial summary and recorded change details.Summary
Fact-check pass against Google's live documentation: fixed a stale faceted-navigation doc URL (it 301s to a new crawling.google.com path), corrected a spliced/inaccurate crawl-budget quote, attributed two previously-unsourced statistics (Gary Illyes' ~60% duplicate-content estimate, my 51.3% multiple-H1 study figure), and rewrote the Product schema audit steps to separate merchant-listing required properties (name, image, offers) from product-snippet eligibility (name plus one of review/aggregateRating/offers) instead of treating them as one required list.
Change details
- Before
Managing crawling of faceted navigation URLs](https://developers.google.com/search/docs/crawling-indexing/crawling-managing-faceted-navigation)AfterUpdated the faceted-navigation documentation link to https://developers.google.com/crawling/docs/faceted-navigation, the URL Google's old doc now 301-redirects to, in both the Official Docs and Quotes lenses. - Before
"Consolidate duplicate content to focus crawling on unique content rather than unique URLs."AfterCorrected the crawl-budget quote from a spliced heading+sentence ("Consolidate duplicate content to focus...") to Google's actual verbatim sentence ("Eliminate duplicate content to focus crawling on unique content rather than unique URLs"), and fixed the matching #:~:text= deep link so it actually anchors. -
Attributed the '60% of the internet is duplicate content' figure to Google's Gary Illyes instead of presenting it as an unsourced fact.
-
Replaced the vague 'about half of sites have them' multiple-H1s claim with the sourced figure (51.3%, my own million-domain study).
-
Rewrote the Product schema audit steps (Step 6 and the Checklists lens) to state Google's actual required-vs-recommended split: merchant listings require name, image, and offers; product snippets require name plus at least one of review, aggregateRating, or offers. The article previously omitted image entirely and implied aggregateRating/review were required.
Full comparison unavailable — no prior snapshot was archived for this revision.