Ecommerce SEO Audit
How I audit an online store — and why I start with Google Search Console and crawl data, not on-page tweaks. A prioritized, scale-first ecommerce SEO audit framework: crawlability, indexation, duplicate content, on-page at scale, technical, structured data, internal links, out-of-stock handling, and off-page — sorted by impact and effort.
An ecommerce SEO audit isn't a 500-point checklist — it's the work of finding the small set of issues holding a store's rankings and revenue back, then prioritizing them by impact and effort. Because ecommerce sites hit problems at scale, I start the audit in Google Search Console (the Page Indexing report) and crawl data, not on-page tweaks: the biggest wins are usually crawling and indexation issues affecting thousands of URLs at once. Duplicate content from facets, collections, and variants is the #1 ecommerce-specific issue audits uncover; thin category pages are #2. The goal is to fix the stuff that matters most and skip the busywork.
Evidence for this claim Search Console's Page indexing report identifies indexed and non-indexed URLs and groups reasons pages are not indexed. Scope: Google Search Console audit data. Confidence: high · Verified: Search Console Help: Page indexing report Evidence for this claim Google's Rich Results Test and product structured-data requirements can be used to validate product markup eligibility. Scope: Google product structured-data validation. Confidence: high · Verified: Google Search Central: Product structured dataTL;DR — An ecommerce SEO audit is a health check for your online store — finding the things that stop Google from showing your product and category pages, and fixing the ones that matter most. The trick: don’t start by tweaking titles and adding keywords. Start in Google Search Console and find out which pages Google can even find and index, because on a store the biggest problems hit thousands of pages at once.
What an ecommerce SEO audit is
An audit is just a structured look at your store to answer one question: what’s holding it back in search? It’s the same idea as a checkup at the doctor — you’re looking for problems before they cost you, and you fix the serious ones first.
The reason ecommerce stores need their own kind of audit is scale. A blog has a few hundred pages. A store can have tens of thousands — one page per product, plus a category page for every way to group them, plus a near-identical copy of each page for every color, size, and sort order. That’s a lot of pages, and a lot of them look almost the same to Google.
Where to start (and where not to)
Most people start an audit by rewriting title tags and stuffing in keywords. That’s backwards. If Google can’t crawl or index a page, no amount of keyword work will help — the page isn’t in the race at all.
So I start in Google Search Console, which is Google’s free dashboard for site owners. The most useful screen is the Page Indexing report: it tells you which of your pages Google has actually indexed, and gives you a reason for every page it hasn’t. On a store, that report usually surfaces the big money problems right away — thousands of duplicate pages, pages Google found but never bothered to crawl, or (worst case) pages accidentally marked “do not index.”
The two problems almost every store has
- Duplicate content. This is the big one. Your filters (color, size, price) create a brand-new web address every time someone clicks one, and most of those addresses show nearly the same products. Variants (the same shirt in five colors) often each get their own page too. Google ends up wading through thousands of near-copies.
- Thin category pages. A category page that’s just a grid of products with no description is hard for Google to rank — there’s not enough on it to tell Google what it’s about.
What you’ll need
- Google Search Console — free, and the single most important tool here.
- A site crawler — something that visits all your pages like Google does and lists the problems. I use Ahrefs Site Audit (this is my site, so that’s the honest answer), but Ahrefs Webmaster Tools gives you a free crawl of sites you verify, and Screaming Frog is a popular desktop option.
- PageSpeed Insights and the Rich Results Test — both free from Google — for checking speed and your product markup.
Want the full, prioritized checklist — every area to audit, in order, with the ecommerce gotchas and how to rank the fixes by impact? Switch to the Advanced tab.
Evidence for this claim Search Console's Page indexing report identifies indexed and non-indexed URLs and groups reasons pages are not indexed. Scope: Google Search Console audit data. Confidence: high · Verified: Search Console Help: Page indexing report Evidence for this claim Google's Rich Results Test and product structured-data requirements can be used to validate product markup eligibility. Scope: Google product structured-data validation. Confidence: high · Verified: Google Search Central: Product structured dataTL;DR — I audit a store in roughly this order: crawlability → indexation → duplicate content → on-page at scale → technical (CWV, HTTPS, canonicals) → structured data → internal links → out-of-stock handling → off-page. I start in the GSC Page Indexing report and a fresh crawl, not on-page, because on a store the highest-leverage problems hit thousands of URLs at once: faceted-nav duplication, accidental sitewide
noindex, “Discovered – currently not indexed” at scale, sitemaps full of redirects. The deliverable is a prioritized impact/effort list, not a 500-point report. Most stores never need to worry about crawl budget — but the ones that do really do.
Audit philosophy: client-first, not issue-first
A crawler will hand you 170+ issue types. If you dump all of them into a report, you’ve produced a document, not an audit. The job is to find the handful that move rankings and revenue and ignore the rest.
The starting point isn’t the tool — it’s the pain. As I’ve put it before: “If clients are coming to you asking for an audit, they already have a pain point. Talk to them. Solve that one thing and they’ll be happy with the audit.” On a store the pain is usually concrete: traffic dropped, a category stopped ranking, a migration went sideways, new products aren’t getting indexed. Anchor the audit to that, then expand.
And know when to stop. From our study of over a million domains: “Sometimes, the best course of action is to do nothing because the costs outweigh the benefits.” Not every flagged issue is worth a developer ticket.
Why I start with Search Console and crawl data, not on-page
The instinct to open the audit with title tags and headings is the most common mistake I see. On an ecommerce site the math is against it: tweaking a title helps one page; fixing an indexation pattern helps thousands. The biggest wins on a store are almost always crawling and indexation problems at scale, and the stakes are highest exactly here — “one mistake can keep millions of pages out of the index or remove an entire site from search results.”
So the audit runs in priority order, top to bottom. Here’s the whole thing.
Step 1 — Crawlability
Can bots reach what matters, and are they wasting their time on what doesn’t?
The ecommerce-specific traps:
- Faceted navigation URL explosion. Filters (color, size, price range) combine into thousands of unique URLs. Google is candid about the cost: “the crawlers will typically access a very large number of faceted navigation URLs before the crawlers’ processes determine the URLs are in fact useless.”
- Parameter variants — session IDs, tracking codes (
?utm_source=), sort orders — multiply the same content across URLs. - Crawl traps — internal search results, infinite calendars, unbounded filter stacking.
robots.txtover- or under-blocking — accidentally blocking CSS/JS needed to render, or a valid product path; or failing to block a junk parameter space.- JS-gated navigation. Google wants real links: “use
<a href>tags when creating links to other content. Don’t use JavaScript events on other HTML DOM elements for navigation.”
Audit steps:
- Fetch and read
robots.txt; confirm nothing important is disallowed and that low-value parameter spaces are. - In the crawler, pull “Blocked by robots.txt” — any important pages caught?
- GSC Crawl Stats: look for spikes in crawled URLs against a flat indexed count — that gap is crawl waste.
- Export all crawled URLs and cluster by path/parameter pattern to size the faceted/parameter problem.
On crawl budget — de-escalate first. Most stores don’t have a crawl-budget problem, and I’ll say so in the audit rather than invent one: “Most sites don’t need to worry about crawl budget, but there are few cases where you may want to take a look.” It becomes real around Google’s thresholds — “Large sites (1 million+ unique pages)…” changing weekly, or 10,000+ pages changing daily — or when “Discovered – currently not indexed” is large. When it is real, you fix it by removing waste, not by asking Google to crawl more. (Deep dive: crawl budget.)
Step 2 — Indexation
Crawlable isn’t indexed. The GSC Page Indexing report is the center of gravity for the whole audit — “see which pages Google can find and index on your site, and learn about any indexing problems encountered.”
The statuses that matter most on a store:
- Duplicate without user-selected canonical / Google chose a different canonical — the faceted/variant duplication problem, made visible.
- Crawled – currently not indexed — usually thin product or near-duplicate filter pages.
- Discovered – currently not indexed — Google knows the URL but hasn’t crawled it; on large catalogs this is a crawl-priority signal.
- URL marked ‘noindex’ — verify it’s intentional. The accidental sitewide version is the “one mistake” scenario.
- Soft 404 — a 200 response on a page that’s effectively empty (out-of-stock, zero-result search). Google warns these “will continue to be crawled, and waste your budget.”
Audit steps:
- Export “Not indexed” URLs by reason; quantify each bucket.
- Compare GSC indexed count against your known catalog size — a large gap means Google isn’t finding pages.
- Run URL Inspection on a sample from each reason bucket.
- Audit the XML sitemap: it should list only canonical, indexable, 200-status URLs. Redirected, noindexed, or canonicalized-away URLs in a sitemap send conflicting signals — pull them.
Step 3 — Duplicate content (the #1 ecommerce issue)
This is the single most common thing an ecommerce audit uncovers, and it’s worth its own step. Google’s Gary Illyes has estimated that roughly 60% of the internet is duplicate content — and stores are overrepresented.
Where it comes from:
- Faceted navigation —
?color=blue,?size=S&color=blue,?size=S&color=blue&sort=priceall serve near-identical content. - Parameter-order permutations —
?color=blue&size=Sand?size=S&color=blueare the same page twice. - Product variants as separate URLs —
?variant=…without a self-referencing canonical or a canonical to the parent. - A product living under multiple category paths, both indexable.
- HTTP/HTTPS, www/non-www, trailing-slash inconsistency that doesn’t 301 to one canonical.
The Shopify gotcha worth calling out by name. Shopify does canonicalize
/collections/{collection}/products/{product} to the clean /products/{product}
URL — so people assume it’s handled. It isn’t fully: the internal links from
collection pages still point at the non-canonical /collections/... version, so
the canonical product URL receives no internal link equity through normal
navigation, and it can show up as an orphan. Fixes: add an “All products” page that
links the canonical /products/ URLs, or edit collection templates to link the
canonical URL directly.
Audit steps:
- GSC: pull both “Duplicate” buckets from Page Indexing.
- Crawler Content Quality / duplicate-cluster report: find clusters with no canonical specified.
- Confirm faceted parameter URLs are handled — Google’s recommended prevention
is
robots.txtblocking of filter parameters, or fragment-based (#) filtering, which “will have no impact on crawling.” Don’t reach fornoindexhere (see Myths). - Spot-check that Shopify
/collections/*/products/*URLs canonicalize to/products/*— and that something links the canonical version.
(Background: duplicate content and canonicalization.)
Step 4 — On-page at scale
On a store, on-page is a templating problem, not a copywriting one — you’re auditing patterns, not pages.
The recurring issues:
- Thin category pages (the #2 ecommerce issue). John Mueller is direct about it: “When the ecommerce category pages don’t have any other content at all, other than links to the products, then it’s really hard for us to rank those pages.” But don’t overcorrect into a wall of text — “maybe 90%, 95% of that text is unnecessary,” and “our algorithms sometimes get confused when they have a list of products on top and essentially a giant article on the bottom.” The target is a short, useful block (sourcing, materials, sizing, popularity) near the top — not filler that buries the grid.
- Templated title collisions —
Buy {Product} | Storepatterns that go near-identical across products, or two products sharing a name. - Missing/empty H1s and titles at the template level.
- Meta descriptions — worth doing on your top category and product pages, less worth it on the long tail (Google rewrites them most of the time anyway).
Audit steps:
- Crawler On-Page report: filter for missing title, missing H1, missing meta description, duplicate titles, duplicate H1s.
- Flag category pages with little to no unique body copy (thin-content candidates) and cross-reference against “Crawled – currently not indexed.”
- Export titles and check for template collisions.
A note on what not to spend audit time on: multiple H1s. They’re valid HTML5 and near-irrelevant to ranking — 51.3% of sites have them somewhere, per my own million-domain study. Skip it.
Step 5 — Technical (CWV, HTTPS, canonicals, mobile)
Core Web Vitals. Google’s passing thresholds: LCP < 2.5s, INP < 200ms, CLS < 0.1. Ecommerce failure modes are predictable — unoptimized hero images and third-party scripts (chat, A/B, pixels) hurt LCP; add-to-cart and checkout JS hurt INP; images without explicit dimensions and injected promo bars hurt CLS. Audit by URL group in the GSC CWV report so you can see whether it’s product, category, or checkout templates failing, then confirm on representative templates in PageSpeed Insights. (See Core Web Vitals.)
HTTPS. Everything — product, cart, checkout — on HTTPS, no mixed content (HTTP images/resources on HTTPS pages), and canonicals/redirects pointing at the HTTPS version.
Canonicals. The common ecommerce errors: canonicals pointing at 4XX pages;
non-canonical URLs in the sitemap; paginated pages canonicalized to page one
(don’t); and — the quiet one — internal links pointing at non-canonical versions so
the canonical never accumulates link equity. Multiple rel=canonical tags on one
page get ignored.
Mobile-first. Google indexes the mobile version. Make sure the mobile product page isn’t a stripped-down one missing the description, the structured data, or full-size images.
Raw versus rendered product evidence. On representative PDPs, save the initial HTML and a rendered capture after a fresh navigation. Do not treat resizing an already-loaded desktop page as a mobile-rendering test. The raw response should expose the core product identity, selected/default SKU and attributes, price, currency, availability, crawlable variant links where required, and matching Product/Offer data. Then compare the rendered DOM and visible selection. Google may render JavaScript, but other crawlers and agents vary; a successful render also does not cure contradictory product facts.
For each sampled variant, continue the comparison through the feed, cart, and checkout. The URL, visible PDP, rendered JSON-LD, feed item, and transaction should agree on product identity, SKU, group ID, selected attributes, price, currency, and availability. Record intentional location- or customer-specific revalidation separately from unexplained mismatches. See Product Page SEO for the raw/rendered boundary and Product Variant SEO for the full selected-offer contract.
Step 6 — Structured data
For a store, the priority types are Product / ProductGroup, BreadcrumbList, Organization (with return policy), and Review / AggregateRating. Adding more valid properties widens eligibility — Google: “adding the more properties you can add, the more enhancements your page can be eligible for.” (The one myth to keep in your head: schema makes you eligible for rich results; it doesn’t make you rank.)
Audit steps:
- Run the Rich Results Test on a representative product and category page.
- GSC Enhancements: check Product Snippets and Merchant Listings for errors/warnings.
- Check required vs. recommended separately per feature — don’t collapse them.
For merchant listings, Google’s required set is
name,image, andoffers(anOffer, withprice,priceCurrency, andavailability). For product snippets,nameis required and you need at least one ofreview,aggregateRating, oroffersto be eligible — Google listsaggregateRating,offers, andreviewas recommended, not required, on that feature. - Common faults: missing
image, missingoffers, missingavailability, schema prices that don’t match the visible price (a policy violation), breadcrumb schema that doesn’t match the visible trail.
Step 7 — Internal linking
The issues: orphan product pages (classic Shopify symptom, above); important category/product pages buried 5+ clicks deep; internal links pointing at redirects or 404s; best sellers not linked from high-authority hubs. Google: “add links from menus to category pages, from category pages to sub-category pages, and finally from sub-category pages to all product pages.”
Audit steps:
- Crawler Links report: pull orphan pages and crawl depth; confirm key pages sit within ~3 clicks of the homepage.
- Pull links to redirects and broken links.
- Check that your highest-revenue products are linked from the main nav and editorial/hub pages.
Step 8 — Out-of-stock and discontinued products
A genuinely ecommerce-only section. My honest answer here is the SEO cliché for a reason: “though it’s a joke in the SEO community, ‘it depends’ is really the answer when dealing with out-of-stock products.” And “ultimately, there’s no perfect solution.” The decision turns on whether it’s temporary vs. permanent and whether the page has traffic or links:
| Scenario | What I do | Why |
|---|---|---|
| Temporarily OOS | Keep the page live | Add restock dates, waitlist, notify-me; don’t throw away ranking |
| Permanently discontinued, close replacement | 301 to the similar product | Preserves link equity if they’re genuinely similar |
| Discontinued, no match, has links/traffic | Keep live with “related products” | Retains ranking potential; route users onward |
| Discontinued, no links/traffic | 404 or 410 | Clean it up; fix internal links pointing at it |
Do not audit availability from the visible label alone. Sample an ordinary in-stock SKU, a temporary stockout, a recently restocked item, one unavailable variant inside an available product group, and a postcode-restricted offer. For each, reconcile the authoritative backend state, selected variant, visible page, Product/Offer markup, merchant feed row, cart line, and checkout outcome. Record the market, postcode, channel, collection time, and feed-processing time so a personalized fulfillment result is not mistaken for the catalog’s general state. The full mapping is in the Product Schema availability contract.
Watch for soft 404s: out-of-stock pages, zero-result searches, or empty cart/account pages returning 200. Pull the Soft 404 bucket in GSC and resolve by pattern. (See out-of-stock products for the full framework.)
Step 9 — Off-page
Lighter for most stores, but worth a pass:
- GSC Performance: branded vs. non-branded click share — heavy branded reliance means weak non-branded discovery.
- Backlink profile by page type — are category/product pages earning links, or only the homepage and blog?
- Find 404 pages that have backlinks and 301 them to the closest live page (reclaim the equity).
- Competitor link gap on category pages.
How to prioritize the findings
Everything above produces a list. The list is not the deliverable — the sorted list is. I score each finding on an impact/effort matrix: “anything high-impact and low-effort is a quick win, so those tasks should be tackled first.” For a store:
- High impact / low effort — do first: a stray sitewide
noindex; a sitemap full of redirect/noindex URLs; missing canonicals on variant URLs; internal links pointing at redirects; out-of-stock soft 404s. - High impact / high effort — plan and schedule: faceted-nav architecture; Core Web Vitals work; rolling Product schema across templates; crawl-depth / architecture fixes.
- Low impact / low effort — when time allows: meta descriptions on long-tail products; minor title-template tidy-ups.
- Low impact / high effort — skip: redirect-chain cleanup on zero-traffic pages; Open Graph tags (social, not ranking).
The whole point of the matrix is permission to not do things.
An illustrative cohort matrix compares thin content, non-canonical pages, deep URLs, and schema errors. Product pages score 22, 11, 36, and 48 percent; category pages 18, 8, 54, and 12 percent; facet URLs 71, 83, 64, and 5 percent; blog pages 9, 3, 14, and 2 percent. These are synthetic rates, not customer or site data.
Common myths to clear up in the report
- “More indexed pages = better SEO.” No — Google’s own advice is to “eliminate duplicate content to focus crawling on unique content rather than unique URLs.” Inflating the index with thin variant pages usually hurts.
- “Use
noindexto save crawl budget on facets.” No: “don’t use noindex, as Google will still request, but then drop the page… wasting crawling time.” Block the fetch (robots.txt) or use fragment filtering instead. - “Shopify handles all canonicalization.” It sets the canonical tag, but it doesn’t fix the internal-linking-to-non-canonical problem. See Step 3.
- “Crawl budget affects every store.” It doesn’t — most stores never need to think about it.
AI summary
A condensed take on the Advanced version:
- An ecommerce SEO audit finds the few issues holding a store back — it’s a prioritized list, not a 500-point checklist. The philosophy is client-first (anchor to the actual pain) and “sometimes the best course of action is to do nothing.”
- Start in Google Search Console and crawl data, not on-page. On a store, the biggest wins are crawling/indexation problems at scale; “one mistake can keep millions of pages out of the index.”
- Audit order: crawlability → indexation (GSC Page Indexing report) → duplicate content → on-page at scale → technical (CWV LCP<2.5s / INP<200ms / CLS<0.1, HTTPS, canonicals, mobile) → structured data → internal links → out-of-stock → off-page.
- #1 ecommerce issue: duplicate content from faceted navigation, parameter
permutations, and variants. Fix with robots.txt blocking or fragment filtering —
not noindex. The Shopify catch: it canonicalizes
/collections/.../products/...but internal links still point at the non-canonical URL (orphans). - #2 issue: thin category pages. Mueller: hard to rank category pages that are just product grids — but a short useful block beats a giant article at the bottom.
- Structured data (Product, Breadcrumb, Org, Review) makes you eligible for rich results; it doesn’t make you rank.
- Prioritize on impact/effort: quick wins (stray noindex, dirty sitemaps, variant canonicals) first; architecture (facets, CWV, schema rollout) planned; busywork skipped.
- Crawl budget is a non-issue for most stores — it matters at ~1M+ pages changing weekly or 10k+ daily, or with a big “Discovered – not indexed” bucket.
Official documentation
The primary sources behind each audit step.
Google — crawling & indexing
- Optimize your crawl budget — who needs it, soft-404 waste, why not to use
noindexto save budget. - Managing crawling of faceted navigation URLs — the URL-explosion problem and the robots.txt / fragment fixes.
- Page indexing report (Search Console Help) — every “not indexed” status, including the duplicate-parameter note.
- URL Inspection tool (Search Console Help) — checking a single URL’s indexed state and structured data.
Google — ecommerce specialty
- Ecommerce URL structure best practices — minimizing alternate URLs, descriptive paths,
&separators, canonical parameter handling. - Help Google understand your ecommerce site structure — the menu → category → sub-category → product hierarchy and
<a href>requirement. - Structured data for ecommerce sites — the recommended schema types.
- Product snippet structured data — required and recommended Product fields.
Google — technical
- Understanding Core Web Vitals and Google Search results — the LCP/INP/CLS thresholds and the CWV report.
- Rich Results Test — validate Product, Breadcrumb, and other markup.
Bing / Microsoft
- bingbot Series: Maximizing Crawl Efficiency — Bing’s “crawl efficiency” framing; relevant to large-catalog stores.
Quotes from the source
On-the-record statements behind the audit. The deep links jump to the quoted passage on the source page.
Me — audit methodology (from my Ahrefs writing)
- “If clients are coming to you asking for an audit, they already have a pain point. Talk to them. Solve that one thing and they’ll be happy with the audit.” — Patrick Stox, Free SEO Audit Template.
- “Anything high-impact and low-effort is a quick win, so those tasks should be tackled first.” — Patrick Stox, Enterprise Technical SEO.
- “One mistake can keep millions of pages out of the index or remove an entire site from search results.” — Patrick Stox, Enterprise Technical SEO.
- “Sometimes, the best course of action is to do nothing because the costs outweigh the benefits.” — Patrick Stox, We Studied Over 1 Million Domains.
- “Most sites don’t need to worry about crawl budget, but there are few cases where you may want to take a look.” — Patrick Stox, When Should You Worry About Crawl Budget?
- “Though it’s a joke in the SEO community, ‘it depends’ is really the answer when dealing with out-of-stock products on e-commerce websites.” … “Ultimately, there’s no perfect solution.” — Patrick Stox, How Should You Handle Out-of-Stock Products?.
The four Ahrefs-blog quotes above are from my own published articles; their
#:~:text= deep links resolve on the live pages but a couple were flagged for
browser confirmation during research — worth a spot-check before treating the
fragments as final.
John Mueller, Google Search Advocate — thin category pages (relayed)
- “When the ecommerce category pages don’t have any other content at all, other than links to the products, then it’s really hard for us to rank those pages.”
- “Maybe 90%, 95% of that text is unnecessary. But some amount of text is useful to have on a page so that we can understand what this page is about.”
- “Our algorithms sometimes get confused when they have a list of products on top and essentially a giant article on the bottom.”
Mueller’s quotes are relayed via Ahrefs’ 11 Ways to Improve E-commerce Category Pages, which sourced them from his Search Central office-hours; confirm against the original hangout before quoting as primary.
Google — official docs
- “Eliminate duplicate content to focus crawling on unique content rather than unique URLs.” Jump to quote
- “Don’t use noindex, as Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time.” Jump to quote
- “The crawlers will typically access a very large number of faceted navigation URLs before the crawlers’ processes determine the URLs are in fact useless.” — Managing crawling of faceted navigation URLs.
- “Use
<a href>tags when creating links to other content. Don’t use JavaScript events on other HTML DOM elements for navigation.” — Help Google understand your ecommerce site structure.
The ecommerce SEO audit checklist
Run top to bottom — the order is the prioritization. Don’t start at on-page.
1. Crawlability
- Read
robots.txt; nothing important disallowed, low-value parameter spaces blocked. - “Blocked by robots.txt” report reviewed for false positives.
- GSC Crawl Stats checked for crawled-vs-indexed gaps (crawl waste).
- All URLs exported and clustered by parameter/path to size the faceted problem.
- Crawl-budget concern confirmed real (1M+ weekly / 10k+ daily / big “Discovered”) before acting.
2. Indexation
- GSC Page Indexing “Not indexed” exported by reason and quantified.
- Indexed count compared against known catalog size.
- URL Inspection run on a sample from each reason bucket.
- XML sitemap contains only canonical, indexable, 200-status URLs.
- No accidental sitewide
noindex.
3. Duplicate content
- Both GSC “Duplicate” buckets reviewed.
- Duplicate clusters with no canonical identified in the crawler.
- Faceted parameters handled via robots.txt or
#fragments (not noindex). - Shopify
/collections/*/products/*canonicalized — and the canonical is internally linked. - HTTP/HTTPS, www/non-www, trailing-slash all 301 to one canonical.
4. On-page at scale
- Missing/duplicate titles, H1s, and meta descriptions filtered in the crawler.
- Thin category pages flagged; short useful copy block, not a wall of text.
- Title templates checked for collisions.
5. Technical
- CWV by URL group: LCP < 2.5s, INP < 200ms, CLS < 0.1.
- HTTPS everywhere; no mixed content.
- Canonicals: no 4XX targets, no non-canonical sitemap URLs, no paginated-to-page-one, single tag per page.
- Mobile version isn’t stripped of description/schema/images.
6. Structured data
- Rich Results Test passes on product + category templates.
- GSC Product Snippets / Merchant Listings reports clean.
- Merchant listing required set present:
name,image,offers(price/currency/availability). - Product snippet eligibility present:
nameplus at least one ofreview,aggregateRating, oroffers. - Schema prices match visible prices; breadcrumb schema matches visible trail.
7. Internal linking
- Orphan pages and crawl depth reviewed; key pages within ~3 clicks.
- Links to redirects and broken links pulled.
- Best sellers linked from nav and hub/editorial pages.
8. Out-of-stock / discontinued
- Soft 404 bucket pulled and resolved by pattern.
- Temporary vs. permanent decisions applied (keep / 301 / 404-410).
- Internal links to deleted products removed or updated.
9. Off-page
- Branded vs. non-branded click share reviewed.
- Links by page type assessed.
- 404s with backlinks 301’d to live pages.
10. Prioritize & deliver
- Every finding scored on impact/effort.
- Quick wins first; architecture planned; busywork skipped.
- Report focused on the few issues that matter, quantified in business impact.
The frameworks behind the audit
1. Scale-first ordering. The audit runs in priority order because leverage differs by area: an on-page tweak fixes one page; an indexation pattern fixes thousands. So: crawlability → indexation → duplicate content → on-page → technical → structured data → internal links → out-of-stock → off-page. Resist the urge to open with title tags.
2. The impact/effort matrix. Score every finding on two axes and act by quadrant:
| Low effort | High effort | |
|---|---|---|
| High impact | Quick wins — do first (stray noindex, dirty sitemap, variant canonicals, soft-404 OOS) | Plan & schedule (faceted-nav architecture, CWV, schema rollout, crawl depth) |
| Low impact | When time allows (long-tail meta descriptions, title tidy-ups) | Skip (zero-traffic redirect chains, Open Graph) |
The matrix’s real value is permission to skip — “sometimes the best course of action is to do nothing.”
3. Client-first scoping. Don’t audit everything. Start from the pain the store actually has (traffic drop, a category that stopped ranking, a bad migration), solve that, then widen. A focused report of 5–10 quantified issues beats a 200-row crawl export every time.
4. The three “not equals” (inherited from technical SEO). Crawling ≠ indexing (a blocked page can still be indexed), crawling ≠ ranking (more crawl ≠ higher positions), crawling ≠ rendering (JS runs separately). Most ecommerce confusion — “why is my product not showing up?” — resolves once you locate which stage it’s failing at.
5. Out-of-stock decision tree. Temporary → keep live (restock/waitlist). Permanent + close match → 301. Permanent + links/traffic, no match → keep with related products. Permanent + nothing → 404/410. “It depends” is the honest default; the variables are permanence and equity.
Patrick's relevant free tools
- Scout Site Audit Free — Run a bounded same-site raw-HTML crawl, compare detector and inferred-cohort changes across snapshots or sites, and export provenance-rich CSV and finding-packet reports.
- Crawl Export Comparator — Normalize two Screaming Frog, Sitebulb, Scout, or generic crawl exports and review added, removed, and field-changed URLs by path template without uploading the files.
- SEO Incident Simulator — Practice thirty deterministic technical SEO incident investigations — indexability, crawl controls, redirects, sitemaps, markup, caching, DNS, bot verification, rendering, hreflang, and faceted navigation — with clearly labeled fixture evidence and Find → Fix → Verify handoffs.
Tools for an ecommerce SEO audit
Mentioned in this audit
- Google Search Console — the backbone: Page Indexing report (every “not indexed” reason), Crawl Stats, Core Web Vitals, Sitemaps, URL Inspection, and Enhancements (Product Snippets, Merchant Listings). Free, and where the audit starts.
- Ahrefs Site Audit — my primary crawler: crawlability, indexability, duplicate clusters, canonicals, orphan pages, internal links, crawl depth, performance, and structured-data checks across 170+ issue types.
- Ahrefs Site Explorer — backlink profile by page type, broken pages with links (redirect-reclamation targets), organic keywords by page, competitor link gaps.
- Google Rich Results Test — validate Product, BreadcrumbList, and other schema on a live URL.
- PageSpeed Insights — field (CrUX) + lab CWV data on representative templates.
Free options
- Ahrefs Webmaster Tools — free crawl + Site Audit for sites you verify; the no-cost way to get most of the above.
- Bing Webmaster Tools — crawl info, Site Scan, and IndexNow (push price/stock changes instead of waiting for a recrawl).
Other crawlers
- Screaming Frog SEO Spider — desktop bulk export of URLs, status codes, canonicals, meta data, hreflang; handy for very large custom crawls.
- Chrome DevTools — confirm whether product content is in the initial HTML or injected by JavaScript (“View source” vs. “Inspect”).
Playbook: organic traffic drops across an ecommerce site
- Confirm the scope. Split Search Console data by page type, country, device, and query class. If only one template or market moved, keep the investigation there; if all segments moved, continue sitewide.
- Align the timing. Mark deployments, migrations, feed changes, inventory events, seasonality, and known search changes on the same timeline. If the drop begins with a release, inspect that release before compiling a generic issue list.
- Check access and response changes. Compare current and prior robots rules, status codes, canonicals, sitemaps, and rendered internal links. If important URLs are blocked, redirected, noindexed, or orphaned, contain that failure first.
- Reconcile indexation. Join intended URLs from the sitemap/catalog with crawl results and Search Console Page Indexing reasons. If the loss clusters in duplicate or canonical states, inspect URL signals; if it clusters in crawled-not-indexed, examine content and page value.
- Test representative templates. Validate one known-good and one affected category, product, facet, and editorial page for rendering, schema, internal linking, and Core Web Vitals. If a shared component fails, widen the sample before changing individual pages.
- Prioritize by affected value and confidence. Ship the smallest reversible fix for the highest-value confirmed cause. If evidence is inconclusive, gather a targeted sample rather than bundling speculative changes.
- Verify and annotate. Record the exact change, test its technical output, and monitor the affected segment against its own baseline. If the intended signal did not change, roll back or reopen the diagnosis.
Audit practices that waste the team’s attention
Export every crawler warning and call it an audit
Why it fails: severity labels do not know the site’s templates, traffic, revenue, or intended URL policy. Do instead: connect each finding to affected URLs, evidence, business impact, and a specific fix owner.
Start with title-tag rewrites during an indexation loss
Why it fails: on-page polish cannot fix blocked, redirected, canonicalized, or unrendered pages. Do instead: establish crawlability and indexation scope before moving down the stack.
Treat every excluded URL as a problem
Why it fails: alternate variants, filtered URLs, and noncanonical duplicates may be intentionally excluded. Do instead: compare the observed state with the documented indexation policy for each URL class.
Change several systems before measuring anything
Why it fails: combined template, canonical, content, and linking changes destroy causal clarity and complicate rollback. Do instead: group related fixes, state the expected signal, and validate each release.
Diagnose audit evidence that does not line up
Use the shared SEO audit evidence package for capture, timestamp, environment, scope, limitations, reproduction, and acceptance tests. Extend it here with product ID, SKU, selected variant, market, postcode, inventory/fulfillment state, feed source and processing time. Keep the original catalog, page, and feed captures separate from the analyst’s diagnosis.
Crawl totals are far above the catalog size
Likely cause: faceted, sorting, tracking, search, or session parameters are generating URL combinations. Fix: classify the parameter patterns, inspect internal discovery and canonical behavior, and crawl a bounded sample before recommending controls.
Search Console and the crawler disagree on indexability
Likely cause: the crawl sees today’s response while Search Console reflects an earlier crawl, or rendering changes directives after raw HTML. Fix: compare timestamps, raw and rendered output, canonical signals, and representative URL Inspection results.
Structured-data errors appear on only some products
Likely cause: optional catalog fields, variant logic, or out-of-stock states enter a different template branch. Fix: segment errors by template and data condition, reproduce one affected record, and correct the shared mapping rather than hand-editing URLs.
Recommendations keep growing but nothing ships
Likely cause: findings lack impact, ownership, dependencies, or acceptance tests. Fix: convert each confirmed issue into a ticket with affected scope, evidence, proposed change, expected signal, owner, and validation step.
Prompts for organizing audit evidence
Cluster findings without inventing severity
Paste a sanitized crawl/issues export after this prompt.
Group these ecommerce SEO findings by root cause and affected template. Preserve the
original evidence and URL counts. For each group, return: observed signal, likely
system owner, evidence still needed, affected page type, reversible first test, and
validation method. Do not assign business impact or severity unless the input
contains traffic, revenue, or indexation evidence supporting it.
[PASTE AUDIT EXPORT]Turn confirmed findings into implementation tickets
Convert only the confirmed findings below into engineering-ready tickets. Each ticket
must include current behavior, intended behavior, affected URL pattern, reproduction
steps, proposed acceptance tests, monitoring window, and rollback condition. Separate
facts from hypotheses and place unresolved questions in a final section.
[PASTE CONFIRMED FINDINGS AND EVIDENCE] Ecommerce audit signal map
| Observed signal | First evidence to inspect | Avoid assuming | Useful next cut |
|---|---|---|---|
| Important pages not discovered | Internal links, sitemap membership, rendered navigation | The sitemap alone supplies enough context | Template, depth, orphan status |
| Duplicate/canonical exclusions grow | Canonicals, redirects, sitemap URLs, internal-link targets | Every excluded variant should be indexed | URL pattern and product family |
| Indexed pages lose impressions | Query/page cohorts, content changes, inventory, competitors | A crawl warning caused the loss | Category, query intent, stock state |
| Product enhancement errors | Visible facts, raw/rendered Product markup, catalog fields | One valid sample proves the template | Error type and data condition |
| Crawl volume explodes | Parameter patterns, facets, calendar/search URLs, logs | More crawling means more indexing | Parameter and bot/user-agent |
| Core Web Vitals regress | CrUX page groups, template releases, lab traces | One lab score represents the field | Template, device, metric |
| Revenue falls without click loss | Landing-page/checkout behavior, price, stock, analytics | SEO visibility is the cause | Product group and conversion path |
Measure whether the audit program improves the site
Confirmed high-impact findings resolved
Metric: count and share of evidence-backed priority findings shipped and validated, not merely closed. What it tells you: whether audit work reaches production and produces the expected technical signal. How to pull it: join the audit ledger with issue-tracker status and acceptance-test evidence. Benchmark / realistic range: establish a baseline by team capacity and dependency class; do not reward closure of low-value findings to inflate the percentage. Cadence: every sprint and quarterly by root cause.
Intended index coverage by page type
Metric: intended canonical URLs represented in the expected indexation state, segmented by product, category, editorial, and approved facet pages. What it tells you: whether the searchable inventory matches policy. How to pull it: reconcile sitemap/catalog URLs, crawl states, and Search Console Page Indexing exports. Benchmark / realistic range: the target depends on explicit URL policy; excluded variants should not be counted as failures. Cadence: monthly and after platform releases.
Organic performance of affected cohorts
Metric: clicks, impressions, and qualified organic outcomes for the exact URL/query cohorts tied to shipped fixes. What it tells you: whether technically validated work corresponds with durable search and business improvement. How to pull it: save pre-change Search Console cohorts and join them to analytics or commerce outcomes where governance allows. Benchmark / realistic range: compare each cohort with its own seasonal baseline and an unaffected comparison group when possible. Cadence: annotate at release, then review after sufficient recrawling and monthly thereafter.
Recurrence rate
Metric: validated issues that reappear on the same template or URL class after remediation. What it tells you: whether the root cause was fixed or only the current symptoms were patched. How to pull it: compare scheduled crawl detectors and audit ledger fingerprints across runs. Benchmark / realistic range: use the first two comparable audits to set a baseline and treat repeated systemic defects as prevention work. Cadence: each audit cycle.
Test yourself: ecommerce SEO audits
Five questions on audit sequencing, evidence, and prioritization.
Ecommerce SEO Audit
An ecommerce SEO audit is a systematic review of an online store's crawlability, indexation, duplicate content, on-page, technical, and link health — designed to surface the small set of issues that actually hold rankings and revenue back, not to produce a 500-point checklist.
Related: Ecommerce SEO, Faceted Navigation, Crawl Budget, Canonicalization
Ecommerce SEO Audit
An ecommerce SEO audit is a structured evaluation of an online store’s technical, on-page, content, and off-page health, run to find the issues that keep its product and category pages from being crawled, indexed, ranked, and clicked. It uses the same toolkit as any SEO audit — but ecommerce sites hit those problems at scale: large catalogs, faceted navigation generating thousands of near-duplicate URLs, product-variant canonicalization, thin category pages, out-of-stock handling, and Product structured data.
The output is not a 500-point checklist. The point is to find and fix the stuff that matters most. The biggest wins on a store are usually crawling and indexation problems affecting thousands of URLs at once — which is why an ecommerce audit starts with the Google Search Console Page Indexing report and crawl data, not with on-page tweaks. Duplicate content (from facets, collections, and variants) is the single most common ecommerce-specific issue an audit uncovers; thin category pages are a close second.
Findings get prioritized on an impact/effort basis: high-impact, low-effort fixes (a stray sitewide noindex, a sitemap full of redirects) get done first; architectural overhauls get planned; low-value busywork gets skipped.
Related: Ecommerce SEO, Faceted Navigation, Crawl Budget, Canonicalization
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 29, 2026.
Editorial summary and recorded change details.Summary
Added raw-versus-rendered PDP testing and a selected-offer parity check to the ecommerce audit.
Change details
-
Required fresh-navigation comparisons across raw HTML, rendered DOM, visible selection, schema, feed, cart, and checkout.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 29, 2026.
Editorial summary and recorded change details.Summary
Added ecommerce-specific applications of the shared audit-evidence and availability contracts.
Change details
-
Added representative SKU/market/fulfillment reconciliation across page, markup, feed, cart, and checkout.
-
Linked ecommerce findings to the shared reproducible evidence package and defined the catalog-specific extension fields.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 28, 2026.
Editorial summary and recorded change details.Summary
Added a relationship disclosure beside the article's explicit Ahrefs Site Audit recommendation.
Change details
-
Disclosed that Patrick is a former Ahrefs employee and receives complimentary Ahrefs access.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 27, 2026.
Editorial summary and recorded change details.Summary
Added a page-segment issue matrix so ecommerce findings can be prioritized by affected template rather than by isolated URLs.
Change details
-
Added an illustrative product/category/facet/blog cohort matrix to the Advanced lens.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 25, 2026.
Editorial summary and recorded change details.Summary
Aligned legacy tool display names with returntag and Scout Site Audit Free.
Change details
-
Updated the linked tool names to match their current public labels.
Full comparison unavailable — no prior snapshot was archived for this revision.
Updated Jul 19, 2026.
Editorial summary and recorded change details.Summary
Fact-check pass against Google's live documentation: fixed a stale faceted-navigation doc URL (it 301s to a new crawling.google.com path), corrected a spliced/inaccurate crawl-budget quote, attributed two previously-unsourced statistics (Gary Illyes' ~60% duplicate-content estimate, my 51.3% multiple-H1 study figure), and rewrote the Product schema audit steps to separate merchant-listing required properties (name, image, offers) from product-snippet eligibility (name plus one of review/aggregateRating/offers) instead of treating them as one required list.
Change details
- Before
Managing crawling of faceted navigation URLs](https://developers.google.com/search/docs/crawling-indexing/crawling-managing-faceted-navigation)AfterUpdated the faceted-navigation documentation link to https://developers.google.com/crawling/docs/faceted-navigation, the URL Google's old doc now 301-redirects to, in both the Official Docs and Quotes lenses. - Before
"Consolidate duplicate content to focus crawling on unique content rather than unique URLs."AfterCorrected the crawl-budget quote from a spliced heading+sentence ("Consolidate duplicate content to focus...") to Google's actual verbatim sentence ("Eliminate duplicate content to focus crawling on unique content rather than unique URLs"), and fixed the matching #:~:text= deep link so it actually anchors. -
Attributed the '60% of the internet is duplicate content' figure to Google's Gary Illyes instead of presenting it as an unsourced fact.
-
Replaced the vague 'about half of sites have them' multiple-H1s claim with the sourced figure (51.3%, my own million-domain study).
-
Rewrote the Product schema audit steps (Step 6 and the Checklists lens) to state Google's actual required-vs-recommended split: merchant listings require name, image, and offers; product snippets require name plus at least one of review, aggregateRating, or offers. The article previously omitted image entirely and implied aggregateRating/review were required.
Full comparison unavailable — no prior snapshot was archived for this revision.