Programmatic SEO
What programmatic SEO actually is, when it works and when it's spam, and how to build pages at scale that get indexed — from Patrick Stox.
1 evidence signal on this page
- Related live toolGoogle Index Checker
Programmatic SEO (pSEO) is one template plus a data source generating many pages for similar queries. It's legitimate when each page genuinely answers its query with unique data, and it's spam when it stamps a thin template across a shallow dataset — which is what trips Google's scaled-content-abuse and doorway policies (and now Bing's). My contrarian take: thin content at scale is a data problem, not a template problem. And the first thing that actually decides whether any of it works is indexing — publish in staged batches, validate indexation and impressions before you scale, and treat crawl budget, internal linking, sitemaps, and index bloat as first-class. The people trying to fully automate this aren't doing well; the ones winning have proprietary data and real oversight.
TL;DR — Programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. is building lots of pages from one template and a pile of data instead of writing each page by hand. Done right (think currency converters or “things to do in [city]” pages backed by real data) it can publish useful coverage efficiently. Done mainly to manipulate rankings with thin permutations, it can violate scaled-content or doorway-abuse policies. Evidence for this claim Google defines scaled content abuse as generating many pages primarily to manipulate rankings, regardless of whether automation, humans, or both created them. Scope: Current Google spam policy; scale itself is not the violation. Confidence: high · Verified: Google Search Essentials: Scaled content abuse Evidence for this claim Programmatic pages should provide original value for an intended audience rather than thin permutations created mainly for search traffic. Scope: Current Google helpful-content self-assessment. Confidence: high · Verified: Google Search Central: Creating helpful content
What programmatic SEO is
Programmatic SEO — people often shorten it to pSEO — means generating a large number of pages from a single template plus a data source, rather than writing each one individually. You design the page layout once, point it at a spreadsheet or database, and it fills in thousands of variations.
The core idea is simple: one template + one good dataset = many pages, each aimed at a slightly different search. If you’ve ever searched and landed on a page like:
- Wise — currency converter pages (one for every “[currency] to [currency]” pair). By some estimates Wise runs millions of these pages.
- Zapier — “connect [App A] to [App B]” integration pages, reportedly hundreds of thousands of them.
- Zillow — a page for essentially every property listing.
…you’ve used programmatic SEO. (Those page counts and traffic numbers are third-party estimates, so treat them as ballpark, not gospel.) The reason these work is that every page has something genuinely useful and different on it — a live exchange rate, a real integration, an actual listing.
When it’s great vs. when it’s spam
Here’s the honest part most “how to do pSEO” guides rush past: the technique is neutral. It scales good pages and bad pages equally well.
- Great: each page answers a real question with data a person actually wants.
- Spam: the only thing that changes between pages is a word in the title, and the rest is filler. Search engines have explicit policies against this (scaled content abuse, doorway pagesDoorway pages are pages or sites built to rank for specific search queries that then funnel users to a different destination instead of being useful in their own right. Google treats them as spam., thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count.), and they’re effective at spotting it.
A quick gut check: take any one of your planned pages and mentally delete the keyword you’re targeting. If what’s left is a generic, could-be-anything page, your data isn’t deep enough yet — and that’s exactly what gets these projects in trouble.
Want the full playbook — how to pick a data source, how to keep pages out of trouble, and (the part I care most about) how to actually get thousands of pages indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — switch to the Advanced tab.
TL;DR — Programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. is one modular template plus a structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. source generating pages across a set of similar queries (a core + modifier model). It’s legitimate when each page genuinely answers its query with unique data; it’s spam when execution is thin — and that’s a data problem, not a template problem. The part competitors skip is the technical-at-scale layer: indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. is the first thing that decides whether any of this works, so publish in staged batches and validate indexation and impressions before scaling, and treat crawl budgetThe number of URLs an engine will crawl in a timeframe., internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segmentation, schema, and index bloat as first-class. Automation does not excuse thin or unhelpful output. Evidence for this claim Google defines scaled content abuse as generating many pages primarily to manipulate rankings, regardless of whether automation, humans, or both created them. Scope: Current Google spam policy; scale itself is not the violation. Confidence: high · Verified: Google Search Essentials: Scaled content abuse Evidence for this claim Programmatic pages should provide original value for an intended audience rather than thin permutations created mainly for search traffic. Scope: Current Google helpful-content self-assessment. Confidence: high · Verified: Google Search Central: Creating helpful content
What it actually is
Programmatic SEO is the systematic creation of pages at scale by combining a single modular template with a structured data source, to target a large set of related queries. You build the page model once; the data populates the variations.
The standard mental model is core + modifier. The core is the repeatable page concept (“currency converter,” “X vs Y comparison,” “things to do in”); the modifier is the dimension your data varies along. The modifiers worth knowing:
- Geographic —
[service] in [city],things to do in [place]. - Comparison —
[A] vs [B],[A] alternatives. - Attribute —
[product] for [use case],best [thing] for [audience]. - Format —
[topic] template,[topic] calculator,[topic] examples. - Question —
how to [task],what is [thing].
That [service] in [city] pattern is the most useful one to flag early, because
it’s also the classic doorway-page trap — more on that below.
How to build it
1. The data source is the whole game — rank your options. In order of defensibility:
- Proprietary data you own and nobody else has. This is the moat. At Ahrefs we lean on our own index data across these pages — we’re showcasing our data throughout, not just pushing out automated informational content.
- Public APIs / licensed datasets — usable, but if it’s available to you it’s available to your competitors, so the value has to come from how you present and combine it.
- Scraped feeds — the bottom of the barrel. Republishing someone else’s content without adding value is literally one of Google’s named spam examples.
2. The template must leave room for genuinely unique per-page data. A good template is mostly scaffolding around data that differs meaningfully page to page — not a paragraph of boilerplate with one variable swapped in.
3. CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms., renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., and delivery. Most teams generate these from a database via a CMS or a static-site build. Prefer server-side renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. (SSR) or static-site generation (SSG) so the unique content is in the initial HTML — don’t make Google render client-side JavaScript to see the one thing that makes the page worth indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.. Emit the pages into segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. (see below).
My central thesis: thin content at scale is a data problem, not a template problem
This is the line I keep coming back to. When a programmatic project produces thin pages, people blame the template or the word count and try to “beef up” each page with more text. Wrong fix. If removing the modifier leaves a generic page, your dataset is too shallow. No amount of template polish saves a page that has nothing unique to say. Fix the data — add depth, add dimensions, add things only you know — or don’t publish that page.
Faking it doesn’t work either. To create quality content you need real expertise, and in a lot of cases people are just faking expertise, or have writers faking it. The way you differentiate at scale is by getting real knowledge from the experts and putting in data that’s only available to you.
When it works vs. when it’s spam
It works when: there’s genuine search demand across the modifier set; each page materially answers its query; the data is unique or uniquely presented; and the pages connect to a real business goal, not just a traffic chart.
It’s spam when it’s unoriginal content generated mainly to manipulate rankings — “no matter how it’s created,” as Google’s scaled-content-abuse policy puts it. I’ll be blunt about the automation fantasy: the people that are trying to automate this are not doing well — a lot have fallen. We built something like 300 websites last year, mostly tool sites, specifically to test whether AI systems are good enough to do this. Some stuff works; some works for a while and then falls off. Google isn’t going to reward something you didn’t put real effort into. And the lazy patterns are the obvious targets — when people decided “let me just make an FAQ and put 50 or 100 FAQs on it,” that was never going to work; it’s an obvious thing to be penalized.
The part everyone skips: making it actually rank at scale
Most pSEO guides stop at “publish and monitor.” That’s where the real technical work starts. This is my wheelhouse, so here’s the layer that competitors miss.
Indexing is the first thing that matters
The biggest one is just indexing — is the page indexed or not? It doesn’t matter what else you do if the page isn’t indexed. When you’re publishing thousands of pages at once, indexing is not a given; Google decides what it wants to keep, and thin variants get dropped (or never picked up). So:
- Never publish all of it at once. Roll out in staged batches and validate indexation and impressions before scaling. Publish 10–20, confirm they get indexed and earn impressions, then 50–100, then the full set. If batch one doesn’t index well, batch ten thousand won’t either — and you’ll have learned it cheaply.
- Watch GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.’s Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. for “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” and “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” creeping up. That’s Google telling you the pages aren’t worth its space — usually a data-depth problem, not a tag problem.
Crawl budget and Crawl Stats
For most sites crawl budgetThe number of URLs an engine will crawl in a timeframe. is a non-issue — it starts to matter at large scale, which is exactly where pSEO lives. More crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. doesn’t mean you’ll rank better, but pages that aren’t crawled and indexed won’t rank at all. Use GSC’s Crawl Stats report to watch response codes and average response time, and don’t let parameter explosions and duplicates waste crawl on junk URLs instead of your real pages.
Internal linking — no orphans
Thousands of pages with nothing linking to them are orphans, and orphans don’t get discovered or indexed well. Build a real hub-and-spoke structure: category/hub pages that link to the programmatic pages, and programmatic pages that link laterally to relevant siblings. This is also what Google’s old doorway guidance asks about — whether your pages live as an “island” you can’t navigate to from the rest of the site.
Index bloat and thin variants
Not every cell in your data grid deserves a page. Combinations with no demand or no
real data produce thin pages that dilute the whole project. noindex the thin
variants (or don’t generate them), and prune underperformers over time. This is
closely related to faceted-navigation index bloatAn SEO term for when a search engine has indexed a lot of low-value, thin, or duplicate URLs that don't serve search demand. It's a quality and crawl-efficiency problem, not a penalty. — the same problem of
machine-generated URL combinations multiplying past anything useful.
Sitemaps and schema
- Segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.. At scale, split your URLs across many sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. under a sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file.. Wise’s many-sitemap pattern is the obvious example — segmentation lets you monitor indexation by segment in GSC, so you can see which slice of pages is and isn’t getting indexed.
- Schema where it genuinely fits:
ItemListfor list/aggregation pages,FAQPageonly where there are real FAQs (not the spam pattern above),LocalBusinessfor genuine location entities. Schema doesn’t make a thin page good — it just helps a good page be understood.
Programmatic sets need intent, crawl state, and Search Console evidence compared at the same URL grain. A mismatch is a review cue, not a live-index verdict.
Reconcile a rollout sample with my free Indexation Reconciler Free
- Export the staged sitemap set, current crawl signals, and matching GSC URL state.
- Group mismatches by template, data source, and deployment cohort.
- Verify samples in URL Inspection, repair the generating rule, and expand only after the cohort is clean.
The sample shows one URL aligned across sitemap, current crawl, and GSC indexed export. A second URL is in the sitemap and indexed in the export but currently noindex, prompting review of the blocking directive and export freshness.
Where Bing stands now
Worth knowing: Bing softened its stance in 2026. The old guidelines called machine-generated content “malicious” “garbage” that “will result in penalties.” The updated wording says large-scale content generated without oversight, quality control, or editorial review “may be excluded from indexing.” That’s the same destination Google reached — the standard is editorial oversight plus added value, not whether a machine touched the page.
The accuracy spine — get these right
- Google’s scaled-content-abuse policy targets content made primarily to manipulate rankings that lacks value — “no matter how it’s created.” Automation and AI are not inherently against policy; the line is value + intent + oversight.
- Bing converged on the same conclusion in 2026: value over method.
[service] in [city]templated funnels are a doorway risk, full stop.- Every case-study page count and traffic figure floating around (Wise, Zillow, Zapier, etc.) is a third-party estimate — hedge it.
- Programmatic SEO is legitimate when each page genuinely answers the query with unique data. The execution is spam-or-not; the technique isn’t.
Bottom line
Programmatic SEO is a great way to scale if you have the data and the technical discipline to back it. If you can create good pages programmatically using your data, it can be a great way to scale quickly. If you’re hoping automation will do the thinking for you, you’re building the thing search engines spent the last few years learning to ignore.
AI summary
A condensed take on the Advanced version:
- Programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. = one template + a structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. source generating pages across a core + modifier query set (geographic, comparison, attribute, format, question modifiers).
- The technique is neutral. It’s legitimate when each page answers its query with unique data; it’s spam when execution is thin.
- Central thesis: thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count. at scale is a data problem, not a template problem — if deleting the modifier leaves a generic page, the dataset is too shallow. More template polish won’t save it.
- Data source ranked: proprietary > public API/licensed > scraped (scraping without added value is a named spam example).
- IndexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. comes first. “Is the page indexed or not?” decides everything. Publish in staged batches (10–20 → 50–100 → full), validating indexation and impressions between batches.
- Technical layer competitors skip: crawl budgetThe number of URLs an engine will crawl in a timeframe. + GSC Crawl StatsA Google Search Console report (under Settings) that shows how Google has crawled your site over the last 90 days — total requests, download size, and average response time, broken down by response code, file type, Googlebot type, and purpose. It's only available for root-level properties (a Domain property or a URL-prefix property verified at the site's root).; hub-and-spoke
internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. (no orphans);
noindex/prune thin variants (index bloatAn SEO term for when a search engine has indexed a lot of low-value, thin, or duplicate URLs that don't serve search demand. It's a quality and crawl-efficiency problem, not a penalty., faceted nav); segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. (Wise pattern); schema (ItemList/FAQPage/LocalBusiness) where it genuinely fits. - Policy: Google’s scaled-content-abuse applies “no matter how it’s created”;
[service] in [city]funnels = doorway risk; Bing softened to “may be excluded from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” if oversight/value is missing — same conclusion as Google. - Reality check: the people trying to fully automate this aren’t doing well; winners have proprietary data and real editorial oversight. All case-study numbers are third-party estimates.
Official documentation
The primary-source policies that decide whether a programmatic project is fine or a problem.
- Spam policies — Scaled content abuse — the core policy for pSEO: many low-value pages made to manipulate rankings, “no matter how it’s created.”
- Spam policies — Doorway abuse — why
[service] in [city]funnel pages are risky. - Spam policies — Scraping — republished feeds/data without added value.
- Creating helpful, reliable, people-first content — the “Who, How, Why” self-assessment and search-engine-first red flags.
- Using generative AI content — automation is fine; using it to generate many pages without value is not.
Bing / Microsoft
- Bing Webmaster Guidelines — the 2026 update: large-scale content without oversight/quality control “may be excluded from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..”
Quotes from the source
On-the-record statements that frame the line between legitimate programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. and scaled spam — plus a few of my own positions.
Google — programmatic SEO and scaled content
- “I love fire, but also programmatic SEO is often a fancy banner for spam.” (and, to be fair, his follow-up: “programmatic SEO is not always spam but hey: Forever the optimist.”) — John Mueller, Google. Jump to quote
- “We don’t really care how you’re doing this scaled content, whether it’s AI, automation, or human beings. It’s going to be an issue.” — Danny Sullivan, Google (April 2025). Jump to quote
- “The key things are, large amounts of unoriginal content and also no matter how it’s created.” — Danny Sullivan, Google. Jump to quote
- “As said before when asked about AI, content created primarily for search engine rankings, however it is done, is against our guidance. If content is helpful & created for people first, that’s not an issue.” — Danny Sullivan, @searchliaison (January 2023). Jump to quote
- “We focus on the quality of content, not who produced it. Use AI to provide people with unique, satisfying information.” — Danny Sullivan, Google (brightonSEO 2023). Jump to quote
Google — automation, AI, and oversight
- “Scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. is when many pages are generated for the primary purpose of manipulating search rankings and not helping users… no matter how it’s created.” — Google Search Central, spam policies. Jump to quote
- “I think the word human created is wrong. Basically, it should be human curated. So basically someone had some editorial oversight over their content and validated that it’s actually correct and accurate.” — Gary Illyes, Google (August 2025). Jump to quote
- The Helpful Content language shift: Google’s August 2022 guidance described helpful content as “written by people, for people”; in September 2023 it changed to “created for people” — formally acknowledging that AI-assisted content can be fine when it’s made for users, not ranking manipulation.
Bing / Microsoft
- Old guidelines (pre-2026): machine-generated content “is considered malicious and usually contains garbage text only created to garnish a higher ranking… This type of content will result in penalties.”
- New guidelines (February 2026): “Large-scale content generated without oversight, quality control, or editorial review often lacks usefulness, accuracy, and originality, and may be excluded from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..”
Me, on programmatic SEO
- “The people that are trying to automate this are not doing well… a lot have fallen.” — Patrick Stox (PageTraffic podcast). Read the source
- “The biggest one is just indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — is the page indexed or not? It doesn’t matter what else you do if the page isn’t indexed.” — Patrick Stox (PageTraffic podcast). Read the source
- “If you have the ability to create good pages programmatically using your data, it can be a great way to scale quickly.” — Patrick Stox, Ahrefs. Jump to quote
- “To create quality content, you need real expertise. The problem is that, in many cases, we’re just faking expertise, or we have writers who are faking expertise.” — Patrick Stox, Search Engine Land. Jump to quote
Before-you-scale QA checklist
Run this before you publish a programmatic set — most failures are baked in at this stage, not discovered later.
- Validate data depth. Pick three sample pages, delete the modifier, and confirm what’s left is still genuinely useful. If it’s generic, the dataset is too shallow — fix the data before generating pages.
- Dedupe / cannibalization check. Make sure pages don’t compete for the same query or duplicate near-identical content across the grid.
- Internal-linking plan. Every page is reachable via real
<a href>links from a hub; no orphans; programmatic pages link laterally to relevant siblings. - Indexation pilot batch. Publish 10–20 pages first; confirm they get indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. and earn impressions before generating the rest.
- Schema. Add
ItemList/FAQPage/LocalBusinessonly where it genuinely matches the page; never fake FAQs to trigger markup. - SitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segmentation. Split URLs into segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. under an index so you can monitor indexation by segment in GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results..
- Pruning plan. Decide upfront which thin/no-demand combinations get
noindex’d or never generated, and set a cadence to prune underperformers. - RenderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. check. The unique, page-defining content is in the initial HTML (SSR/SSG), not injected later by client-side JavaScript.
The frameworks
1. Should you do programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. at all?
Answer all four “yes” before you start:
- Business-goal alignment — do these pages serve a real objective, or just a traffic number? (My own example of a good OKR: use our data to create 2,000 programmatic pages in six months to show the value of our data and platform — note it’s anchored to the data and the platform, not raw page count.)
- Conversion connection — is there a plausible path from these pages to something that matters (sign-ups, leads, revenue)?
- Accessible unique data — do you have data that’s yours, or that you can present in a way nobody else does? If the only data is scraped or commodity, stop here.
- Real demand — is there genuine search volume across the modifier set, or are you manufacturing pages for queries nobody runs?
2. The staged-rollout framework
Never ship the whole set at once. Validate indexation and impressions between batches:
- Pilot — 10–20 pages. Publish, then confirm in GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. that they indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. and start earning impressions. If they don’t index, stop and fix the data/template — the problem will only multiply.
- Expand — 50–100 pages. Re-check indexation rate and early impression trends across the larger set. Watch “Crawled/Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal..”
- Full rollout. Only once batches one and two index cleanly. Keep monitoring by sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segment, and prune the combinations that never index or never earn impressions.
The principle: each batch is a cheap experiment that tells you whether the next, bigger batch is worth generating.
Spam policy → pSEO mistake — cheat sheet
Each search-engine concept maps to a specific programmatic failure mode. If your project does the thing in the right column, the policy in the left column is the one that bites you.
| Search-engine concept | The pSEO mistake that triggers it |
|---|---|
| Scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. (Google) | Many unoriginal, templated pages where only the modifier changes; “write me 100 pages on 100 topics” output with nothing original. |
| Doorway abuse (Google) | [service] in [city] pages that funnel users to one destination; substantially similar pages closer to search results than a real browseable hierarchy. |
| Scraping (Google) | Republished data feeds or other sites’ content with no added value or unique benefit. |
| Thin / unhelpful content (Google Helpful Content; Bing) | Modifier-only pages with no genuine per-page data; pages that leave readers needing to search again. |
| No oversight / no editorial review (Bing, 2026) | Large-scale generated pages published with no quality control — “may be excluded from indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” |
| Index bloatAn SEO term for when a search engine has indexed a lot of low-value, thin, or duplicate URLs that don't serve search demand. It's a quality and crawl-efficiency problem, not a penalty. (technical) | Generating every grid combination regardless of demand or data; faceted-nav-style URL explosion. |
The one-line test: delete the modifier from a page. If what’s left is generic, the page is thin — fix the data, not the template.
Prompts for pressure-testing a programmatic set
Run the delete-the-modifier test across a dataset
Paste a sample of your rows plus the page template or rendered page fields. Include the primary modifier column. Expect a row-level risk review, not generated filler.
Audit this proposed programmatic SEO dataset and template for distinct per-page value.
For each sample row:
1. Identify the primary modifier.
2. Describe what useful information remains if that modifier and its direct mentions
are removed from the rendered page.
3. Mark the row as distinct, borderline, or generic.
4. Name the supplied fields that create real page-specific value.
5. If it is borderline or generic, say whether the honest fix is deeper data,
consolidation into a broader page, noindex, or not generating the URL.
Then flag rows likely to cannibalize one another, pages that funnel to the same final
destination without standalone value, and fields that merely restate commodity or
scraped data. Do not write extra paragraphs to disguise shallow data. Do not invent
demand, proprietary fields, or conversion value that I did not provide.
[PASTE DATA SAMPLE AND TEMPLATE/RENDERED FIELDS]Build a staged rollout review
Paste the proposed URL pattern, data source, internal-link plan, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segmentation, and results from the current batch. Expect a go, stop-and-fix, or do-not-generate call.
Review this programmatic SEO rollout using these gates:
- The pages support a business goal and a plausible conversion path.
- The dataset supplies useful, distinct information for each modifier.
- Real search demand exists across the intended combinations.
- Every page is reachable through crawlable internal links and a segmented sitemap.
- The unique content is present in the rendered HTML.
- The pilot is 10–20 pages; expansion is 50–100 pages; full rollout waits until the
earlier batches index and begin earning impressions.
Return:
1. A verdict: proceed to the next batch, stop and fix, or do not generate.
2. Evidence for each gate using only the supplied material.
3. Any scaled-content, doorway, scraping, cannibalization, or index-bloat risk.
4. The smallest next batch and the GSC/sitemap evidence required before scaling again.
Do not infer that indexed pages are valuable merely because they indexed, and do not
invent an acceptable indexing-rate benchmark.
[PASTE PROJECT PLAN AND CURRENT BATCH RESULTS] Patrick's relevant free tools
- Canonicalization Checker — Audit HTML and HTTP canonical signals, test the canonical target, and identify observable conflicts that can cause Google to choose a different URL.
Tools for preflight and staged rollout checks
Start with the on-site tools
- Google Index Checker — checks observable status,
redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't.,
noindex, and canonical blockers on pilot URLs, then points you to GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. URL Inspection for Google’s actual answer. Use a representative sample from each batch. - XML Sitemap Validator — validates the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segments used to monitor each template or rollout cohort, with errors and warnings tied to the XML rather than a guessed indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. outcome.
- XML Sitemap Generator — creates a capped, robots-respecting sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. from a same-site crawl and keeps noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., off-canonical, failed, and uncertain URLs separate. Use it to compare crawlable output with the URL set your generator intended to publish.
Complete the evidence chain
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Page IndexingThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. and Performance reportsThe Google Search Console report that shows how your site actually performed in Google Search, built from real impressions and clicks. It reports four metrics — clicks, impressions, average CTR, and average position — and keeps the most recent 16 months of data. — filter by sitemap or URL pattern to see whether each staged batch is indexed and begins earning impressions.
- Google Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. — confirms Google’s reported state for a representative URL when a batch-level report needs a concrete example.
- A full-site crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — checks status, canonicals, directives, render-visible unique content, internal-link depth, orphan candidates, and duplicate page patterns before scale multiplies them.
- Server access logs — show whether GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. reaches the programmatic section and whether crawl capacity is being consumed by unwanted parameter or facet combinations.
- Schema validation — use it only for markup that truthfully matches the page; valid syntax cannot make a shallow dataset useful.
Resources worth your time
My related writing
- The Beginner’s Guide to Technical SEO — where crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and architecture fit, all of which programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. leans on.
- Enterprise SEO — scaling SEO with data and process, including programmatic pages done from your own data.
- What is quality content? — why real expertise and unique data are the differentiators, and why faked expertise fails at scale.
- When Should You Worry About Crawl Budget? — most sites don’t, but programmatic projects are exactly the case where you might.
Examples I’ve used
- The “SEO for x” page pattern — re-using components to create pages where x is a different type of business — is a small-effort, real-data example of programmatic pages working.
- The 2,000 programmatic pages in six months OKR (from my SEO OKRs piece) — a goal framed around proving the value of your data and platform, not just publishing volume.
I’ve written more about the technical side — crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and crawl budgetThe number of URLs an engine will crawl in a timeframe. — across these; the programmatic-specific lessons (indexing first, staged rollouts, data over templates) come from watching scaled projects succeed and fail in practice.
From around the industry
- Programmatic SEO, Explained for Beginners (Ryan Law, Ahrefs) — solid primer on the definition, real-world examples (Wise, Zapier, Webflow), and the “or spam?” question; predates the May 2024 scaled-content-abuse update.
- Programmatic SEO: Scale content, rankings & traffic fast (Search Engine Land) — covers template design, avoiding index bloatAn SEO term for when a search engine has indexed a lot of low-value, thin, or duplicate URLs that don't serve search demand. It's a quality and crawl-efficiency problem, not a penalty., and tracking performance at scale.
- Programmatic SEO: What It Is + Tips & Examples for 2026 (Backlinko) — good rundown of when pSEO makes sense vs. when it doesn’t, with traffic estimates for Wise, Zapier, TripAdvisor, and Zillow.
- Google On Scaled Content: ‘It’s Going To Be An Issue’ (Search Engine Journal) — Danny Sullivan’s clearest on-record statement that method doesn’t matter; intent and value do.
- Bing Adds GEO To Official Guidelines, Expands AI Abuse Definitions (Search Engine Journal, Feb 2026) — side-by-side of old vs. new Bing policy wording on large-scale generated content.
- Google’s Spam Policies — Scaled Content Abuse — the canonical policy text; “no matter how it’s created.”
- Google Crawl Budget documentation — official guidance on when crawl budgetThe number of URLs an engine will crawl in a timeframe. matters and how to manage it; directly relevant for any large-scale pSEO project.
Test yourself: Programmatic SEO
Five questions on data depth, policy risk, and staged rollout. Pick an answer for each, then check.
Programmatic SEO
Programmatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset.
Related: Scaled Content Abuse, Doorway Pages, Faceted Navigation, Crawl Budget, Thin Content
Programmatic SEO
Programmatic SEO — often shortened to pSEO — is the systematic creation of pages at scale by combining one modular template with a structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. source. Instead of writing each page by hand, you write the template once and let the data populate thousands (sometimes millions) of variations targeting related long-tail queries, like “[currency A] to [currency B]” or “[tool A] vs [tool B].”
The technique itself is neutral. Wise’s currency pages, Zapier’s app-integration pages, and Zillow’s listings are all programmatic, and they work because every page carries real, unique data a searcher actually wants. The same approach produces spam when the only thing that changes between pages is the modifier in the title — that’s where it collides with Google’s scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. and doorway policies, and with Bing’s stance against large-scale content published without oversight.
The honest test is a data test, not a template test: if you strip the modifier out of a page and what’s left is a generic, interchangeable shell, the dataset behind it is too shallow to support pages at scale. The other thing that decides whether programmatic SEO works is whether the pages actually get indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. — which is why crawl budgetThe number of URLs an engine will crawl in a timeframe., internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segmentation, and index-bloat control matter as much as the content model.
Related: Scaled Content Abuse, Doorway Pages, Faceted Navigation, Crawl Budget, Thin Content
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Revision history
Compare the published article with an archived editorial snapshot. Added and removed words are shown only after you open a comparison.
Updated Jul 22, 2026.
Editorial summary and recorded change details.Summary
Moved evidence notes from detached positions to the exact policy claims they support.
Change details
-
Attached the scaled-content-abuse and helpfulness citations directly to the relevant Beginner and Advanced claims.