Programmatic SEO

What programmatic SEO actually is, when it works and when it's spam, and how to build pages at scale that get indexed — from Patrick Stox.

First published: Jun 24, 2026 · Last updated: Jul 22, 2026 · Advanced
demand #1 in Programmatic SEO#67 on the site
1 evidence signal on this page

Programmatic SEO (pSEO) is one template plus a data source generating many pages for similar queries. It's legitimate when each page genuinely answers its query with unique data, and it's spam when it stamps a thin template across a shallow dataset — which is what trips Google's scaled-content-abuse and doorway policies (and now Bing's). My contrarian take: thin content at scale is a data problem, not a template problem. And the first thing that actually decides whether any of it works is indexing — publish in staged batches, validate indexation and impressions before you scale, and treat crawl budget, internal linking, sitemaps, and index bloat as first-class. The people trying to fully automate this aren't doing well; the ones winning have proprietary data and real oversight.

TL;DR — Programmatic SEOProgrammatic SEO (pSEO) is the practice of generating many pages from a single template plus a data source to target large sets of similar queries. It's powerful when each page genuinely answers its query with unique data, and spam when it just stamps a thin template across a shallow dataset. is one modular template plus a structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. source generating pages across a set of similar queries (a core + modifier model). It’s legitimate when each page genuinely answers its query with unique data; it’s spam when execution is thin — and that’s a data problem, not a template problem. The part competitors skip is the technical-at-scale layer: indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. is the first thing that decides whether any of this works, so publish in staged batches and validate indexation and impressions before scaling, and treat crawl budgetThe number of URLs an engine will crawl in a timeframe., internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them., sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. segmentation, schema, and index bloat as first-class. Automation does not excuse thin or unhelpful output. Evidence for this claim Google defines scaled content abuse as generating many pages primarily to manipulate rankings, regardless of whether automation, humans, or both created them. Scope: Current Google spam policy; scale itself is not the violation. Confidence: high · Verified: Google Search Essentials: Scaled content abuse Evidence for this claim Programmatic pages should provide original value for an intended audience rather than thin permutations created mainly for search traffic. Scope: Current Google helpful-content self-assessment. Confidence: high · Verified: Google Search Central: Creating helpful content

What it actually is

Programmatic SEO is the systematic creation of pages at scale by combining a single modular template with a structured data source, to target a large set of related queries. You build the page model once; the data populates the variations.

The standard mental model is core + modifier. The core is the repeatable page concept (“currency converter,” “X vs Y comparison,” “things to do in”); the modifier is the dimension your data varies along. The modifiers worth knowing:

  • Geographic[service] in [city], things to do in [place].
  • Comparison[A] vs [B], [A] alternatives.
  • Attribute[product] for [use case], best [thing] for [audience].
  • Format[topic] template, [topic] calculator, [topic] examples.
  • Questionhow to [task], what is [thing].

That [service] in [city] pattern is the most useful one to flag early, because it’s also the classic doorway-page trap — more on that below.

How to build it

1. The data source is the whole game — rank your options. In order of defensibility:

  • Proprietary data you own and nobody else has. This is the moat. At Ahrefs we lean on our own index data across these pages — we’re showcasing our data throughout, not just pushing out automated informational content.
  • Public APIs / licensed datasets — usable, but if it’s available to you it’s available to your competitors, so the value has to come from how you present and combine it.
  • Scraped feeds — the bottom of the barrel. Republishing someone else’s content without adding value is literally one of Google’s named spam examples.

2. The template must leave room for genuinely unique per-page data. A good template is mostly scaffolding around data that differs meaningfully page to page — not a paragraph of boilerplate with one variable swapped in.

3. CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms., renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., and delivery. Most teams generate these from a database via a CMS or a static-site build. Prefer server-side renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. (SSR) or static-site generation (SSG) so the unique content is in the initial HTML — don’t make Google render client-side JavaScript to see the one thing that makes the page worth indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.. Emit the pages into segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. (see below).

My central thesis: thin content at scale is a data problem, not a template problem

This is the line I keep coming back to. When a programmatic project produces thin pages, people blame the template or the word count and try to “beef up” each page with more text. Wrong fix. If removing the modifier leaves a generic page, your dataset is too shallow. No amount of template polish saves a page that has nothing unique to say. Fix the data — add depth, add dimensions, add things only you know — or don’t publish that page.

Faking it doesn’t work either. To create quality content you need real expertise, and in a lot of cases people are just faking expertise, or have writers faking it. The way you differentiate at scale is by getting real knowledge from the experts and putting in data that’s only available to you.

When it works vs. when it’s spam

It works when: there’s genuine search demand across the modifier set; each page materially answers its query; the data is unique or uniquely presented; and the pages connect to a real business goal, not just a traffic chart.

It’s spam when it’s unoriginal content generated mainly to manipulate rankings — “no matter how it’s created,” as Google’s scaled-content-abuse policy puts it. I’ll be blunt about the automation fantasy: the people that are trying to automate this are not doing well — a lot have fallen. We built something like 300 websites last year, mostly tool sites, specifically to test whether AI systems are good enough to do this. Some stuff works; some works for a while and then falls off. Google isn’t going to reward something you didn’t put real effort into. And the lazy patterns are the obvious targets — when people decided “let me just make an FAQ and put 50 or 100 FAQs on it,” that was never going to work; it’s an obvious thing to be penalized.

The part everyone skips: making it actually rank at scale

Most pSEO guides stop at “publish and monitor.” That’s where the real technical work starts. This is my wheelhouse, so here’s the layer that competitors miss.

Indexing is the first thing that matters

The biggest one is just indexing — is the page indexed or not? It doesn’t matter what else you do if the page isn’t indexed. When you’re publishing thousands of pages at once, indexing is not a given; Google decides what it wants to keep, and thin variants get dropped (or never picked up). So:

  • Never publish all of it at once. Roll out in staged batches and validate indexation and impressions before scaling. Publish 10–20, confirm they get indexed and earn impressions, then 50–100, then the full set. If batch one doesn’t index well, batch ten thousand won’t either — and you’ll have learned it cheaply.
  • Watch GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results.’s Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. for “Crawled – currently not indexedA Google Search Console Page Indexing status meaning Googlebot fetched the page but Google decided not to index it — usually a content- or site-quality signal, not a technical error.” and “Discovered – currently not indexedA Google Search Console Page Indexing status meaning Google knows the URL exists but hasn't crawled it yet — the Last Crawl date is empty. Often a crawl-capacity or crawl-demand (site-quality) signal.” creeping up. That’s Google telling you the pages aren’t worth its space — usually a data-depth problem, not a tag problem.

Crawl budget and Crawl Stats

For most sites crawl budgetThe number of URLs an engine will crawl in a timeframe. is a non-issue — it starts to matter at large scale, which is exactly where pSEO lives. More crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. doesn’t mean you’ll rank better, but pages that aren’t crawled and indexed won’t rank at all. Use GSC’s Crawl Stats report to watch response codes and average response time, and don’t let parameter explosions and duplicates waste crawl on junk URLs instead of your real pages.

Internal linking — no orphans

Thousands of pages with nothing linking to them are orphans, and orphans don’t get discovered or indexed well. Build a real hub-and-spoke structure: category/hub pages that link to the programmatic pages, and programmatic pages that link laterally to relevant siblings. This is also what Google’s old doorway guidance asks about — whether your pages live as an “island” you can’t navigate to from the rest of the site.

Index bloat and thin variants

Not every cell in your data grid deserves a page. Combinations with no demand or no real data produce thin pages that dilute the whole project. noindex the thin variants (or don’t generate them), and prune underperformers over time. This is closely related to faceted-navigation index bloatAn SEO term for when a search engine has indexed a lot of low-value, thin, or duplicate URLs that don't serve search demand. It's a quality and crawl-efficiency problem, not a penalty. — the same problem of machine-generated URL combinations multiplying past anything useful.

Sitemaps and schema

  • Segmented XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.. At scale, split your URLs across many sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. under a sitemap indexA sitemap index is a sitemap of sitemaps — a single file that lists your other sitemap files instead of listing URLs directly. It's how large sites stay under the 50,000-URL / 50MB-per-sitemap limit while submitting just one file.. Wise’s many-sitemap pattern is the obvious example — segmentation lets you monitor indexation by segment in GSC, so you can see which slice of pages is and isn’t getting indexed.
  • Schema where it genuinely fits: ItemList for list/aggregation pages, FAQPage only where there are real FAQs (not the spam pattern above), LocalBusiness for genuine location entities. Schema doesn’t make a thin page good — it just helps a good page be understood.
TIP Catch template-wide indexability drift during a staged rollout

Programmatic sets need intent, crawl state, and Search Console evidence compared at the same URL grain. A mismatch is a review cue, not a live-index verdict.

Reconcile a rollout sample with my free Indexation Reconciler Free

  1. Export the staged sitemap set, current crawl signals, and matching GSC URL state.
  2. Group mismatches by template, data source, and deployment cohort.
  3. Verify samples in URL Inspection, repair the generating rule, and expand only after the cohort is clean.
One conflicting row can identify a generating-rule regression before the next rollout cohort multiplies it.

The sample shows one URL aligned across sitemap, current crawl, and GSC indexed export. A second URL is in the sitemap and indexed in the export but currently noindex, prompting review of the blocking directive and export freshness.

Where Bing stands now

Worth knowing: Bing softened its stance in 2026. The old guidelines called machine-generated content “malicious” “garbage” that “will result in penalties.” The updated wording says large-scale content generated without oversight, quality control, or editorial review “may be excluded from indexing.” That’s the same destination Google reached — the standard is editorial oversight plus added value, not whether a machine touched the page.

The accuracy spine — get these right

  1. Google’s scaled-content-abuse policy targets content made primarily to manipulate rankings that lacks value — “no matter how it’s created.” Automation and AI are not inherently against policy; the line is value + intent + oversight.
  2. Bing converged on the same conclusion in 2026: value over method.
  3. [service] in [city] templated funnels are a doorway risk, full stop.
  4. Every case-study page count and traffic figure floating around (Wise, Zillow, Zapier, etc.) is a third-party estimate — hedge it.
  5. Programmatic SEO is legitimate when each page genuinely answers the query with unique data. The execution is spam-or-not; the technique isn’t.

Bottom line

Programmatic SEO is a great way to scale if you have the data and the technical discipline to back it. If you can create good pages programmatically using your data, it can be a great way to scale quickly. If you’re hoping automation will do the thinking for you, you’re building the thing search engines spent the last few years learning to ignore.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.