Site Architecture
Silo structure, hub and spoke, and topic clusters compared — what's actually different, what Google recommends, and why internal links beat URL folders.
Site architecture is the set of crawlable discovery and navigation paths between your pages, plus supporting entry points like sitemaps — it's not the same thing as your URL folders. Silo, hub and spoke, and topic clusters are overlapping practitioner labels for that structure, not documented Google categories — hub and spoke and topic clusters share the same underlying pattern more than they genuinely differ. Google's documented minimum is a crawlable link to every page you care about; that supports discovery, it doesn't guarantee crawling, indexing, ranking, traffic, sitelinks, or AI citations. No Google source requires banning links between topic areas — evaluate strict-silo rules against user navigation and relevance, not assumed authority mechanics. What matters is contextual internal linking, not URL folders.
Evidence for this claim Google uses links to discover pages and as a relevance signal, so crawlable internal navigation supports discovery and understanding. Scope: Current Google internal-link guidance. Confidence: high · Verified: Google Search Central: Link best practices Evidence for this claim Google recommends navigation paths from menus to categories and subcategories to products, with direct links to important pages. Scope: Current Google ecommerce site-structure guidance, broadly applicable to hierarchical sites. Confidence: high · Verified: Google Search Central: Ecommerce site structureTL;DR — Site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. is the crawlable network of links between your pages — not just your URL folders. The goal is simple: every important page should have a link pointing to it, and related pages should link to each other. You’ll hear three names for organizing that network — silos, hub and spoke, and topic clusters — they’re overlapping practitioner labels for roughly the same idea, not three separate systems. None of this guarantees rankings or traffic by itself — it’s what makes pages findable and understandable. Google likes a pyramid: home page at the top, broad category pages in the middle, specific articles at the bottom.
What site architecture is
Site architecture is the network of crawlable links that connects your pages —
the paths search engines (and readers) use to discover and move between them —
plus supporting entry points like an XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags.. It’s not the same thing as your
URL folders: a page’s path (/blog/category/post/) doesn’t by itself establish
where that page sits in the architecture. It’s the difference between a tidy
library where every book has a shelf and a sign pointing to it, and a pile of
books on the floor.
It matters for search for two reasons:
- Discovery. Search engines find pages mainly by following standard
<a href>links from pages they already know about. If nothing links to a page, it’s hard for Google to find it at all. - Context. When a page links to another with relevant words around the link, it helps Google understand what the linked page is about.
Discovery isn’t the finish line, though — a page being linked and crawlable doesn’t guarantee it gets indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. or ranks. Architecture gets you foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. and understood; it doesn’t promise an outcome past that.
The three models you’ll hear about
- Silo structureA silo structure is a site organized by grouping related content into topic areas and concentrating internal links within each topic so a hub page and its supporting pages reinforce each other. Modern practice keeps the topical concentration and drops the old rule against cross-linking between topics. — you split your content into separate topic “silos” and keep each silo’s pages linking mostly to each other. The old, strict version says never link between silos.
- Hub and spoke — one broad overview page (the hub) links out to a set of detailed pages (the spokes), and each spoke links back to the hub.
- Topic clusters — a “pillar” page covers a big topic, and “cluster” pages cover the smaller subtopics, all interlinked.
Here’s the part most guides bury: hub and spoke and topic clusters describe the same underlying structure. One name came from library/information-architecture people; the other was popularized by HubSpot in 2017 as a content-marketing idea. None of these three are Google-documented categories — they’re practitioner labels with overlapping implementations, not an official taxonomy. I’ll explain that more in the Advanced version.
What actually matters
You don’t need to pick a brand name. You need:
- A clear hierarchy: home page → category/hub pages → specific pages.
- Every important page linked from at least one other page (no “orphans”) — Google’s own documentation puts it plainly: every page you care about should have a link from at least one other page on your site.
- Real
<a href>links, not click-only navigation Google can’t reliably follow. - Related pages linking to each other with descriptive link textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page..
- Nothing important buried so deep that crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. struggle to reach it.
None of this is a guarantee of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., or ranking — it’s what makes a page reachable and legible in the first place.
The thing most people get wrong
You don’t have to avoid linking between sections. No Google source requires banning links between topic areas. The strict silo rule — “never link between silos” — trades away links that would genuinely help readers navigate and help Google understand the destination page, for the sake of a rule that isn’t documented anywhere. Judge cross-links by whether they’re relevant and useful to a reader, not by whether they cross a silo boundary. The valuable idea from silos is grouping related content together, not walling it off.
Want the deeper version — how link equityPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. flows, flat vs. deep hierarchies, and when each model fits — switch to the Advanced tab.
Evidence for this claim Google uses links to discover pages and as a relevance signal, so crawlable internal navigation supports discovery and understanding. Scope: Current Google internal-link guidance. Confidence: high · Verified: Google Search Central: Link best practices Evidence for this claim Google recommends navigation paths from menus to categories and subcategories to products, with direct links to important pages. Scope: Current Google ecommerce site-structure guidance, broadly applicable to hierarchical sites. Confidence: high · Verified: Google Search Central: Ecommerce site structureTL;DR — Silo, hub and spoke, and topic clusters are overlapping practitioner labels, not documented Google architecture categories — but they converge on one structure Mueller has described favorably: a pyramid / top-down hierarchy. Hub and spoke and topic clusters share the same underlying pattern more than they genuinely differ — one comes from information architecture, the other from HubSpot’s 2017 content-marketing rebrand. The part of silo thinking worth keeping is topical concentration of internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.; the part to drop is the strict “no cross-silo links” rule — no Google source requires it, so judge cross-links by relevance and user navigation instead. What Google leans on is internal linkingLinks between pages on the same site. and context, not URL folder structure. Keep hierarchies reasonably shallow, link related pages across clusters when it’s relevant, and remember architecture enables outcomes like topical authoritySemantic search is meaning-based retrieval — matching what a user means, not just the words they typed. Search engines detect entities, expand synonyms, infer intent, and rank by conceptual relevance, which is why keyword stuffing lost its power and topical depth gained it. and AI-search visibility — it doesn’t guarantee them, and it can’t rescue thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count..
Why SEOs argue about this at all
Site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. is one of those topics where three communities invented overlapping vocabularies and then spent fifteen years insisting their word was the real one. I’ve spent six-plus years at Ahrefs on the product side around Site Audit, looking at what site structuresWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders. actually look like in crawl data across a huge number of sites — and the gap between the dogma and what works is wide. So let me try to collapse the confusion.
There are three named models:
- Silo structureA silo structure is a site organized by grouping related content into topic areas and concentrating internal links within each topic so a hub page and its supporting pages reinforce each other. Modern practice keeps the topical concentration and drops the old rule against cross-linking between topics. — content grouped into isolated thematic sections. The strict interpretation (popularized by Bruce Clay) prohibits cross-silo internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to keep each silo’s “link equityPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems.” concentrated.
- Hub and spoke — a broad overview hub page links out to detailed spoke pages, which link back. The hub is sometimes called a pillar page.
- Topic clusters — HubSpot’s 2017 framing: a pillar page targets a broad keyword, cluster pages target long-tail subtopics, everything interlinked.
The clarification I’ll plant my flag on
None of “silo,” “hub and spoke,” or “topic cluster” is a Google-documented architecture category. They’re practitioner labels, and in practice their implementations overlap heavily. “Hub and spoke” is the older information-architecture term: a central overview page linking to detailed subtopic pages with reciprocal links. “Topic clusters” is HubSpot’s 2017 content-marketing rebrand of essentially that same pattern — a pillar page plus interlinked cluster content. “Pillar page” is just the marketing word for the hub. Calling them flatly identical overstates it — HubSpot’s framing leans more on planned topical coverage as a content strategy, while “hub and spoke” is older, more general information-architecture vocabulary — but if you’re choosing how to structure links, you’re solving the same problem either way: a central page, detailed subtopic pages, and reciprocal links between them.
And once you allow cross-silo links — which nothing Google has published tells you not to — a “silo” is a hub-and-spoke cluster in practice. So the three names mostly converge on one workable structure: grouped, hierarchical, internally linked, with sensible cross-links. Treat them as overlapping vocabulary for that structure, not as three competing systems to choose between.
What Google actually recommends
Google’s recommendation is a pyramid / top-down hierarchy: the home page covers the broadest topic, category/hub pages sit in the middle, and specific content pages live at the bottom. John Mueller put the why plainly: the top-down approach or pyramid structure “helps us a lot more to understand the context of individual pages within the site.” That’s the payoff — structure isn’t a ranking trick, it’s how Google figures out what each page is about and how pages relate.
Worth noting: Google uses the term “hub page” in its own documentation — describing how “a hub page, such as a category page, links to a new blog post” for discovery. So the hub-and-spoke framing isn’t an SEO invention; it’s Google’s own language.
Internal links beat URL structure
This is the single most misunderstood part of architecture. A lot of silo advice is
really about URL folders — putting /category-a/ pages under one path and
forbidding links to /category-b/. But Google focuses on internal linking
signals, not URL path segments. Mueller has repeatedly noted that some SEOs
over-focus on URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly.; Google works out hierarchy from how pages link to each
other, not from folder names.
The implication is big: a page at /blog/technical-seo/site-architecture/ signals
nothing to Google beyond what its content and inbound links say about it. Logical
URLs are good for humans and for managing the site — but they aren’t the
architecture. The link graph is the architecture. This is also why “virtual
silos” (concentrating links by topic regardless of folder) work fine, and why
strict “physical silos” (folder-based isolation with no cross-links) buy you crawl
and UX problems for no offsetting benefit.
Google’s documentation is specific about what counts as a link it can reliably
follow: a standard <a href> element with a resolvable URL. Navigation built
only from JavaScript click handlers or non-standard markup doesn’t get you the
same reliable discovery path, no matter how clean your URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. looks. Get
the crawlable-link contract right first — folder naming is secondary.
Where strict silos go wrong
The legitimate insight in silo thinking is topical concentration — grouping related content and linking generously within a topic area genuinely helps readers navigate and helps Google understand what a page is about. The failure mode is the strict rule: never link between silos. No reviewed Google source requires that prohibition — so judge it on its own terms, not as an assumed authority mechanic:
- it prevents natural, contextually relevant links readers would benefit from,
- it hurts UX by dead-ending users at artificial silo boundaries, and
- a page can still be relevant to more than one topic area, and a link that helps a reader find it shouldn’t be blocked because it crosses a label you drew.
Shari Thurow has been making this point for years (“stop the silo madness”), and Ahrefs’ own contrarian take (Joshua Hardwick’s “why it makes no sense”) lands in the same place. Keep the concentration; drop the wall — and make the call based on whether a link genuinely serves the reader and the destination page, not on unverified claims about how strictly Google weighs isolation.
Flat vs. deep hierarchies
Two failure modes at the extremes:
- Too flat — everything one click off the home page. Link equity and topical signal get diluted; the home page can’t meaningfully vouch for hundreds of equal children, and you lose the topical grouping that helps context.
- Too deep — important pages many clicks from home. Mueller’s framing: going too deep “makes it harder for us to crawl and harder for us to pass the signals around.” Deep pages tend to get crawled less and inherit less internal authority. Google hasn’t published a specific click-depth number; don’t treat any fixed count as a requirement — the right depth depends on your site’s size and how distinct its categories are, which is its own topic (see the flat vs. deep and crawl-depth deep dives in this cluster for the fuller decision framework).
The target is a shallow hierarchy with strong contextual linking: keep important pages reachable in as few hops as your site’s scale reasonably allows, grouped into hubs, with cross-links where topics genuinely relate. This is the same crawl-depth concern that shows up whenever you’re auditing how far botsA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. have to travel to reach your money pages.
What internal linking actually does
Architecture’s real job is shaping the internal link graph. Three things to get right:
- Linkability. Every important page should have a link from at least one other page — ideally several. Google’s own guidance is explicit about this minimum. Pages with no inlinks are harder to find and harder to rank, but a link doesn’t guarantee either outcome.
- Context. The words before and after a link, and the anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. itself, can help people and Google understand what the target is about. This is how a hub “explains” its spokes — it’s not a promise of a specific ranking effect.
- Concentration. Linking densely within a topic cluster puts relevance signals where they belong — the kernel of truth silos were always reaching for. Whether that concentration translates into something practitioners call “topical authority” isn’t something Google documents directly; treat it as a reasonable hypothesis worth testing on your own site, not a guaranteed mechanism.
This site is a live example. patrickstox.com runs a pillar → cluster → article →
sub-article hierarchy across a few hundred pages, using pillar, cluster,
clusterSelf, subcluster, and subsubcluster taxonomy — and it deliberately
cross-lists articles across clusters (the alsoIn pattern) rather than walling
them off. That’s hub-and-spoke with intentional cross-linking: exactly the model I’m
describing, not a strict silo.
When to use which model
Honestly? Build the hub-and-spoke / topic-cluster structure and stop worrying about the labels:
- Pick a pillar/hub topic with real breadth and search demand.
- Identify the subtopics that have their own demand — those become spokes.
- Interlink hub ↔ spokes and spoke ↔ spoke where relevant.
- Cross-link to other clusters when the context is genuinely related.
There’s no magic number of cluster pages per hub. The right count is however many distinct subtopics with real demand exist — not an arbitrary target of 5, 10, or 30. Wikipedia is the canonical example: broad overview pages linking deeply to detailed subtopic pages, all densely interlinked, no silo walls anywhere.
Architecture and AI search
AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. and AI assistants use query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics. — they decompose a question into multiple related sub-queries. The reasonable hypothesis, not a documented guarantee: a site with organized, interlinked coverage across a topic’s subtopics has a better chance of surfacing across more of those sub-queries, since more of the subtopics are covered by a findable, crawlable page. I haven’t seen controlled evidence isolating architecture as the cause here, so treat it as something to test on your own content rather than an established mechanism. Gary Illyes’ public position is that AI search optimizationAI search optimization is the practice of making your brand and content visible, citable, and accurately represented across AI-powered search — Google AI Overviews, ChatGPT, Perplexity, Copilot. It's built on traditional SEO plus a heavier emphasis on off-site brand mentions and content AI systems can cite. needs normal SEO — well-structured, crawlable, high-quality content — not a special architecture. The structure that already serves traditional search is the same one you’d build for AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.; there’s no separate playbook.
Auditing your architecture
A practical pass, mostly with crawl data:
- Find orphans — pages with no internal links in. Crawl the site (Ahrefs Site Audit, Screaming Frog) and look for zero-inlink pages.
- Find over-deep pages — important URLs buried many clicks from home; no fixed number is the rule, but if crawl data shows them getting crawled less, dig in.
- Map topics and find hub candidates — clusters of related pages missing a central overview.
- Run an internal-link gap analysis — related pages that should link to each other but don’t.
Candidate scores surface related pages and explain why they were paired. They do not establish the right architecture or replace a crawl-based orphan and depth audit.
Prioritize supplied-page connections with my free Internal Link Cluster Visualizer Free
- Supply a representative hub, its spokes, and their current internal links.
- Review high-scoring pairs whose targets have few supplied inlinks.
- Approve only contextually useful links, then recrawl to verify the resulting graph and click depth.
The row scores a proposed link at 76.6 from Internal linking explained to Topic cluster strategy. Reasons include shared architecture and internal-link terms, a common blog URL segment, and low or absent supplied inlinks to the target.
The honest caveat
Architecture is infrastructure, not a ranking shortcut. It enables crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., controls how link equity flows, and helps engines understand context — but a beautiful hub-and-spoke structure wrapped around thin, low-value content still won’t rank. Illyes has noted it’s rare to see two results from one domain in a SERP; structure serves content quality and relevance, it doesn’t substitute for them. Get the structure right so that good content can do its job — that’s the whole point.
The neighboring topics in this cluster — how internal links pass signals, how crawl depth affects discovery, and faceted navigationFaceted navigation (faceted search, product filtering) lets visitors refine a list of products or content by attribute — price, color, size, brand, rating. The SEO problem: each filter combination can spawn a distinct crawlable URL, turning a small catalog into millions of near-duplicate pages that waste crawl budget and dilute ranking signals. on large sites — all plug into the same architecture decisions covered here.
AI summary
A condensed take on the Advanced version:
- Site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. = the crawlable link graph, not the URL folders. It shapes discovery and how clearly Google understands each page’s context — it doesn’t guarantee crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., ranking, or AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. by itself.
- Three practitioner labels, one overlapping structure. Silo, hub and spoke, and topic clusters aren’t Google-documented categories; they converge on a pyramid / top-down hierarchy, which Mueller has described favorably: the pyramid structure “helps us a lot more to understand the context of individual pages.”
- Hub and spoke and topic clusters mostly overlap. Same underlying pattern, different communities — IA term vs. HubSpot’s 2017 content-marketing rebrand. “Pillar page” is just the marketing word for the hub.
- Drop strict silos. The valuable bit is topical concentration of links; no Google source requires the “never link between silos” rule, and it hurts UX.
- Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. > URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly.. Google works out hierarchy from links, not path segments; Mueller has said SEOs over-focus on URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly..
- Keep it shallow, but don’t chase a fixed number. Google hasn’t published a click-depth threshold; too deep hurts crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and signal flow, too flat dilutes topical signal.
- No magic cluster count. Cover the subtopics with real demand; Wikipedia is the model.
- AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. needs normal SEO (Illyes) — organized, interlinked coverage is a reasonable hypothesis for helping with query fan-outQuery fan-out is the technique where an AI search system breaks a single user question into multiple related sub-queries, runs those searches concurrently, and synthesizes the retrieved results into one answer. Google confirms AI Overviews and AI Mode 'may use a query fan-out technique' issuing multiple related searches across subtopics., not a documented guarantee.
- Architecture is infrastructure, not a ranking shortcut — it can’t rescue thin content.
Official documentation
Primary-source guidance from the search engines on structure, links, and hierarchy.
- In-Depth Guide to How Google Search Works — URL discoveryURL discovery is how search engines find URLs to crawl — by pull (following links and reading sitemaps) and by push (you notify them via IndexNow, the Indexing API, or WebSub). It's the find step that comes before a page is ever fetched. via links, and Google’s own use of “hub page” / category-page language.
- Crawlable links — Make your links crawlable — every important page should be linked from at least one other page; links must be
<a href>; anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. and surrounding context for understanding linked pages. - Importance of Link Architecture (2008) — Google’s long-standing guidance that internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. structure shapes PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. flow and perceived page importance.
Bing / Microsoft
- Bing Webmaster Guidelines — a clear, flat, crawlable structure; reasonable internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.; heading hierarchy (H1An H1 tag is the HTML `<h1>` element that marks a page's primary heading — the big visible headline at the top of the content. It helps users, search engines, and screen readers understand what the page is about./H2/H3); submit sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing..
Quotes from the source
On-the-record statements from Google. Where the source page supports it, each link is a deep link that jumps to the quoted passage.
John Mueller, Google — on hierarchy / pyramid structure
- “The top-down approach or pyramid structure helps us a lot more to understand the context of individual pages within the site.” — John Mueller, on site structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders.. Coverage (Search Engine Journal) · Coverage (Search Engine Roundtable)
John Mueller, Google — on URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. vs. internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.
- Mueller has cautioned that some SEOs over-focus on URL/folder structure; Google relies on internal linkingLinks between pages on the same site. to understand hierarchy, not URL path segments. Coverage (Search Engine Roundtable)
John Mueller, Google — on crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops.
- “I don’t think it would always have a negative effect. I do think if you make it too deep, then that makes it harder for us to crawl and harder for us to pass the signals around.” — John Mueller, on deep page hierarchies.
Google Search Central docs — on discovery via hub pages
- “Other pages are discovered when Google extracts a link from a known page to a new page: for example, a hub page, such as a category page, links to a new blog post.” Jump to quote
Google Search Central docs — on linkability
- “Every page you care about should have a link from at least one other page on your site.” Jump to quote
Gary Illyes, Google — on crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index. and AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.
- “MAKE THAT DAMN SITE CRAWLABLE.” — Gary Illyes (Reddit AMA), underscoring that discoverability via links is the priority over any specific architecture brand.
- Illyes has said AI search optimizationAI search optimization is the practice of making your brand and content visible, citable, and accurately represented across AI-powered search — Google AI Overviews, ChatGPT, Perplexity, Copilot. It's built on traditional SEO plus a heavier emphasis on off-site brand mentions and content AI systems can cite. requires only normal SEO — well-structured, crawlable, high-quality content — with no special architecture changes needed.
Site-architecture checklist
A pass to confirm your structure helps crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., context, and link flow:
- Clear top-down hierarchy: home page → hub/category pages → specific pages.
- Every important page is linked from at least one other page (no orphans).
- Important pages aren’t buried needlessly deep (no fixed click-depth rule — check crawl data for pages getting crawled less as they get deeper).
- Each topic area has a hub/pillar page that links to its subtopic pages.
- Subtopic pages link back to their hub and to relevant sibling pages.
- Cross-cluster links exist where topics are genuinely related (no strict silo walls).
- Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. use descriptive anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. with relevant surrounding context.
- No important page relies on click-only navigation — links are real
<a href>. - URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. is logical for humans, but you’re not relying on folders to signal architecture (links do that).
- Crawl data reviewed for over-deep pages, orphans, and internal-link gaps.
The mental models
1. The link graph IS the architecture. URL folders are for humans and housekeeping. Google works out hierarchy and context from how pages link to each other. Audit and design the link graph, not the folder tree.
2. Three names, one structure. Silo, hub and spoke, topic clusters → all converge on a grouped, hierarchical, internally linked pyramid once you allow sensible cross-links. Stop shopping for the “right” model; build the cluster and link it well.
3. Keep the kernel, drop the wall. From silos, keep topical concentration (link densely within a topic). Drop the isolation rule (never link between silos) — no Google source requires it, and it hurts users.
4. The depth dial. Too flat dilutes topical signal; too deep starves pages of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and link authority. Google hasn’t published a click-depth number, so aim for shallow-with-strong-linking judged against your own crawl data, grouped into hubs, rather than a fixed target.
5. Architecture is a multiplier, not a source. Good structure multiplies the value of good content by making it discoverable and contextually clear. Multiply zero (thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count.) and you still get zero.
Architecture model reference
| Model | Useful idea | Mistake to avoid | Practical implementation |
|---|---|---|---|
| Pyramid | Broad pages lead to increasingly specific pages | Burying important pages too deep | Home → hub/category → detail page |
| Hub and spoke | A central overview organizes related subtopics | Linking only outward from the hub | Hub ↔ spokes, plus relevant spoke links |
| Topic cluster | Content coverage is planned around a subject | Treating the label as a different architecture model | Use the same hub-and-spoke graph |
| Silo | Related pages receive concentrated topical links | Forbidding every cross-topic link | Keep topical grouping; allow useful cross-links |
What each layer should do
| Layer | Primary job | Audit question |
|---|---|---|
| Home | Route users and crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. into major areas | Are the site’s real priorities represented? |
| Hub or category | Explain the group and link to its members | Can every important spoke be reached? |
| Detail page | Satisfy a specific intent and reinforce its context | Does it link back and to genuinely related pages? |
| Cross-link | Connect related needs across groups | Would a reader naturally follow it? |
Test yourself: Site Architecture
Five quick questions on silos, hubs, clusters, and what Google actually recommends. Pick an answer for each, then check.
Resources worth your time
My related writing
- The Beginner’s Guide to Technical SEO — where site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. and internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. fit in the bigger picture.
- Internal Links for SEO: An Actionable Guide — how internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. pass signals and shape the link graph that is your architecture.
- SEO Silo Structure: Why It Makes No Sense — the Ahrefs contrarian take on strict silos (the case for dropping the wall).
- How to Build a Topic Cluster — the practical hub-and-spoke / cluster build, step by step.
My speaking
- How Search Works (SlideShare) — my walkthrough of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and ranking, the pipeline architecture has to serve. (Standing disclaimer applies: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From around the industry
- Stop the Silo Madness: Effective Site Architecture for SEO and Findability (Shari Thurow, Search Engine Land) — the definitive case against strict silos.
- Site Architecture for SEO: Structure That Ranks & Scales (Search Engine Land) — a thorough structure guide.
- Complete Guide to Topic Clusters (Search Engine Land) — the cluster model in depth.
- Topic Clusters: The Next Evolution of SEO (HubSpot) — the 2017 origin of the “topic cluster” terminology.
- SEO Content Strategies: The Hub and Spoke Model (Botify) — the same structure under its information-architecture name.
- John Mueller Recommends Pyramid Site Structure (Search Engine Journal) — the source for the pyramid/context quote.
Site Architecture
Site architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models.
Related: Internal Link, Website Structure
Site Architecture
Site architecture (also called website structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders. or information architecture) is how the pages on a site are grouped, categorized, and linked together. A good architecture makes every important page reachable, concentrates internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. within topic areas, and gives search engines clear context for what each page is about. Google recommends a pyramid / top-down hierarchy — home page at the top, category or hub pages in the middle, specific content pages at the bottom.
Three organizational models dominate SEO discussion:
- Silo structureA silo structure is a site organized by grouping related content into topic areas and concentrating internal links within each topic so a hub page and its supporting pages reinforce each other. Modern practice keeps the topical concentration and drops the old rule against cross-linking between topics. — content grouped into isolated thematic sections; the strict version prohibits cross-silo links. Popularized by Bruce Clay.
- Hub and spoke (also: content hub, pillar-and-cluster) — a broad overview “hub” page links out to detailed “spoke” pages and back.
- Topic clusters — a content-marketing model (popularized by HubSpot in 2017) where a pillar page targets a broad term and cluster pages target subtopics, all interlinked.
The key thing most write-ups miss: hub and spoke and topic clusters are functionally the same model — different labels from the information-architecture and content-marketing communities. And what actually moves the needle is internal linkingLinks between pages on the same site., not URL folder structure: Google leans on contextual links to understand hierarchy, so strict physical silos that block relevant cross-links tend to hurt more than help.
Related: Internal Link, Website Structure
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Reworked the article against the structured evidence record: led with a crawlable-discovery-and-navigation definition instead of URL folders, labeled silo/hub-and-spoke/topic-cluster as overlapping practitioner terms rather than a single Google-endorsed model, removed invented click-depth thresholds (Google hasn't published a specific number), converted topical-authority and AI-search benefits from stated outcomes to hedged hypotheses, and regrounded the strict-silo critique in user navigation and relevance instead of assumed authority mechanics.
Change details
-
Added the crawlable-link contract (standard <a href> elements with resolvable hrefs) and separated discovery from guaranteed crawling/indexing/ranking outcomes in the beginner and advanced lenses.
-
Removed the ~3-clicks framing (an invented threshold Google hasn't published) from the beginner, advanced, checklist, frameworks, ai-summary, and audit sections; replaced with crawl-data-driven guidance.
-
Reframed topical-authority and AI-search query-fan-out benefits as hypotheses to test rather than established mechanisms, and updated the quiz explanations to match.
Full comparison unavailable — no prior snapshot was archived for this revision.