Website Structure
How a site's hierarchy, navigation, breadcrumbs, and URLs are organized — and why internal linking is the primary signal Google reads to understand it. The hub.
1 evidence signal on this page
- Related live toolGoogle Index Checker
Website structure is your site's hierarchy, navigation, breadcrumbs, and URL organization — how pages relate to each other and how people and search engines move through them. The thing most guides underweight: internal linking is the primary signal Google reads to understand that structure, not the slashes in your URLs. Click depth matters more than URL depth. A reasonable pyramid beats both an artificially flat site and a needlessly deep one, and none of this guarantees a specific crawl frequency or ranking outcome by itself. This hub explains how Google actually reads structure and points you to the deep dives: URL structure, internal links, orphan pages, breadcrumbs, pagination, crawl depth, site-architecture models, and subdomain vs subdirectory.
Evidence for this claim Google recommends organizing a site logically so users and search engines can understand relationships between pages. Scope: Current Google SEO Starter Guide. Confidence: high · Verified: Google Search Central: SEO Starter Guide Evidence for this claim Crawlable internal links provide discovery paths; sitemaps can supplement discovery but are not a substitute for site navigation. Scope: Current Google crawling and sitemap guidance. Confidence: high · Verified: Google Search Central: Make links crawlableTL;DR — Website structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders. is your hierarchy, navigation, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and URL organization — how your pages are grouped and how people and search engines move between them. Search engines find your pages mainly by following links, so the big lever isn’t your URL folders — it’s your internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. Group related content, link to your important pages, and keep them a few clicks from the homepage.
What website structure is
Website structure (a lot of people call it site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models.) is the whole visible shape of your site: your hierarchy, your navigation and menus, your breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and how your URLs are organized. Think of it two ways at once:
- For people — can a visitor land on your site and quickly find what they came for, using menus, categories, and breadcrumbs?
- For search engines — can a botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. like GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. follow links from page to page and discover everything you want foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore.?
The classic shape is a pyramid: your homepage at the top, broad categories under it, narrower subcategories under those, and individual pages at the bottom. A shop might go Home → Shoes → Running Shoes → a specific shoe. That grouping helps both humans and Google understand how your content fits together.
The one idea most people miss
Here’s the thing that surprises people: your URL folders don’t decide your
structure by themselves — your internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. do the heavy lifting. It’s
tempting to think a page at example.com/shoes/running/model-x/ is “deeper”
than example.com/model-x/, but Google doesn’t count slashes to decide what’s
important. Internal linkingLinks between pages on the same site. is the main signal it reads to understand how your
pages relate to each other.
So the question that actually matters is: how many clicks does it take to get from your homepage to a page? That’s click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., and it’s the thing to optimize — not the number of slashes in the URL.
How to get it right (the simple version)
- Link to your important pages from your homepage and main navigation.
- Group related content into clear categories.
- Don’t bury pages — important ones should be reachable in a few clicks.
- Use real links (
<a href="…">), not buttons or menus that only work with JavaScript clicks — bots may not follow those. - Don’t leave orphans — a page nothing links to may never be found.
- Add a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. as a backup so search engines have a list of your URLs.
The thing most people get wrong
Flatter is not automatically better. A common myth is that you should squash everything close to the homepage. But if every page is one click away, you’ve also told Google that everything is equally important — which removes the grouping signals it uses to understand your site. A sensible pyramid beats a pancake.
One honest caveat: cleaning up your structure doesn’t guarantee faster crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., better rankings, or more traffic on its own. It removes friction and makes your important pages easier to find — the rest still depends on the content on those pages.
Want the deeper version — Mueller’s actual quotes, click depth vs URL depth, the mega-menu trap, and where each subtopic lives? Switch to the Advanced tab.
Evidence for this claim Google recommends organizing a site logically so users and search engines can understand relationships between pages. Scope: Current Google SEO Starter Guide. Confidence: high · Verified: Google Search Central: SEO Starter Guide Evidence for this claim Crawlable internal links provide discovery paths; sitemaps can supplement discovery but are not a substitute for site navigation. Scope: Current Google crawling and sitemap guidance. Confidence: high · Verified: Google Search Central: Make links crawlableTL;DR — Website structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders. is your site’s hierarchy, navigation, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and URL organization. Internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. is the primary signal Google reads to understand it — not URL folders. Mueller is explicit: “we don’t care so much about the folder structure.” Click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. (links from the homepage) matters; URL depth (slash count) doesn’t, and an artificially flat URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. buys you nothing. A reasonable pyramid beats both extremes — over-flattening (e.g. mega menus) strips the grouping signals Google uses for context. Google infers a page’s relative importance from internal linkingLinks between pages on the same site. patterns. Orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. get no internal discovery path and may never be crawled. BreadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results. should reflect a typical user path, not mechanically mirror the URL. Choose subdomains or subdirectories based on organizational needs rather than assuming the URL form alone creates a penalty. SitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. help discovery but don’t replace links, and none of this guarantees a specific crawl frequencyCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial., ranking, or traffic outcome.
Internal linking is how Google reads your structure
Website structure itself is broader than any one signal: it’s your site’s hierarchy, your navigation and menus, your breadcrumbs, and how your URLs are organized — the whole visible shape a person or a crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. encounters. But the thing most guides underemphasize is how Google actually reads that shape: primarily by analyzing your internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. graph, not by parsing URL paths. John Mueller has said it about as plainly as it gets: “For us, we don’t care so much about the folder structure, we really essentially focus on the internal linking.” And on the slashes specifically: “Just looking at the number of slashes, for example, in a URL doesn’t tell us that this is lower level or higher level.”
URLs often reflect architecture — a tidy /category/subcategory/page/ path
usually mirrors the link structure — but that’s correlation, not causation. As
Mueller put it, “a lot of times the architecture of the website is visible in the
URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly., but it doesn’t have to be the case.” Google’s ecommerce docs say
the same thing from the engine’s side: “Google tries to find the best content on
your site by analyzing the relationship between pages based on their linkages.”
Once you internalize this, most “structure” decisions get clearer: you’re not designing folders, you’re designing a link graph.
Click depth, not URL depth
Because Google reads links, the metric that matters is click depth — how many links a crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. follows to reach a page from the homepage — not how many segments the URL has. Mueller again: “It’s really like from the homepage or from the primary page, how quickly can we reach that specific page?”
The corollary is counterintuitive and worth saying out loud: there is no SEO benefit to an artificially flat URL structure. Stripping folders out of your URLs while the link graph stays deep changes nothing for Google. If you want a buried page to count for more, link to it from closer to the homepage — don’t rewrite its URL.
None of that makes URLs pointless. Google’s own URL-structure guidance still recommends simple, descriptive URLs — readable words instead of long ID strings, hyphens instead of underscores — because they help people understand a page before they click, and because parts of a URL can display as breadcrumbs in search results. That’s a real, documented role for URLs: readability and maintainability for humans. It’s just a different measurement from click depth, and shouldn’t be confused with it.
Pyramid beats both extremes
Flat is not the goal. Neither is deep. Mueller has been explicit that a hierarchical structure is what helps: “a pyramid structure helps us a lot more to understand the context of individual pages,” and “it’s not the case that a super flat structure is going to be better than a reasonable pyramid.” He also flags the over-deep failure mode: “you don’t want it to be such that it’s like you have to click through a million times.”
The over-flattening trap is real and underdiscussed. Mega menusA mega menu is a large, categorized navigation panel — usually opened by hover or click on a top-level nav item — that surfaces many links at once, grouped into columns instead of the single list a standard dropdown shows. that surface hundreds of links one click from the homepage can flatten your structure so much that, in Mueller’s words, “Google can’t recognize which parts of the site belong together” — so there’s value in reducing crawl height without pancaking the site. A reasonable pyramid (home → category → subcategory → page) keeps the grouping signals intact.
Importance flows through links
Google infers a page’s relative importance from your internal links. Its ecommerce guidance is direct: “the more links a page has to it within a site, the higher the relative importance of the page,” and Google uses “the number of links it needs to follow to reach a page and the number of links to a page to infer the relative importance of a page.” Two practical levers fall out of that: link to priority pages more, and link to them from closer to the homepage.
Two implementation details Google calls out: use real <a href> links —
“don’t use JavaScript events on other HTML DOM elements for navigation” — and
don’t rely on internal site search for discovery, because “GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. generally
doesn’t try to submit searches into a search box as part of crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. a site.”
Where directories still earn their keep
URL folders aren’t pure decoration. For large sites, Google does learn
crawl behavior at the directory level — it can crawl /news/ more often than
/archive/ because it learns how often URLs in each directory change. Gary Illyes
has noted Google “prefers a hierarchical structure for large sites,” and that
clean architecture matters because Google is crawling fewer pages and you need to
signal which are the priorities. So organize by topic and, where you can, by
update frequency — but remember the crawl-importance signal still comes from the
link graph, not the path.
Sitemaps help — but they don’t replace links
XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. are a real discovery channel; Illyes has called them the second most important way Google finds URLs — with internal links first. A sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. is a list for discovery, not a stand-in for navigation or linking — Google is explicit that submitting a sitemap doesn’t replace crawlable links to your important pages. If a page is reachable only via the sitemap, it’s still missing the navigation and internal-link context that helps people and Google understand where it fits and why it matters. Submit the sitemap as a backup, but fix the links.
This works on the page set and inlinks you supply. Use the score to prioritize human review, then use a crawler to confirm reachability and click depth across the full site.
Build an internal-link candidate queue with my free Internal Link Cluster Visualizer Free
- Supply priority pages, likely hubs, and their current internal links.
- Review candidates where topical reasons are visible and the target has few supplied inlinks.
- Add only contextually justified links and verify the new structure with a fresh crawl.
The row scores a proposed link at 76.6 between an internal-linking page and a topic-cluster page. It lists shared terms, a common URL segment, and low or absent supplied inlinks as the reasons.
Breadcrumbs: user path, not URL mirror
Breadcrumbs are one of the places people over-mechanize structure. Google’s own breadcrumb guidance describes them as showing a page’s position in the site hierarchy and recommends representing a typical user path — it does not say a breadcrumb trail has to mechanically reproduce the URL segments. In practice that matters most on pages reachable through more than one legitimate path (a product filed under two categories, say): pick the path that best represents how a typical visitor got there, mark it up consistently with structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding., and don’t treat “does it match the URL folders” as the test.
What structure doesn’t guarantee
Worth saying plainly: reorganizing your menus, directories, breadcrumbs, or internal links doesn’t by itself guarantee a specific crawl frequencyCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial., faster indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., more PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems., higher rankings, more traffic, or citations in AI answers. Even Google’s own sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. — arguably the most visible reward for a clean structure — are generated automatically and not guaranteed; a logical, well-linked site makes them more likely, but there’s no lever that forces them. Good structure removes friction and makes priority pages easier to find and understand. What happens after that still depends on the content itself.
Bing recommends shallow reachability
Where Google talks specifically in terms of link-graph proximity, Bing’s webmaster guidance is more general: it recommends a well-organized hierarchy and keeping important pages easily reachable from the homepage, and it leans heavily on XML sitemaps for discovery. I wasn’t able to recover a current, readable primary-source page with a specific click-count figure for Bing (its webmaster guidelines page is a JavaScript application shell with no fetchable click-depth text as of this update), so treat any specific number you see elsewhere as a secondary claim, not a documented Bing rule. For your own site, the practical target is the same one that applies to Google: keep priority pages reachable in a small, deliberate number of clicks rather than chasing an exact count. Exact click-depth distributions and audit thresholds are covered on the crawl-depth page.
Where to go next
This page is the map. Each subtopic below is its own deep dive (they’re in the sidebar too):
- URL Structure — readable, descriptive URLs; hyphens not underscores; case sensitivity; and why the URL describes hierarchy rather than creating it.
- Internal Links — the primary structural signal Google reads, how anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. communicates a page’s topic, and how to route importance to priority pages. (Cross-listed from On-Page SEO.)
- Orphan PagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. — pages no internal link points to: why they go undiscovered, earn no link equity, and quietly waste crawl budgetThe number of URLs an engine will crawl in a timeframe. — plus how to find and fix them. (Nested under Internal Links.)
- Breadcrumbs — reinforcing hierarchy for users and search engines, the structured-data markup, and why Google recommends a typical user path rather than a mechanical mirror of the URL.
- PaginationPagination splits a large set of content — product listings, blog archives, search results — across multiple sequentially numbered URLs. For SEO, each paginated page should be crawlable, indexable, and self-canonical; Google no longer uses rel=prev/next, but Bing still does. — how paginated sets (category listings, archives) fit into structure and how Google handles them today.
- Site ArchitectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. (models) — the pyramid baseline, plus silo, hub-and-spoke, and topic-cluster models, and when each makes sense.
- Subdomain vs SubdirectoryA subdomain (blog.example.com) is a separate hostname; a subdirectory (example.com/blog/) is a path on the same hostname. Google has no blanket ranking preference — it decides per site whether a subdomain is treated as part of the site, based on integration signals. — Google treats them the same algorithmically; Mueller’s practical advice to keep related content together unless it’s “really kind of slightly different.”
- Crawl DepthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. — click depth vs crawl-traversal depth and why important pages belong close to the homepage. (Cross-listed from How Search WorksSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank..)
For the wider context, see Technical SEOTechnical SEO is the practice of making a site easy for search engines to crawl, render, index, and (now) be eligible for AI answers. It's the foundation that lets your content and links rank — not a ranking trick of its own. and How Search WorksSearch works in three stages — crawling, indexing, and serving (ranking). A page has to clear each one to appear in results: getting crawled doesn't mean you're indexed, and getting indexed doesn't mean you rank..
AI summary
A condensed take on the Advanced version:
- Website structureWebsite structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders. = the visible IA. Hierarchy, navigation, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and URL organization. Internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. is the primary signal Google reads to understand it, not URL folders. Mueller: “we don’t care so much about the folder structure… we really essentially focus on the internal linkingLinks between pages on the same site..”
- Click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., not URL depth. What matters is how many links from the homepage reach a page. An artificially flat URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. provides no SEO benefit — link closer to the homepage instead of rewriting URLs. URLs still have a real, documented role: readable and descriptive for people, and shown as breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results. in search results — just a different measurement from click depth.
- A reasonable pyramid beats both extremes. Flat is not automatically better; Google reads excessive flatness as “everything is equally important,” removing context. Over-deep is bad too. Mega menusA mega menu is a large, categorized navigation panel — usually opened by hover or click on a top-level nav item — that surfaces many links at once, grouped into columns instead of the single list a standard dropdown shows. can over-flatten and hide which pages belong together.
- Importance flows through links. The more internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to a page (and the
closer it sits to the homepage), the higher its inferred importance. Use real
<a href>links; don’t rely on site search or JS-event navigation for discovery. - Directories still matter at scale for directory-level crawl frequencyCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial., but the importance signal is the link graph, not the path.
- Breadcrumbs represent a typical user path, not a mechanical mirror of the URL — that’s Google’s own framing, useful when a page has more than one legitimate path to it.
- SitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. help but don’t replace links — Illyes ranks internal links first, sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. second. A sitemap is a discovery backup; it isn’t a substitute for navigation or internal links.
- No guarantees. Reorganizing structure doesn’t by itself guarantee crawl frequency, indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. speed, rankings, traffic, or AI citationsAn AI citation is the visible source link an AI answer engine shows next to its generated text — the clickable reference that credits the web page it used. A citation's presence is a separate thing from whether the cited page actually supports the statement, and from being retrieved (read behind the scenes) or merely mentioned (named without a link); citation is driven more by brand mentions and being retrievable than by traditional ranking. — even sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. are automatic and not guaranteed.
- Subdomain vs subdirectoryA subdomain (blog.example.com) is a separate hostname; a subdirectory (example.com/blog/) is a path on the same hostname. Google has no blanket ranking preference — it decides per site whether a subdomain is treated as part of the site, based on integration signals.: treated the same algorithmically; decide on organization. Bing recommends shallow reachability generally; I couldn’t recover a current primary source with a specific click-count figure.
- Subtopics: URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly., internal links, orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site., breadcrumbs, paginationPagination splits a large set of content — product listings, blog archives, search results — across multiple sequentially numbered URLs. For SEO, each paginated page should be crawlable, indexable, and self-canonical; Google no longer uses rel=prev/next, but Bing still does., site-architecture models, subdomain vs subdirectory, crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops..
Official documentation
Primary-source guidance from the search engines.
- SEO Starter Guide — directory-based organization, descriptive URLs, and internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. with meaningful anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page..
- Help Google understand your ecommerce site structure — how Google infers structure from linkages, the hierarchical linking pattern, and using
<a href>tags. - URL structure — hyphens vs underscores, readable words, case sensitivity, and parameter handling.
- Sitelinks — how a logical, well-linked structure improves sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. quality.
- Build and submit a sitemap — the backup discovery channel.
- Optimize your crawling and indexing (2009, still referenced) — the one-URL-to-one-content ideal.
Bing / Microsoft
- Bing Webmaster Guidelines — well-organized hierarchy, easily reachable important pages, clean URLs, and sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.-first discovery. (This page is a JavaScript application shell; I couldn’t recover fetchable primary-source text on a specific click-count figure as of this update.)
Quotes from the source
On-the-record statements from Google. Each link is a deep link that jumps to the quoted passage on the source page.
John Mueller — internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. vs URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly.
- “For us, we don’t care so much about the folder structure, we really essentially focus on the internal linkingLinks between pages on the same site..” Read the coverage
- “Just looking at the number of slashes, for example, in a URL doesn’t tell us that this is lower level or higher level.” Read the coverage
- “It’s really like from the homepage or from the primary page how quickly can we reach that specific page?” Read the coverage
John Mueller — pyramid vs flat
- “A pyramid structure helps us a lot more to understand the context of individual pages.” Read the coverage
- “It’s not the case that a super flat structure is going to be better than a kind of reasonable pyramid.” Read the coverage
John Mueller — artificially flat URLs
- Google says there is “no SEO benefit for having an artificially flat URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly.” compared with a site that has directory depth. Read the coverage
John Mueller — subdomains vs subdirectories
- “In general, we see these the same… I would personally try to keep things together as much as possible… Use subdomains where things are really kind of slightly different.” Read the coverage
Gary Illyes — hierarchy and discovery
- Google “prefers a hierarchical structure for large sites,” and clean architecture matters because Google is crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. fewer pages — with XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. the second most important discovery method after internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. Read the coverage
#:~:text= fragments target the secondary coverage and should be confirmed against the original source before being treated as final. Website-structure checklist
A pass to confirm your structure works for both visitors and crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.:
- Important pages are linked from the homepage and/or main navigation.
- Content is grouped into clear categories (a sensible pyramid, not a pancake and not a maze).
- Priority pages sit a small, deliberate number of clicks from the homepage — click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. is low (no site-wide fixed number; see the crawl-depth page for measuring your own distribution).
- No orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. — every page you want foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. has at least one internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. pointing to it.
- Navigation uses real
<a href>links, not JS-event clicks or buttons. - You’re not relying on internal site searchEcommerce site search is the on-site search box that lets shoppers query a store's catalog directly. For SEO it's a two-sided topic: the feature itself is a conversion tool, but the results pages it generates are a classic source of crawl waste, index bloat, and near-duplicate URLs that should usually be kept out of Google's index. for discovery of important pages.
- Internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. describes the destination’s topic.
- Mega menusA mega menu is a large, categorized navigation panel — usually opened by hover or click on a top-level nav item — that surfaces many links at once, grouped into columns instead of the single list a standard dropdown shows. aren’t over-flattening the structure (grouping signals intact).
- BreadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results. are present, marked up with structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding., and represent a typical user path rather than a mechanical copy of the URL.
- Subdomain vs subdirectoryA subdomain (blog.example.com) is a separate hostname; a subdirectory (example.com/blog/) is a path on the same hostname. Google has no blanket ranking preference — it decides per site whether a subdomain is treated as part of the site, based on integration signals. chosen on content organization, not an imagined ranking penalty (keep related content together by default).
- XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. submitted in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. — as a backup to links, not a replacement.
The mental models
1. Structure is visible; internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. are how Google reads it. Your structure is the hierarchy, navigation, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and URLs people actually see. Before you touch folders, look at what links to what — Google reads relationships between pages “based on their linkages.” Your URL path can describe the hierarchy, but it doesn’t create it.
2. Optimize click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., not URL depth. The question is “how many links from the homepage to reach this page?” — not “how many slashes?” Want a page to count for more? Link to it from closer in. Rewriting its URL flatter does nothing.
3. The pyramid, between two failure modes. Aim for home → category → subcategory → page. Too flat (mega menusA mega menu is a large, categorized navigation panel — usually opened by hover or click on a top-level nav item — that surfaces many links at once, grouped into columns instead of the single list a standard dropdown shows., everything one click away) strips grouping/context signals; too deep (click through a million times) buries pages. The pyramid is the middle path Google actually rewards.
4. Importance is a function of links. More internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. to a page, and shorter link distance from the homepage, = higher inferred importance. That’s your steering wheel: point links at what matters.
5. Discovery = links first, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. second. Internal links are how pages get foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. and how importance is signaled. A sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. is a backup list — necessary, but it carries no importance signal. Orphans fall through both: link to them.
6. Same-or-different decision (subdomain vs subdirectoryA subdomain (blog.example.com) is a separate hostname; a subdirectory (example.com/blog/) is a path on the same hostname. Google has no blanket ranking preference — it decides per site whether a subdomain is treated as part of the site, based on integration signals.). Google treats them the same, so decide on organization: keep related content together in a subdirectory by default; reach for a subdomain only when the content is genuinely “slightly different.”
7. BreadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results. answer “how did a typical visitor get here,” not “what does the URL say.” When a page has more than one legitimate path to it, pick the one that best represents a normal visit and mark it up consistently — don’t force the trail to match URL folders that don’t reflect real navigation.
8. Structure lowers friction; it doesn’t buy outcomes. A clean structure makes priority pages easier to find and understand for people and crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.. It doesn’t guarantee a crawl frequencyCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial., a ranking, traffic, or an AI citation — those still depend on the content itself.
Patrick's relevant free tools
- Internal Link Cluster Visualizer — Analyze a bounded supplied internal-link graph, orphans, PageRank, and lexical missing-link suggestions.
- SEO Opportunity Finder — Choose an evidence-backed content, link, or technical SEO workflow, then open the free tool that can evaluate it without turning every warning into an opportunity.
- Semantic Site Map — Explore this site's build-time semantic center, topic coherence, article outliers, and nearest editorial neighbors.
Tools for seeing the structure search engines receive
- Google Index Checker — review observable status, robots, and canonical blockers on important pages, then hand off to Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. for Google’s recorded state.
- XML Sitemap Validator — catch malformed sitemap filesA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and URLs that conflict with the crawlable structure you intended to publish.
- Ahrefs Site Audit or another link crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — inspect crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., inlinks, outlinks, orphan candidates, redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., and the actual graph produced by templates.
- A crawl visualization — use a directory tree for operational URL organization and a force-directed link graph for architecture; do not confuse the two pictures.
- Analytics and Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance. — join demand and performance with the crawl so high-value pages buried in the graph receive priority.
Start every audit from normal entry points such as the home page. A list crawl tests URLs, but it does not reveal whether users or bots can reach them through internal links.
Test yourself: Website structure
Five quick questions on how search engines actually read your site’s structure. Pick an answer for each, then check.
Resources worth your time
My related writing
- Internal Links for SEO: An Actionable Guide — internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. as the primary mechanism that transmits structure signals.
- What Are Sitelinks, Their Benefits & How To Influence Them — how site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. feeds Google’s sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them., and why flat structures make sitelinksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. harder.
- What Is an Enterprise SEO Audit & How To Do One — auditing site structure at scale.
- The Beginner’s Guide to Technical SEO — where structure sits in the bigger technical picture.
My speaking
- Patrick Stox on SlideShare — talks including A Crash Course in Technical SEOTechnical SEO is the practice of making a site easy for search engines to crawl, render, index, and (now) be eligible for AI answers. It's the foundation that lets your content and links rank — not a ranking trick of its own., Website MigrationsA site migration is any significant change to a website's URL structure, domain, platform, protocol, or hosting that can affect how search engines crawl, index, and rank it. The risk scales with how much you change at once. (SMX Munich), and International SEOInternational SEO is the practice of optimizing a site so search engines understand which countries and/or languages it targets, and serve the right version to each user. It spans URL structure, hreflang, and on-page localization.: The Weird Technical Parts (Pubcon), all of which touch site-architecture audits and redesigns. (My standing disclaimer applies: this is my understanding of these systems, not an official spec.)
From around the industry
- How to Structure Your Website Architecture for SEO (Ahrefs, Chris Haines) — types, tools, and technical elements of structure.
- Site structure: the ultimate guide (Yoast) — comprehensive, WordPress-centric walkthrough.
- Website Architecture (Backlinko) — hub-style reference on architecture concepts.
- How to build a better, smarter, more discoverable site architecture (Search Engine Land, Anna Crowe) — detailed how-to with a case study, and Illyes context on hierarchy and crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor..
- Site architecture: Creating a website structure that ranks (Search Engine Land) — industry reference guide.
- Pyramid Site Structure (Search Engine Journal) — Mueller’s pyramid-vs-flat and mega-menu guidance.
- Google Treats Subdomains & Subdirectories The Same, John Mueller Says (Search Engine Journal) — the subdomain vs subdirectoryA subdomain (blog.example.com) is a separate hostname; a subdirectory (example.com/blog/) is a path on the same hostname. Google has no blanket ranking preference — it decides per site whether a subdomain is treated as part of the site, based on integration signals. question.
- Site Architecture for SEO: The Definitive Guide (Impression Digital) — comprehensive implementation guide.
Website Structure
Website structure (site architecture) is a site's visible hierarchy, navigation, breadcrumbs, and URL organization — how pages relate and how people and search engines move between them. Internal linking is the primary signal Google reads to understand that structure, not URL folders.
Related: URL Structure, Internal Link, Crawl depth
Website Structure
Website structure — also called site architectureSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models. — is the way a website’s pages are organized, interconnected, and presented: its hierarchy, navigation and menus, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., and URL organization. It determines how visitors navigate and, just as importantly, how search engine crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. discover, traverse, and interpret your pages. A good structure groups related content into logical hierarchies, keeps important pages a short click away, and surfaces internal-link signals that tell search engines which pages matter most.
The single most misunderstood thing here: internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. is the primary signal Google reads to understand that structure. Google’s John Mueller has been explicit that the slashes in a URL don’t define hierarchy — what Google reads is the link graph. “We don’t care so much about the folder structure, we really essentially focus on the internal linkingLinks between pages on the same site..” Two sites with identical URLs can have completely different structures depending on how their pages link to each other.
That reframes most “structure” decisions. Click depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. (how many links from the homepage to a page) matters; URL depth (slash count) does not. A pyramid of home → category → subcategory → page helps Google understand context; both an artificially flat structure and a needlessly deep one work against you. Orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. with no internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. may never be crawled. And subdomains vs subdirectories are treated the same algorithmically — the choice is about organization, not a ranking penalty.
Related building blocks include URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly., internal links, breadcrumbsBreadcrumbs are a secondary navigation trail (Home > Category > Page) that shows where a page sits in a site's hierarchy. They create internal links that pass PageRank, and when marked up with BreadcrumbList structured data they can drive the path Google shows in desktop search results., paginationPagination splits a large set of content — product listings, blog archives, search results — across multiple sequentially numbered URLs. For SEO, each paginated page should be crawlable, indexable, and self-canonical; Google no longer uses rel=prev/next, but Bing still does., crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops., and the architecture models (silo, hub-and-spoke, topic clustersSite architecture is how a website's pages are organized, categorized, and interlinked. It controls how crawlers discover pages, how link equity flows, and how clearly search engines understand each page's topical context. Silo structure, hub and spoke, and topic clusters are the three common models.) that shape how all of those fit together.
Related: URL Structure, Internal Link, Crawl depth
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Broadened the definition so website structure covers the visible hierarchy, navigation, breadcrumbs, and URLs (internal linking is now framed as the primary signal Google reads, not the whole definition); removed the unverifiable 'Bing three-click rule' everywhere it appeared after a raw-fetch check found no recoverable primary-source text; added explicit breadcrumb (typical user path, not URL mirror) and no-guaranteed-outcomes guidance; and tightened the sitemap and URL-role passages to match documented Google guidance.
Change details
- Before
Bing is stricter on click depth: it recommends keeping important pages within three clicks of the homepage... the three-click rule is a useful hard target.AfterRemoved the specific three-click figure (no recoverable current primary source); the article now says Bing recommends shallow reachability generally and routes exact click-depth thresholds to the crawl-depth page. -
Added a 'Breadcrumbs: user path, not URL mirror' section stating Google's guidance to represent a typical user path rather than mechanically mirror the URL, plus the multiple-valid-paths edge case.
-
Added a 'What structure doesn't guarantee' section clarifying that reorganizing structure doesn't by itself guarantee crawl frequency, indexing, rankings, traffic, or AI citations.
- Before
If a page is reachable only via the sitemap, you've signaled almost nothing about its importance, and orphaned pages get no link equity.AfterIf a page is reachable only via the sitemap, it's still missing the navigation and internal-link context that helps people and Google understand where it fits and why it matters.
Full comparison unavailable — no prior snapshot was archived for this revision.