Orphan Pages
Orphan pages have no internal links, so search engines struggle to find them and they get zero internal PageRank. How to spot and fix them.
1 evidence signal on this page
- Related live toolGoogle Index Checker
An orphan page is one nothing on your site links to internally. Because Google discovers pages mainly by following links, an orphan with no sitemap entry and no backlinks is effectively invisible — and even when it's found via a sitemap, it gets zero internal PageRank and often lands in 'Discovered — currently not indexed.' A crawler alone can't find orphans (it only follows links it can reach); you have to compare the crawl against your sitemap, analytics, and backlink data. Fix the ones you want ranked by adding contextual internal links; noindex the intentional ones (PPC and thank-you pages); redirect or delete the rest.
Evidence for this claim Google discovers pages through links and recommends that every page you care about have a link from at least one other page on the site. Scope: Current Google link and discovery guidance. Confidence: high · Verified: Google Search Central: Link best practices Evidence for this claim Sitemaps can help search engines discover URLs but do not guarantee crawling or indexing and do not replace navigational links. Scope: Current Google sitemap behavior. Confidence: high · Verified: Google Search Central: Learn about sitemapsTL;DR — An orphan page is a page on your site that nothing else links to. Search engines find pages by following links, so an orphan is easy to miss — it often doesn’t get crawled, indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., or any traffic. The fix is usually simple: link to it from a few relevant pages on your site.
What an orphan page is
Picture your website as a network of rooms connected by doorways. The doorways are your internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. An orphan page is a room with no doorways leading into it — you can only get there if you already know the exact address.
That’s a problem because search engines don’t know the address ahead of time. They find pages mainly by following links: a botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. lands on a page, grabs the links on it, and visits those next. If no page links to a page, the bot has no trail to follow, so the page can go unnoticed — and a page that never gets foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. can’t show up in search.
Why orphan pages hurt
- They’re hard to discover. No internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. means no link trail for Google to follow.
- They miss out on “link juicePageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems..” Links pass authority around your site. An orphan gets none of it, so even good content struggles to rank.
- They often sit unindexed. A common Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. status, “Discovered — currently not indexed,” frequently shows up for orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site.: Google knows the URL exists but doesn’t think it’s worth indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..
How a page ends up orphaned
It’s rarely on purpose. The usual culprits:
- You redesigned the site or changed the menu and a page got dropped from the navigation.
- A blog post never got assigned to a category or tag.
- A product was discontinued but the page was left live.
- A page only ever existed as a one-off (a campaign or ad landing page).
How to find and fix them
You can’t find orphans just by crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. your site — a crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. only sees pages it can reach by following links, which by definition excludes orphans. Instead you compare two lists: the pages a crawler found, and every page you know exists (your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., your analytics, Search ConsoleGoogle's free tool for monitoring crawling, indexing, and search performance.). Anything on the second list but not the first is an orphan.
To fix one you actually want to rank: add a few internal links to it from
related pages, using clear anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. that describes the page. If it’s a page
you don’t want in search (like a thank-you page after a purchase), the answer
isn’t links — it’s a noindex tag. And if it’s old and useless, redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. it to
something relevant or remove it.
Want the data, the Google quotes, and the exact audit workflow? Switch to the Advanced tab.
Evidence for this claim Google discovers pages through links and recommends that every page you care about have a link from at least one other page on the site. Scope: Current Google link and discovery guidance. Confidence: high · Verified: Google Search Central: Link best practices Evidence for this claim Sitemaps can help search engines discover URLs but do not guarantee crawling or indexing and do not replace navigational links. Scope: Current Google sitemap behavior. Confidence: high · Verified: Google Search Central: Learn about sitemapsTL;DR — Google discoversGoogle Discover is a personalized, mobile-first content feed built into the Google app, Chrome's mobile New Tab page, and google.com that surfaces articles and videos based on a user's interests and activity — not a response to a search query. There's nothing to 'rank' for in the traditional sense; eligibility is governed by Discover's content policies plus the same helpful-content, image, and page-experience signals Google Search already uses. URLs three ways — following links (primary), sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. (supplemental), and re-crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. pages it already knows. An orphan page (no internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.) with no sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. entry and no backlinks has no discovery path at all. Even when a sitemap gets it crawled, it receives zero internal PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems., which is why orphans cluster in “Discovered — currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” A crawl alone can’t find orphans — you compare the crawl against a superset of known URLs (sitemap, analytics, GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., backlink index). Fix the ones you want ranked with contextual internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.;
noindexthe intentional ones; redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. or delete the rest. In our 1M-domain study, 66.2% of sites had pages with only one dofollow internal link — orphans are the extreme end of a very common internal-link-poverty problem.
What makes a page an orphan
An orphan page is one that no other page on your site linksSitelinks are extra links from the same domain that Google clusters together under a single search result, usually for branded or navigational queries. They're generated entirely algorithmically — there's no way to add, edit, or guarantee them. to internally. That’s the whole definition — it’s about incoming internal links, nothing else. A page can have a pile of external backlinks and still be an orphan internally. A page can sit in your XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. and still be an orphan. Orphan status is measured purely against your site’s internal link graph.
It helps to think of internal linkingLinks between pages on the same site. as a spectrum rather than a binary. Truly orphaned (zero internal links) is the worst case; one internal link is barely better; a well-linked page sits at the healthy end. When we studied over a million domains for the Ahrefs Site Audit study, 66.2% of sites had pages with only a single dofollow incoming internal link. As I put it then: “At least it’s not a completely orphaned page. But if it’s a page that you want to rank, you may want to add some more internal links.” Orphans are just the extreme end of that same under-linking problem.
Why orphan pages are an SEO problem
Search engines can’t reliably discover them
Google is explicit about how it finds URLs. “There isn’t a central registry of all web pages, so Google must constantly look for new and updated pages.” It does that three ways: pages it already knows from prior crawls, “when Google extracts a link from a known page to a new page,” and “when you submit a list of pages (a sitemap) for Google to crawl.”
Link-following is the primary mechanism; sitemaps are supplemental, not a substitute. Google’s own link best-practices guidance is blunt: “Every page you care about should have a link from at least one other page on your site.” An orphan with no sitemap entry and no external backlinks has none of the three discovery paths — it is genuinely invisible to GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer..
John Mueller has called internal linking “super critical for SEO… one of the biggest things that you can do on a website to kind of guide Google and guide visitors to the pages that you think are important.” His framing is about link distance from your important hub pages, not just raw link count — an orphan has infinite link distance, because it isn’t reachable from your link graph at all.
They receive no internal PageRank
PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. still underpins Google’s link-based authority, and internal links are the plumbing that moves it around your site. An orphan page is disconnected from that plumbing, so it receives zero internal PageRank no matter how good its content is.
This is the angle most people miss, and it cuts both ways. A page with strong external backlinks but no internal links isn’t just under-powered — it’s a PageRank leak. Authority flows into that page from the outside and then has nowhere to go, because no internal links carry it onward to the rest of your site. Add internal links and you both strengthen the orphan and let its inbound equity circulate.
They cluster in “Discovered — currently not indexed”
When Google knows a URL exists (from a sitemap or an external link) but decides not to index it, GSC reports “Discovered — currently not indexed.” Orphan pages are a classic cause: Google found the URL, but with no internal links pointing at it, there’s no signal that the page matters, so it gets deprioritized. Adding internal links from relevant pages is one of the most reliable ways to nudge those URLs out of that bucket. (It’s not the only cause — thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count. and crawl budget play in too — but it’s one of the first things I check.)
They can waste crawl budget — but mostly on large sites
For the average site this barely registers. On very large sites (think 100k+ pages, or large numbers of rapidly changing URLs), big pools of orphaned and near-orphaned URLs consume crawl budgetThe number of URLs an engine will crawl in a timeframe. without earning their keep. Most sites genuinely don’t need to think about this; if you’re below that scale, treat orphans as a discovery and PageRank problem, not a crawl-budget one.
Common causes
- Migrations and redesigns — URLs survive the move but the links that pointed to them don’t get rebuilt.
- Navigation changes — a page gets pulled from the menu and nothing replaces the link.
- CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. workflow gaps — a post is published without being assigned to a category or tag, so the only link path (the category archive) never gets created. WordPress and similar CMSes generate a lot of paginationPagination splits a large set of content — product listings, blog archives, search results — across multiple sequentially numbered URLs. For SEO, each paginated page should be crawlable, indexable, and self-canonical; Google no longer uses rel=prev/next, but Bing still does. and archive URLs that can strand content this way.
- Discontinued productsA discontinued product is an item you'll never sell again — the manufacturer stopped making it, or you dropped the line. The SEO decision is end-of-life: 301-redirect the URL to a genuinely similar replacement or the closest relevant category if it earned links or traffic, 404/410 it if it didn't, or keep it live as a Discontinued tombstone page only when it still helps users. This is distinct from a temporary out-of-stock product, which you keep live at 200. — the product page is left live after it’s removed from listings.
- Campaign and PPC landing pages — built to be reached from an ad, never linked from the site (often intentional).
- Staging, test, and siloed microsite content — left publicly accessible and never wired into the main link graph.
How to find orphan pages
The core method is a two-list comparison, because no crawl can find orphans on its own:
- List 1 — reachable pages. Run a site crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. that starts at your homepage and follows internal links. This is everything Google could discover by link-following.
- List 2 — all known pages. Gather every URL you know exists: your XML sitemap, Google Analytics (pages with sessions), GSC pages, your backlink index, and server logs.
- Orphans = in List 2 but not List 1.
In Ahrefs Site Audit, this is built in — it seeds the crawl from your sitemaps, the Ahrefs backlink index, and connected GA/GSC data, then flags any URL it knows about but couldn’t reach via internal links as an “Orphan page.” In Screaming Frog, you upload your sitemap, GA, and GSC data as URL sources, run the crawl in list/crawl mode, then use the crawl analysis to surface URLs that appear in those sources but weren’t found by link-following. With GSC alone you can’t see site structure, but URLs sitting in “Discovered — currently not indexed” that don’t appear in your crawl are strong orphan suspects.
Supply the orphan candidates alongside reachable cluster pages. A similarity score can prioritize possible sources, but only an editor can confirm that a link belongs in the passage.
Create a review queue with my free Internal Link Cluster Visualizer Free
- First prove orphan status by comparing the crawl with sitemaps, analytics, GSC, or backlink data.
- Supply valuable orphan candidates and relevant reachable pages to generate possible connections.
- Add only contextually useful links, then recrawl to confirm each kept page is reachable.
How to fix them
Triage by value first — not every orphan deserves a link.
Before you triage, confirm the candidate itself is worth linking to: check its
response status, its canonical target, and whether it’s set to noindex.
Linking to a URL that redirects, 404s, canonicalizes elsewhere, or is
deliberately excluded from the index isn’t a fix — it just adds a broken or
wasted link. Only pages that resolve cleanly and are meant to be indexed belong
in the “add links” bucket below.
- Valuable pages you want to rank → add internal links. Find topically relevant pages and link from them with descriptive anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page.. Prioritize linking from pages that already have internal authority (hub and cluster pages), and make sure the page is in your XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. too.
- Intentional orphans (PPC, thank-you, some legal pages) →
noindex. These don’t need internal links; they need to be kept out of the index cleanly. Adding links would be the wrong fix. Give each category of intentional orphan a named owner and a documented reason it’s excluded, so the next audit doesn’t re-flag it as an accidental one. - Low-value pages → redirect or delete. 301 to the most relevant alternative, or remove it entirely if there’s nothing relevant and no backlinks worth preserving.
Prevent recurrence: bake an internal-linking step into your publishing workflow, assign categories/tags before publishing, and re-audit after every migration, redesign, or navigation change. This is the same internal-link discipline that governs crawl depthCrawl depth usually means click depth — how many clicks it takes to reach a page from the homepage by following internal links. It can also mean a crawler setting that limits how many levels deep a crawl goes before it stops. and how PageRank flows through your architecture — orphan prevention is just the floor of good internal linking.
A note on the related-but-different term: an orphan page has no incoming internal links; a dead-end page has no outgoing internal links. They’re separate issues, and a page can be both.
AI summary
A condensed take on the Advanced version:
- Definition: an orphan page has no incoming internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. It’s measured only against your internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. graph — a page can have external backlinks or be in your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and still be an orphan.
- Why it hurts: Google discoversGoogle Discover is a personalized, mobile-first content feed built into the Google app, Chrome's mobile New Tab page, and google.com that surfaces articles and videos based on a user's interests and activity — not a response to a search query. There's nothing to 'rank' for in the traditional sense; eligibility is governed by Discover's content policies plus the same helpful-content, image, and page-experience signals Google Search already uses. URLs three ways — following links (primary), sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. (supplemental), prior crawls. An orphan with no sitemap entry and no backlinks has no discovery path. Even when crawled via sitemap it gets zero internal PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems., so orphans cluster in “Discovered — currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..”
- PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. leak: an orphan with external backlinks pulls authority in that then can’t flow onward to the rest of the site.
- Scale of the problem: in Patrick’s 1M-domain Ahrefs Site Audit study, 66.2% of sites had pages with only one dofollow internal link — orphans are the extreme end of widespread internal-link poverty.
- Finding them: a crawl alone can’t find orphans. Compare the crawl (pages reachable by links) against a superset of known URLs (sitemap, analytics, GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., backlinks). Orphans = known but unreachable.
- Fixing them: valuable pages → add contextual internal links; intentional
orphans (PPC, thank-you) →
noindex; low-value → redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. or delete. - Crawl budgetThe number of URLs an engine will crawl in a timeframe. is a real but secondary concern, mostly for very large sites.
Official documentation
Primary-source guidance relevant to orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. and discovery.
- In-Depth Guide to How Google Search Works — the three URL-discovery mechanisms (prior crawls, link extraction, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. submission).
- Make your links crawlable — “Every page you care about should have a link from at least one other page on your site,” and how links are used to find and rank pages.
- SEO Starter Guide — how Google finds pages via known crawls and link-following.
- Page Indexing report — “Discovered — currently not indexed” — Google’s definition of the status orphan pages commonly land in.
- Sitemaps overview — what a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. does (and doesn’t) guarantee; a discovery aid, not a substitute for links.
Bing / Microsoft
- Bing Webmaster Tools — the Site Scan tool surfaces crawlability issuesCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index.; Bing has no documentation that addresses orphan pages by name, but the link-following discovery principle applies to BingbotBingbot is Microsoft Bing's primary web crawler — the bot that discovers, fetches, and renders pages to build the Bing index. That index also powers Yahoo, DuckDuckGo, Ecosia, and Microsoft Copilot, so Bingbot's reach is far wider than Bing's own search-market share. too.
Ahrefs (tool documentation)
- Orphan page error in Site Audit — how Ahrefs defines and seeds the orphan-page check.
Quotes from the source
On-the-record statements from Google. Each link is a deep link that jumps to the quoted passage on the source page.
Google — how URLs are discovered
- “There isn’t a central registry of all web pages, so Google must constantly look for new and updated pages.” — Google Search Central docs. Jump to quote
- “Some pages are known because Google has already visited them. Other pages are discovered when Google extracts a link from a known page to a new page… Still other pages are discovered when you submit a list of pages (a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.) for Google to crawl.” Jump to quote
Google — links are critical
- “Google uses links as a signal when determining the relevancy of pages, and to find new pages to crawl.” — Make your links crawlable. Jump to quote
- “Every page you care about should have a link from at least one other page on your site.” Jump to quote
John Mueller, Google Search Advocate
- “Internal linkingAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. is super critical for SEO. I think it’s one of the biggest things that you can do on a website to kind of guide Google and guide visitors to the pages that you think are important.” — Google Search Central SEO Office Hours, March 2022. Read the coverage
- On link distance over URL depth: “It’s really, from the homepage, or from the primary page, how quickly can we reach that specific page?” — an orphan has infinite link distance. Read the coverage
Orphan-page audit checklist
A repeatable pass to find and resolve orphans:
- Run a full internal-link crawl from the homepage (List 1 — reachable pages).
- Gather all known URLs: XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags., Google Analytics, GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. pages, backlink indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. (List 2).
- Compute the gap: URLs in List 2 but not in List 1 — these are your orphan candidates.
- In Ahrefs Site Audit, connect GA + GSC and check the Orphan page issue.
- In Screaming Frog, upload sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing./GA/GSC as URL sources and run crawl analysis to surface unlinked URLs.
- Cross-check “Discovered — currently not indexed” in GSC against your crawl — unreachable ones are strong orphan suspects.
- Rule out false positives first: auth-gated pages, JS-only links, alternate hosts/protocols, blocked paths, and URL-normalization variants can all look like a zero-inlink orphan without being one.
- Triage each orphan by value: organic trafficVisitors from unpaid search results — it compounds without ad spend.? backlinks? useful content?
- Before linking, confirm the candidate resolves cleanly (200 status, correct
canonical, not
noindex) — don’t link to a redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., duplicate, or excluded URL. - Want to rank it → add contextual internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. with descriptive anchor text from relevant, authoritative pages; confirm it’s in the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing..
- Intentional (PPC / thank-you / some legal) → apply
noindex, don’t add links; assign an owner for each intentional-orphan category. - Low value → 301-redirect to the best alternative, or delete.
- Add an internal-linking step to your publishing workflow and re-audit after every migration or redesign.
The mental models
1. The internal-link spectrum. Orphan (0 links) → single link → well-linked. Don’t treat it as a binary. Patrick’s data: 66.2% of sites have pages with only one dofollow internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. — most orphan work is really “under-linking” work one notch further along.
2. Three discovery paths — and orphans block them all. Google finds URLs via prior crawls, link-following (primary), and sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. (supplemental). An orphan with no sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. entry and no backlinks has zero open paths. Ask: which of the three is actually carrying this page?
3. Orphans = known minus reachable. A crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. can’t find orphans — it only follows links it can reach. Orphans are the set difference between every URL you know exists and every URL a crawl reached. You always need a second list.
4. PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems. is plumbing; orphans are disconnected pipes. Internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. circulate authority. An orphan receives none — and an orphan with external backlinks leaks the authority flowing into it, because nothing carries it onward.
5. The triage decision rule.
Want it to rank? Add internal links. Intentionally hidden (PPC, thank-you)?
noindex. Low value? RedirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't. or delete. Never default to “add links” for
every orphan.
Patrick's relevant free tools
- Internal Link Cluster Visualizer — Analyze a bounded supplied internal-link graph, orphans, PageRank, and lexical missing-link suggestions.
- XML Sitemap Validator — Paste, upload, or fetch a sitemap by URL — errors, warnings, and a health score with line numbers. Pasted and uploaded sitemaps are validated entirely in your browser.
- XML Sitemap Generator — Generate an XML sitemap from a capped, robots-respecting same-site crawl. Noindex, off-canonical, failed, and uncertain URLs remain visibly separate; lastmod dates are emitted only when the page provides evidence.
Tools for finding orphan pages
Remember the rule: a crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. alone can’t find orphans, so every tool below works by combining a crawl with an external list of known URLs.
- Ahrefs Site Audit — flags an Orphan page issue by seeding the crawl from your sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., the Ahrefs backlink indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and connected Google Analytics / Search Console data, then comparing that against what it could reach via internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them..
- Screaming Frog SEO Spider — upload your XML sitemapAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags., GA, and GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. data as URL sources, run crawl analysis, and export URLs that exist in those sources but weren’t foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. by link-following.
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. — no structure view, but the Page Indexing reportThe Google Search Console report (formerly Index Coverage) showing how many of your URLs are indexed vs. not indexed, and grouping the not-indexed ones by reason. (“Discovered — currently not indexed”) surfaces URLs Google knows about; cross- reference against your crawl to spot orphans.
- Google Analytics — pages receiving sessions that don’t appear in your internal-link crawl are orphan candidates.
- Server log analysis — URLs bots hit (or users land on) that aren’t reachable internally point to orphans, especially on large sites.
- Ahrefs Webmaster Tools — free Site Audit access for sites you verify, with the same orphan-page check.
Orphan-page fixes that create new problems
Treating the XML sitemap as an internal-link strategy
A sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. can help discovery, but it does not give a page navigational context or internal linkAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. equity. Keep the URL in the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and link it from a relevant hub, category, or article when the page is meant to rank.
Adding every orphan to the global navigation
Sitewide links are not the answer to every disconnected page. Place each valuable page where a reader would naturally need it; consolidate, redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., or remove pages that have no continuing purpose.
Calling every zero-inlink crawl result an orphan
A crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. sees only the sources and starting URLs you gave it. Compare crawl data with sitemaps, analytics, Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., and backlink data before deciding that the page exists and truly has no internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them.. A “zero inlinks” result can also be a crawl-configuration false positive rather than a true orphan — check for: links gated behind a login or other authentication, links only added by JavaScript the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. didn’t render, links pointing to a different host or protocol (http vs. httpsHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', www vs. non-wwwWWW vs. non-WWW is the choice between serving a site from www.example.com (the www subdomain) or example.com (the bare/apex/root domain). Both point to the same content, so it's a canonicalization and consistency decision, not a ranking factor.) than the one you crawled, mobile or template variants the crawl treated as separate URLs, paths blocked by robots.txtA Google Search Console Page Indexing status: the URL was excluded from indexing because your robots.txt disallows crawling it. Usually intentional and benign — robots.txt blocks crawling, not indexing. or crawler settings, and URL-normalization differences (trailing slashesA trailing slash is the forward slash (/) at the end of a URL — example.com/page/ versus example.com/page. Except at the bare root domain, the two versions are different URLs to search engines, so you pick one format and enforce it., tracking parameters, case) that make the same page look like two. Rule these out and confirm the canonical target before you treat the URL as genuinely orphaned.
Common orphan-page investigation issues
The crawler reports no orphans
Symptom: The crawl looks clean even though analytics contains URLs that users can visit. Likely cause: A link crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. cannot discover a URL with no links, and no secondary URL source was connected. Fix: Import sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., analytics, Search Console, backlink, and CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. URL lists, then compare them with crawled URLs.
A fixed page still shows zero inlinks
Symptom: You added a link, but the next report still classifies the target as an
orphan. Likely cause: The source page was not crawled, the link is injected only
after interaction, or the destination differs after normalization. Fix: Inspect
the source HTML for a real href, crawl the source directly, and compare the exact
final destination after redirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't..
An orphan is indexed but never improves
Symptom: The URL is indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. after being added to a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., yet impressions stay flat. Likely cause: Discovery was solved without adding useful internal context, or the page overlaps another URL. Fix: Add relevant contextual links, verify the canonical, and decide whether consolidation is better than preserving the page.
Classify an orphan-page export
Paste a CSV containing URL, status, canonical, organic sessions, impressions, backlinks, and page type into this prompt:
Act as a technical SEO reviewer. Classify each supplied URL as: link and retain,
consolidate, redirect, intentionally isolated, or investigate. Use only the columns
I provide. For every decision, name the evidence, the most relevant page or section
that should link to it when applicable, and any missing data that prevents a safe
decision. Do not infer traffic, backlinks, or index state. Return a review table and
a separate list of ambiguous cases. Propose contextual link placements
Paste the orphan’s title, short summary, and a list of candidate source pages:
Find the candidate source pages where a link to this target would genuinely help a
reader. For each fit, propose one natural sentence and descriptive anchor text.
Reject weak placements instead of forcing a link. Do not invent claims or add the
link to global navigation unless the target belongs there for users. Compare a crawl with a known-URL inventory
Save normalized absolute URLs from your crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. as crawled.txt and the combined
sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., analytics, Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., backlink, and CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. inventory as known.txt.
Run this on macOS or Linux:
comm -23 <(sort -u known.txt) <(sort -u crawled.txt)PowerShell equivalent:
Compare-Object (Get-Content crawled.txt | Sort-Object -Unique) (Get-Content known.txt | Sort-Object -Unique) -PassThru | Where-Object SideIndicator -eq '=>'The output is a candidate list, not a verdict. RedirectsA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., canonicals, tracking parameters, and intentionally isolated pages still need classification.
Prove that a retained page is no longer orphaned
Test to run: Crawl from the site’s normal entry point and inspect the target’s inlinks. Expected result: At least one relevant, crawlable internal anchor reaches the target’s final URL. Failure interpretation: The source was not reachable, the link is not a real anchor, or it points through a different URL. Monitoring window: Immediate after deployment and crawl completion. Rollback trigger: The new link disrupts navigation or points users to the wrong page; remove or correct it.
Prove that the indexability signals agree
Test to run: Check the final status, robots directives, and canonical with the
Google Index Checker, then use Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. URL
Inspection for Google’s recorded state. Expected result: A retained page returns
200, permits indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and canonicals to itself or the intended equivalent.
Failure interpretation: The page was linked before its technical signals were
fixed. Monitoring window: HTTP and HTML signals are immediate; Google’s recorded
state updates after recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial.. Rollback trigger: The link exposes a private,
duplicate, or intentionally excluded URL.
Orphan-page count by business disposition
Metric: Candidate orphans split into retain, consolidate, redirectA redirect sends browsers and crawlers from a requested URL to a different one. An HTTP redirect specifically is a 3xx status code paired with a Location header; meta refresh and JavaScript redirects achieve a similar navigation without being a 3xx response themselves. Permanent redirects (301/308) are Google's signal the target should be canonical; temporary ones (302/303/307) aren't., intentionally isolated, and investigate. What it tells you: Whether the backlog is a discovery problem or mostly URL hygiene. How to pull it: Join known-URL sources against a fresh link crawl and record the reviewed disposition. Benchmark / realistic range: Establish a site-specific baseline; the target for valuable, indexable pages marked retain is zero true orphans. Cadence: Monthly and after migrations, redesigns, or navigation changes.
Valuable pages restored to the link graph
Metric: Retained candidates that gain at least one relevant crawlable inlink. What it tells you: Whether fixes changed the graph rather than only a spreadsheet. How to pull it: Compare target inlink counts between two crawl snapshots. Benchmark / realistic range: Measure against the approved remediation list; intentional isolated pages do not belong in the denominator. Cadence: After each remediation release, then monthly until the backlog closes.
Test yourself: Orphan pages
Five quick questions on what orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. are, why they hurt, and how to fix them. Pick an answer for each, then check.
Resources worth your time
My related writing
- Ahrefs Site Audit Study: 1M Domains — the source of the 66.2% single-internal-link stat and a broad look at common technical issues.
- Internal Links for SEO: An Actionable Guide — how to build the internal-link structure that prevents orphans.
- How to Find and Fix Orphan Pages — the Ahrefs Blog walkthrough with tool-specific steps.
- The Beginner’s Guide to Technical SEO — where orphan pagesAn orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site. fit in the bigger crawl/indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. picture.
My speaking
- How Search Works (SlideShare) — my walkthrough of discovery, crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and ranking, including how link-following drives discovery. (Standing disclaimer: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From around the industry
- Orphan page error in Site Audit (Ahrefs Help) — the precise definition Site Audit uses and how it seeds the crawl.
- How To Find Orphan Pages (Screaming Frog) — the three URL-source methods for surfacing orphans in a crawl.
- What Are Orphan Pages & How to Find and Fix Them (Semrush) — a detailed causes list with screenshots.
- How to Find Orphan Pages (Search Engine Journal) — the spreadsheet-based comparison method.
- What Are Orphan Pages & How to Find and Fix Them (Conductor) — a clear triage/decision-flow framing.
- Orphan Page (Ahrefs SEO Glossary) — the short reference definition.
Orphan Pages
An orphan page is a page on your site that no other page links to internally. Because crawlers discover pages by following links, an orphan page is effectively invisible to search engines unless it's in an XML sitemap or linked from an external site.
Related: Internal Link, Crawl depth, Discovered – currently not indexed
Orphan Pages
An orphan page is a page on your website that has no incoming internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. from any other page on the site. Search engines discover pages primarily by following links, so a page nothing links to has no reliable discovery path — it’s effectively cut off from the rest of your site.
Google discoversGoogle Discover is a personalized, mobile-first content feed built into the Google app, Chrome's mobile New Tab page, and google.com that surfaces articles and videos based on a user's interests and activity — not a response to a search query. There's nothing to 'rank' for in the traditional sense; eligibility is governed by Discover's content policies plus the same helpful-content, image, and page-experience signals Google Search already uses. URLs three ways: by following links from pages it already knows (the primary mechanism), by reading XML sitemapsAn XML sitemap is a UTF-8 file listing the canonical URLs on your site (with optional lastmod) so search engines can discover and prioritize them. It's a discovery and diagnostic aid, not a guarantee of indexing — and Google ignores its priority and changefreq tags. you submit, and by re-crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. pages it has seen before. An orphan page can still be foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. if it’s in a sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. or has external backlinks, but without internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. it receives zero internal PageRankPageRank is Google's original recursive link-graph algorithm: a page's score depends on the scores of the pages linking to it, and in the published model each page's score is split across its outbound links (the simplified version: links are weighted votes). Google says it's evolved since launch but still part of its core ranking systems., Google has little ongoing reason to re-crawl it, and it frequently lands in the “Discovered — currently not indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.” bucket in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results..
Not every orphan page is a mistake. PPC landing pages, post-purchase thank-you pages, and staging URLs are often deliberately kept out of the internal link graph — the right fix for those is noindex, not adding links. The pages that matter are the ones you want to rank that nothing links to. To find them you compare two lists: every page a crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. reaches by following internal links, and every page you know exists (from your sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing., analytics, GSC, and backlink data). Anything in the second list but not the first is an orphan.
Related: Internal Link, Crawl depth, Discovered – currently not indexed
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Added a pre-link eligibility check and a specific false-positive checklist for orphan candidates, plus an ownership requirement for intentional orphans.
Change details
-
Advanced 'How to fix them' now requires confirming a candidate's response status, canonical, and noindex state before adding a link, and requires a named owner for each intentional-orphan category.
-
Anti-patterns lens now lists specific crawl-configuration false-positive causes (auth-gated pages, JS-only links, alternate hosts, blocked paths, URL-normalization variants) instead of a general warning.
Full comparison unavailable — no prior snapshot was archived for this revision.