Technical SEO
Orphan Pages SEO: What Happens When Neither Google Nor ChatGPT Can Reach Your Content
Most advice about orphan pages was written for a search engine that had effectively unlimited patience. Googlebot is enormous. If a page sat outside your internal link structure but appeared in your sitemap, there was a decent chance Google would eventually wander over, crawl it, and index it anyway. Untidy, but survivable.
That assumption no longer covers the whole picture. There is a second retrieval layer now — the crawlers behind ChatGPT, Perplexity, Claude and Google’s own AI features — and it operates at a fraction of Googlebot’s scale. A crawler with a small budget does not wander. It follows the paths that are obviously there.
Which turns an orphan page from a housekeeping problem into a visibility problem. This piece covers what an orphan page is, what the current crawl data says about why it matters more than it did five years ago, how to find yours, and — the part almost nobody writes about — how to work out which ones are actually worth your time.
Key takeaways
- An orphan page is any page with no in-content internal links pointing at it, regardless of whether it appears in your sitemap.
- OpenAI’s crawlers logged 887 million events in a month against Googlebot’s 18.2 billion, per Botify’s crawler volume comparison — roughly 4% of Google’s volume.
- A smaller crawl footprint makes in-content links more decisive for discovery, not less.
- Across the 15 SaaS blogs I crawled, 38% of posts had no in-content link pointing at them — orphaning is the default state, not an edge case.
- Not every orphan needs fixing. Triage on impressions, conversion role and content quality before you touch anything.
What is an orphan page?
An orphan page is a page on your site that no other page links to in its body content. A visitor cannot reach it by clicking through your navigation, your category pages, or any article. It exists, it usually returns a 200 status code, and structurally it sits outside your site.
That matters because of how discovery works. Google’s own documentation on how Search works describes URL discovery in three routes: pages Google has already visited, pages found by extracting a link from a known page — a hub or category page pointing at a new blog post — and pages submitted in a sitemap. An orphan page has cut off the second route entirely, which is the one that carries context about what the page is and how it relates to everything else you publish.
Orphan page vs. dead-end page vs. noindexed page
Three states get conflated, and they need different responses.
An orphan page has no inbound internal links. Nothing points at it.
A dead-end page has the opposite problem: things link to it, but it links to nothing. Crawlers arrive and stop. That is a link-equity and user-journey issue, not a discovery one.
A noindexed page is deliberately excluded from the index. It may be perfectly well linked. Its exclusion is a decision, not an accident.
One more distinction, because it trips people up constantly: a page in your sitemap with zero in-content links is still an orphan. Google’s sitemap documentation describes a sitemap as telling search engines which pages and files you think are important on your site. That is a declaration of intent, not a structural signal. It does not tell a crawler where the page belongs or what it relates to. Those signals come from links.
This is where most orphan pages SEO advice quietly goes wrong: it treats sitemap coverage as equivalent to internal linking, and the two do different jobs.
The orphan page problem changed when AI started crawling
Here is what shifted. AI crawlers went from a rounding error to a real retrieval layer in about eighteen months, and they are still small enough that reachability genuinely constrains them.
Botify’s 2026 analysis of 7 billion log files, covering November 2024 through March 2026, found OpenAI’s crawl of the web roughly tripled after the GPT-5 launch — OAI-SearchBot events increased 3.5x since August 2025, and GPTBot 2.9x. That is steep growth, and it is the number most people quote.
The more useful number is the one underneath it. In the month analysed, there were 18.2 billion Googlebot events against 887 million OpenAI events. OpenAI’s crawlers represent roughly 4% of Google’s total crawl volume, up from 1.38% a year earlier. Against Bing, the same Botify dataset puts OpenAI at about 14% of the crawl.
A smaller crawl is a pickier crawl
Tripling still leaves you at 4%. And a crawler operating at 4% of Google’s volume cannot afford the exploratory, low-yield fetching that Googlebot does routinely. It goes where the links go.
This is the practical consequence: a page that Googlebot might eventually reach through sitemap discovery, sheer volume and repeated recrawls is a page an AI crawler will plausibly never see. The tolerance that made orphaning survivable in classic search is exactly the thing the smaller crawlers do not have. Sitebulb’s orphan page guide puts it plainly — orphans are invisible to AI search too, and it names this as a newer problem than the guidance most people are working from.
Retrieval favours reinforced pages
I want to be precise about where the evidence stops, because this is an area where confident overclaiming is rife.
What the log data supports is a reachability argument: fewer crawl events mean a crawler covers less of your site, and in-content links are the mechanism that decides what gets covered. That is solid.
What nobody outside these companies can honestly assert is that internal links are a confirmed input to any AI model’s ranking or citation selection. No such disclosure exists. Anyone telling you otherwise is guessing with conviction.
So the honest version is this: internal links determine whether an AI crawler reaches your page at all. Whether they influence what happens after that is unknown. The first half is enough to act on — you cannot be cited from a page that was never fetched.
What orphan pages cost you in classic search
The traditional consequences still apply. They are just less uniformly severe than the standard advice implies.
Discovery and indexation
An unlinked page depends entirely on secondary discovery routes: your sitemap, an external backlink, or a manual submission. Each of those works, and none carries the contextual signal an in-content link does. If you are seeing pages stuck in the status you see when Google fetched the page and skipped it, weak internal linking is one of the first things worth ruling out.
No internal link equity reaches the page
Internal links distribute authority around your site. A page with no inbound internal links receives none of it, and competes on external signals alone. For a blog post with no backlinks — which is most blog posts — that means competing on essentially nothing.
The crawl-budget claim, qualified
Nearly every article on this topic asserts that orphan pages waste crawl budget. It is repeated so consistently that it reads as settled. It mostly is not, and Google says so directly.
Google’s crawl budget guide opens by telling readers that if their site does not have a large number of rapidly changing pages, or if pages are crawled the same day they are published, they do not need to read it. The guide is aimed at sites with 1 million+ unique pages changing weekly, or 10,000+ pages changing daily.
If you run a 200-page marketing site, crawl budget is not your problem and citing it will damage your credibility with anyone technical. If you run a large ecommerce catalogue with faceted navigation, it very much is. Google’s definition of crawl budget combines a crawl capacity limit with crawl demand, and names “perceived inventory” — the volume of duplicate or unwanted URLs Google knows about — as the factor site owners control most. Note that this is an argument about URL bloat, not specifically about orphans. If you want the mechanics in full, I have written separately about what crawl budget actually rations.
Use the crawl-budget argument when it applies. Drop it when it does not — you have better ones.
How common is this, actually?
Worth establishing the scale, because “you might have a few orphan pages” and “roughly a third of your blog is stranded” call for very different levels of urgency.
I crawled 10,426 posts across 15 SaaS blogs and mapped 44,513 in-content links between them. 3,979 posts — 38% — had zero in-content inbound links. The full methodology and per-site results are in my crawl of 15 SaaS blogs, including the spread, which is wide enough to be the real finding: this is a choice sites make, not a rate of natural decay.
The pattern is not confined to blogs. Botify’s log file analysis documents one site where orphan pages made up more than 70% of the pages Google explored, while producing just 5% of organic visits — the remaining 95% came from pages inside the site structure, which accounted for only 30% of total pages. That is one client site rather than a population figure, and the analysis dates from 2020, so treat it as an illustration of how lopsided the ratio can get rather than a benchmark.
Take both together and the picture is consistent: orphaning is the default outcome of publishing at volume without a linking process, and the stranded pages contribute far less than their share of the URL count.
Why orphan pages happen
Nobody sets out to orphan content. Six patterns produce almost all of it.
Site migrations. URLs change, redirects get mapped, but the in-content links inside old articles still point at the previous structure — or get stripped entirely during the content transfer.
CMS templates that never link out. Posts publish into a reverse-chronological feed. They appear on page one of the blog index, slide to page four, and once they are past the pagination depth anything crawls to, nothing links to them at all.
Campaign and landing pages. Built for a paid channel, deliberately kept out of the navigation, then left live long after the campaign ends.
Deprecated products and services. The page for a retired offering gets pulled from the menu. The URL stays up because someone might still need it.
Paginated and filtered archives. Tag pages, author archives and filter combinations generate URLs that no editorial content ever references.
Publishing outside a hub structure. The most common cause by a distance. Posts get written to a keyword list rather than a topic map, so there is no pillar page that would naturally link to them and no editorial habit of linking new posts to old ones.
Recognising which of these applies matters, because it tells you whether you have a one-off cleanup or a process problem that will regenerate orphans every month.
How to find orphan pages
Every tool that finds orphan pages runs the same operation. Learn the operation and the tool choice stops mattering much.
The method underneath every tool: two lists and a subtraction
Build list A: every URL you know exists. Sources are your XML sitemap, Search Console’s page indexing report, Google Analytics landing pages, your CMS export, and server logs if you have them.
Build list B: every URL a crawler can reach by following in-content links from your homepage.
Subtract B from A. What remains is your orphan set. Everything below is an implementation of that subtraction.
Screaming Frog
The most common route, with one step people skip. Enable “Crawl Linked XML Sitemaps” under Configuration → Spider → Crawl, connect the Google Analytics and Search Console APIs, and enable crawling of new URLs discovered through both.
Then run the crawl — and afterwards run Crawl Analysis. This is the step that catches people out. Screaming Frog’s own tutorial is explicit that the Orphan URLs filters only populate after Crawl Analysis has completed; before that they sit empty with a “Crawl Analysis Required” note, and it is easy to conclude you have no orphans when you have simply not run the analysis. Export via Reports → Orphan Pages.
Semrush, Ahrefs and Sitebulb
Semrush Site Audit reports orphaned pages once you connect Google Analytics — without that connection it only compares the sitemap against the crawl, which misses orphans absent from both. Ahrefs surfaces them in Site Audit through its own crawl-versus-known-URL comparison. Sitebulb runs the same subtraction and is the most explicit of the three about which orphans are genuine issues.
All three share a limitation worth knowing: they can only find orphans among URLs they know about. If a page is in neither your sitemap, nor GA, nor Search Console, nor any crawl path, no tool will surface it. Server logs are the only complete source.
The free route
Export your sitemap. Map your in-content internal links. Compare the two.
You can do this with a spreadsheet and patience, or you can use the Internal Link Graph tool I built for exactly this — enter a domain and it maps in-content links as an interactive graph, showing orphans and de-facto hubs. No licence required, and it makes the shape of the problem visible in a way a CSV does not.
Checking whether AI crawlers reach the page
This is the check almost nobody runs, and it is the one that connects directly to the argument above.
Pull your server logs and filter for the AI crawler user agents. OpenAI’s crawler documentation lists three that behave differently and are worth separating: OAI-SearchBot for search results, GPTBot for training, and ChatGPT-User for fetches triggered by a user’s question in a conversation. Add PerplexityBot, which Perplexity’s own bot documentation describes as the agent that surfaces and links sites in its results, plus ClaudeBot. Cross-reference the URLs they fetched against your orphan list. If a page appears in Googlebot’s log lines but in none of the AI crawlers’, you have direct evidence of the reachability gap rather than an inference from it — and a much easier conversation with whoever needs to approve the fix. If you have not worked with logs before, I have covered how to read server logs for crawler behaviour separately.
Which orphans to fix first (and which to leave alone)
Most guides end at “add an internal link.” That is fine advice for a site with nine orphans and useless for one with four hundred. Fixing them all is not a plan, and some of them should not be fixed at all.
Intentional orphans you should leave alone
Some pages are orphans by design. Thank-you and confirmation pages should not be reachable from your navigation. Paid landing pages are often deliberately excluded so they do not compete with organic equivalents. Gated download confirmations, checkout steps and A/B test variants all belong in this category.
Sitebulb’s triage guidance treats these as a distinct class rather than defects, alongside duplicate URLs, which tend to be a canonicalisation question rather than a linking one. If a page is an intentional orphan, the fix is not a link — it is making the intent explicit with a noindex tag so the page stops appearing in reports as an unresolved issue.
The triage table
Score each remaining orphan on four inputs, then assign a disposition.
| Signal | What to check |
|---|---|
| Demand | Impressions and clicks in Search Console over the last 90 days |
| Commercial role | Does it sit on a conversion path, or support one? |
| Quality | Would you publish it today as it stands? |
| Duplication | Does a better page on your site already cover this? |
| Disposition | When it applies | Action |
|---|---|---|
| Adopt | Has impressions, or is decent content on a topic you care about | Add in-content links from relevant pages |
| Redirect | Superseded by a better page covering the same topic | 301 to that page |
| Noindex | Intentional orphan with a functional purpose | Add noindex, leave unlinked |
| Delete | No demand, no commercial role, poor quality, no unique content | 410, or 404 |
Conductor’s orphan page guide frames the same decision as a two-step question — does the page belong in your site structure, and if not, does it carry value worth preserving through a 301 to the closest relevant alternative. That is a good sanity check on any disposition you assign.
Work the adopt list in descending order of impressions. Pages already earning impressions while orphaned are the ones with the most upside — they are ranking on content quality alone, with no internal support at all. The orphan check is one step in a wider process; for the full sequence see where the orphan check sits in a full internal linking audit.
How to fix an orphan page
Adopt it: where the link has to come from
The link must be in the body content of a topically related page. A footer link, a sidebar widget, or a bulk “related posts” module carries far less weight than a contextual link inside a paragraph, and site-wide modules are routinely discounted.
Aim for two to three in-content links from genuinely relevant pages. Your best sources are pages that already rank and get crawled frequently — they pass more authority and get recrawled sooner, so the new link is discovered faster.
Write anchor text that describes the destination. Not “click here”, not the bare URL, and not an exact-match keyword repeated identically across every link.
The mistake that makes this fail: adding the page to your sitemap and calling it done. Google’s documentation on URL discovery does list sitemap submission as a genuine discovery route, so this is not useless — but it is a separate and weaker route from link extraction, and it does nothing about link equity, contextual relevance, or reachability for crawlers operating on a tight budget. The sitemap announces the URL. The link explains it.
Redirect, noindex or delete
301 redirect when another page covers the same topic better. Point it at the closest relevant page, not the homepage — a homepage redirect is treated as a soft 404 and wastes whatever equity the old URL held.
Noindex when the page needs to exist and be reachable by URL but has no business in search results. Leave it unlinked; that is now deliberate.
410 Gone when the content is genuinely finished. It is a cleaner signal than 404 and gets processed faster. Do check for inbound backlinks first — if the page has any, redirect instead.
Confirming the fix
Run URL Inspection in Search Console on the fixed URL and use “Test live URL” to confirm Google can fetch it. Then re-crawl your site and verify the page no longer appears in the orphan report.
After that, watch three things over four to eight weeks: impressions for that specific URL in Search Console, its position for the terms it should serve, and — if you have log access — whether AI crawler hits start appearing against it.
Set a realistic expectation. Discovery can happen within days. Meaningful ranking movement takes weeks, and on a low-authority site it may take longer or not arrive at all if the underlying content is weak. An internal link makes a page reachable and gives it authority. It does not make a thin page good.
Frequently asked questions
Are orphan pages bad for SEO?
Usually, but not universally. An orphan page receives no internal link equity and depends on secondary discovery routes, which typically means poor rankings. Some orphans are intentional and correct — thank-you pages, paid landing pages, checkout steps. The problem is the unintentional ones.
Can Google find orphan pages if they are in the sitemap?
Often, yes. Google lists sitemap submission as a genuine discovery route. But the sitemap only announces that a URL exists; it carries no signal about relevance, importance, or how the page relates to the rest of your site. Sitemap inclusion is a weaker substitute for an in-content link, not an equivalent one.
Do orphan pages hurt AI search visibility?
They restrict it at the discovery stage. OpenAI’s crawlers run at roughly 4% of Googlebot’s volume, so they cover far less of any given site and rely more heavily on clear link paths. Whether internal links influence citation selection after retrieval is undisclosed by every provider — but a page that is never fetched cannot be cited.
What is an example of an orphan page?
A blog post published in 2022 that appeared on the blog index, slid past the pagination depth, and was never linked from any other article. It is in your sitemap, returns a 200, and no page on your site points at it.
Can I find orphan pages without a paid tool?
Yes. Export your sitemap, list your in-content internal links, and subtract the second from the first. A spreadsheet handles this on a small site; a link-graph mapper is faster on a large one. Server logs give the most complete picture of what crawlers actually reach.
How often should I check for orphan pages?
Quarterly for most sites. Monthly if you publish more than about twenty pages a month, and immediately after any migration or CMS change — migrations are one of the largest single generators of orphans.