How do you identify orphan pages in an internal link audit?

Identify orphan pages by crawling your website to find URLs that have no internal links pointing to them, then compare the crawl results with your XML sitemap, CMS URL list and search data to uncover pages that are accessible but disconnected from the site structure. Review each URL for its SEO and business value before adding relevant contextual links, redirecting it or removing it.

Orphan pages are URLs with no internal links pointing to them from the crawlable pages on your website. Identify them by comparing a site crawl with other URL sources, such as the XML sitemap, CMS or database export, Google Search Console data and server logs. The URLs that appear in these sources but are not discovered through internal links are potential orphan pages and should then be assessed for value, relevance and technical status.

1. Crawl the website as a search engine would

Start with a crawl of the live site, following internal HTML links and recording the URLs discovered. Configure the crawl to include the relevant URL variations and to render JavaScript where navigation or content links depend on client-side code. The crawl should normally capture each page’s status code, indexability, canonical URL, page title, depth and number of internal inlinks.

Export the URLs that have no internal inlinks. These are the initial orphan-page candidates. Treat the result as a working list rather than a final diagnosis: a page may appear to be orphaned because the crawler could not access a blocked section, did not render a JavaScript menu or was prevented from following a link by crawl settings.

2. Compare the crawl with every reliable URL source

A crawl only finds pages that are connected to the starting points and links it can access. To find disconnected URLs, compare its results with separate records of pages that exist or have been requested.

  • XML sitemaps: Extract the URLs listed in each sitemap, including image, video or news sitemap files where relevant. A URL that is in a sitemap but absent from the crawl may be disconnected from the site structure.
  • CMS or database exports: Export published pages, products, categories, resources and other content types. This can reveal content that is live but excluded from the sitemap or from normal navigation.
  • Search performance data: Review pages that have received impressions or clicks. A URL can attract search traffic through external links or historical discovery even when no current internal link points to it.
  • Server logs: Where available, analyse requests from search-engine crawlers and users. Log data can identify URLs that are being visited but are not discoverable through the current internal linking structure.
  • Analytics data: Check landing pages and referral paths for URLs receiving visits. Analytics data is useful supporting evidence, although tracking gaps and consent settings mean it should not be treated as a complete URL inventory.
  • Redirect and migration records: Include old URLs, recently migrated content and URLs referenced in redirect rules. These may expose duplicate or obsolete pages that need to be resolved rather than linked.

Normalise all URL lists before comparing them. Account for protocol, host name, trailing slashes, upper- and lower-case characters, URL fragments, tracking parameters, duplicate query strings and redirects. Resolve each URL to its final canonical form where appropriate, while retaining the original URL for investigation.

3. Confirm whether each candidate is genuinely orphaned

For each URL found outside the crawl, check whether a relevant internal link exists but was missed. Inspect navigation, breadcrumbs, HTML sitemaps, pagination, related-content modules, faceted navigation and links rendered after JavaScript execution. Also check whether the link is blocked by robots.txt, requires a form submission, is placed in an excluded template or points to a different URL variant.

Use a clear definition for the audit. A page may have an internal link from a login-only area, a noindex page or a low-value utility page, but that does not necessarily provide useful discoverability. For SEO purposes, record whether the page has an inlink from a crawlable, indexable and contextually relevant page, not merely whether any reference exists somewhere in the source code.

Review the page’s status and indexability before making changes. A URL returning a successful response, containing useful content and intended for organic search is a stronger orphan-page concern than a discontinued page returning a not-found response. Canonicalised pages, noindex pages, parameter variations, search-result pages and private content may be intentionally excluded and should be classified accordingly.

4. Classify orphan pages by purpose and value

Do not add links to every orphan page automatically. Classify each URL using its content quality, search visibility, business purpose, conversion role, freshness and relationship to other pages.

  • Keep and link: The page is useful, accurate, indexable and relevant to an existing topic or user journey. Add links from suitable pages.
  • Improve and link: The page has potential but needs updated content, a clearer purpose, stronger targeting or a better relationship with related pages before it is promoted internally.
  • Consolidate: The page overlaps another URL or is a weak version of a stronger resource. Merge the useful content and redirect the redundant URL where appropriate.
  • Redirect: The URL has been replaced by a closely related page and should pass users and signals to the most relevant destination. Avoid redirecting unrelated URLs to a generic page.
  • Remove: The content is obsolete, duplicated without purpose, empty or not required. Choose the appropriate removal response and update references that still point to it.
  • Leave intentionally unlinked: Some pages, such as transactional confirmation pages, internal utility pages or restricted content, should not be made part of the public information architecture.

5. Choose appropriate internal links

For pages that should remain accessible and indexable, add links from relevant, crawlable pages rather than placing every URL in a broad sitewide module. Good source pages usually include the parent category, a closely related service or product page, a supporting guide, a glossary entry, a relevant case study or a high-quality hub page.

Use descriptive anchor text that accurately indicates the destination. Link where the reader is likely to need the additional information, and make the relationship between the source and destination clear. A page should generally have more than one sensible discovery route when the site structure supports it, but excessive cross-linking can dilute context and make pages harder to use.

Consider the position and prominence of the link. A contextual link in the main content or a clear category structure is usually more informative than an isolated footer link. However, do not force links into copy where they do not help the reader. Internal linking should reflect the information architecture and user journey, not just the outcome of an audit.

6. Validate the changes

After links, redirects or removals have been implemented, recrawl the affected sections and repeat the URL comparison. Confirm that important pages now have internal inlinks, the links return successful responses, canonical signals are consistent and no new orphan pages have been introduced by template or migration changes.

Check that the destination pages can be reached without relying on a sitemap alone. Review the updated sitemap separately: it should contain the preferred, indexable URLs that the site intends search engines to discover, not act as a substitute for internal navigation. Monitor crawl data, search performance and landing-page behaviour over time, particularly after major content or platform changes.

Common causes of orphan pages

  • Publishing a page without adding it to a category, hub or related-content module.
  • Removing a navigation item while leaving the page live.
  • Changing a URL without updating internal references.
  • Launching campaign or landing pages outside the main site template.
  • Filtering, pagination or JavaScript navigation that a crawler cannot access.
  • Deleting parent pages while retaining their child pages.
  • Moving content during a redesign or CMS migration.
  • Using inconsistent URL formats, redirects or canonical tags.

The most reliable audit combines crawl data with site records and behavioural evidence. The final output should distinguish genuinely valuable pages that need stronger internal discovery from intentional, obsolete or technically invalid URLs. This prevents orphan-page remediation from creating unnecessary links, duplicate content or poor user journeys.

An orphan page is a live URL with no discoverable internal link from the rest of the website. Because a standard crawl follows links, it may not find these pages at all, so identifying them requires comparing crawl data with independent URL sources.

Start by comparing the crawl with the XML sitemap, CMS export, search performance data and, where available, server logs. URLs that appear in these sources but not in the crawl are potential orphan pages. Check each candidate for URL variations, redirects, canonical tags, JavaScript navigation and blocked sections before confirming that it is genuinely disconnected.

Once confirmed, assess the page’s purpose and value. Add relevant contextual links to useful, indexable content; consolidate or redirect overlapping pages; and remove obsolete or private URLs rather than linking to them. Recrawl the site afterwards to verify that important pages are reachable through the internal structure.

Start identifying orphan pages

Start identifying orphan pages by comparing your latest site crawl with your XML sitemap, CMS URL list and search data. Review each candidate before adding relevant internal links, consolidating it or removing it.