Technical SEO Mastery · Site architecture, internal linking and XML sitemaps · lesson 9 of 18 · 14 min
Site architecture and internal linking
Architecture is how you tell search engines what matters
Internal links do three jobs at once: they let crawlers discover URLs, they pass signals between pages (Google still uses links, including internal ones, to understand importance and relationships), and they give context through anchor text. A good architecture puts your most valuable pages close to the home page, groups related content, and links it all with descriptive anchors.
Principles
- Shallow click depth for priority pages. Important categories and products should be reachable in a few clicks from the home page. Depth is not a ranking factor in itself, but deep pages tend to be crawled less often and receive less internal link value.
- Topical grouping (hubs and spokes). A hub page covers a topic broadly and links to detailed spoke pages; spokes link back to the hub and to closely related siblings.
- Descriptive anchor text. "Technical SEO audit checklist" beats "click here". Avoid over-engineered exact-match anchors on every link — write for users.
- Every indexable page gets at least one crawlable internal link from an indexable page. No orphans.
- Consistent URLs. Link to the canonical, final URL — not to redirects or parameter variants.
- Breadcrumbs on hierarchical sites, marked up with
BreadcrumbListstructured data.
Typical structures
Home
├── /services/ (hub)
│ ├── /services/technical-seo/ (spoke)
│ ├── /services/local-seo/
│ └── /services/link-building/
├── /guides/ (hub)
│ ├── /guides/core-web-vitals/
│ └── /guides/hreflang/
└── /case-studies/
For e-commerce: Home → Department → Category → Subcategory → Product, supported by cross-links (related products, "shop by brand", recently viewed rendered as real links).
Where internal link value leaks or pools
| Pattern | Effect | Fix | |---|---|---| | Mega-menus linking to hundreds of pages | Dilutes emphasis; can bloat HTML | Keep menus to real priority pages | | Links only in JS widgets without href | Not discovered | Render real anchors | | Sitewide footer links to low-value pages | Wasted emphasis | Trim footers | | Important pages only reachable via internal search | Effectively orphaned | Add category and contextual links | | Blog posts never linked after falling off page 1 of the blog index | Deep, rarely crawled | Topic hubs, related-post links, updated evergreen hubs | | nofollow on internal links | Wastes the link's signal; does not "save" value | Remove internal nofollow |
Measuring your architecture
Use a crawler to export for each indexable URL:
- Crawl depth (clicks from home).
- Unique inlinks and the linking pages.
- Internal link score (tools such as Screaming Frog's Link Score or Sitebulb's equivalents approximate relative internal importance).
- Anchor texts used to link to it.
Then join with Search Console clicks and impressions and business value. The classic finding: high-value pages with low inlinks and high depth, and low-value pages (tags, archives, filters) soaking up most internal links.
Orphan page discovery
Orphans are pages that exist (in sitemaps, logs or analytics) but have no internal links. To find them:
- Crawl from the home page.
- Separately collect URLs from XML sitemaps, Search Console (Performance and Page indexing exports), analytics landing pages and server logs.
- URLs present in step 2 but not in step 1 are orphans. Decide for each: link it, redirect it, or remove it.
Worked example 1: an agency blog
An agency has 400 blog posts. The top converting guide, on Google Business Profile optimisation, sits at depth 7, linked only from paginated blog archives. Fix:
- Create a Local SEO hub at
/guides/local-seo/linked from the main navigation. - Link the hub to the 25 most relevant local posts with descriptive anchors; each post links back to the hub.
- Add contextual links from the three service pages that mention Google Business Profile.
- Retire or merge weak, overlapping posts (with 301s to the best version).
After recrawling, the guide sits at depth 2 with dozens of relevant internal links. Monitor impressions and crawl frequency in logs; do not promise a ranking outcome.
Internal linking audit checklist
- Home page links to top categories/hubs.
- No internal links to 3xx, 4xx or non-canonical URLs.
- No important page deeper than a sensible depth for the site's size.
- Every indexable page has inlinks from indexable pages.
- Anchor text is descriptive and varied.
- Breadcrumbs present and marked up on hierarchical templates.
- Pagination and faceted links behave as designed.
Hands-on: internal PageRank from a crawl export (networkx)
Crawler "link score" metrics are useful; computing your own internal PageRank lets you combine it with business data. Export all internal links (source, destination) for indexable HTML pages:
# pip install pandas networkx
import pandas as pd, networkx as nx
links = pd.read_csv("all_inlinks.csv", usecols=["Source", "Destination", "Type", "Follow"])
links = links[(links.Type == "Hyperlink") & (links.Follow == True)]
G = nx.from_pandas_edgelist(links, "Source", "Destination", create_using=nx.DiGraph)
pr = pd.Series(nx.pagerank(G, alpha=0.85), name="internal_pr")
pages = pd.read_csv("internal_html.csv", usecols=["Address", "Crawl Depth", "Unique Inlinks"])
gsc = pd.read_csv("gsc_pages.csv") # columns: page, clicks, impressions
df = (pages.merge(pr, left_on="Address", right_index=True, how="left")
.merge(gsc, left_on="Address", right_on="page", how="left"))
df["pr_rank"] = df.internal_pr.rank(ascending=False)
df["imp_rank"] = df.impressions.rank(ascending=False)
# High-impression pages with weak internal support = linking opportunities
print(df[(df.imp_rank <= 100) & (df.pr_rank > 500)]
.sort_values("impressions", ascending=False)
[["Address", "Crawl Depth", "Unique Inlinks", "internal_pr", "impressions"]].head(20))
Treat the output as a prioritised list for contextual links from relevant, strong pages — not as a reason to add sitewide footer links.
Worked example 2: a Jeddah e-commerce "new arrivals" trap
A Jeddah fashion retailer (illustrative) links every product from a "New arrivals" carousel on the home page — until each item rotates out after two weeks and loses almost all internal links. Evergreen best-sellers end up at depth 6+. Fixes: permanent category and "best-sellers" modules with real links, related-product links by attribute (fabric, occasion), breadcrumbs with BreadcrumbList markup, and a quarterly check of depth and internal PageRank for the top 200 revenue products.
Common mistakes
- Adding internal nofollow to "sculpt PageRank" — the value is not redistributed.
- Relying on XML sitemaps instead of links for discovery of important pages.
- Automated "related links" that link randomly rather than by topic.
Video lecture: Site architecture and internal linking
Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.
- Site architecture and internal linking
- Three jobs of internal links
- Six principles
- Leaks and pools
- Measure it
- Finding orphans
- Hands-on: internal PageRank
- Example 1: agency blog (illustrative)
- Example 2: Jeddah 'new arrivals' trap (illustrative)
- Typical structures
- The internal linking checklist
- Watch me do it: linking opportunities
- Mistakes and measures
- Recap and try this now
Lecture transcript
Site architecture and internal linking
Internal links are the one ranking-related signal you control completely. No outreach, no budget, just decisions about which pages link to which. And yet most sites leave their internal linking to chance: a mega-menu here, a random related posts widget there. In this lecture you'll learn how architecture tells search engines what matters, the patterns that leak or pool link value, how to find orphans, and how to compute internal PageRank from a crawl to find your best linking opportunities.
Three jobs of internal links
Internal links do three jobs at once. They let crawlers discover URLs. They pass signals, because Google still uses links, including internal ones, to understand importance and relationships. And they give context through anchor text. Think of a city's road network. Main roads lead to the important districts, local roads connect neighbourhoods, and signposts tell you where you're going. A good architecture puts valuable pages close to the centre, groups related content, and uses clear signposts.
Six principles
Six principles. Keep priority pages at a shallow click depth; depth isn't a ranking factor itself, but deep pages tend to be crawled less and receive less internal value. Group content into hubs and spokes: a hub covers a topic broadly and links to detailed spokes, which link back and to close siblings. Use descriptive anchors, not click here. Give every indexable page at least one crawlable link from an indexable page. Link to canonical, final URLs. And use breadcrumbs with BreadcrumbList markup on hierarchical sites.
Leaks and pools
Where does link value leak or pool? Mega-menus linking to hundreds of pages dilute emphasis and bloat HTML. Links inside JavaScript widgets without an href aren't discovered. Sitewide footer links to low-value pages waste emphasis. Pages reachable only through internal search are effectively orphaned. Blog posts that fall off page one of the blog index sink deeper every week. And nofollow on internal links doesn't save value; it just wastes that link's signal. Remove internal nofollow.
Measure it
How do you measure your architecture? Use a crawler to export, for each indexable URL, its crawl depth, unique inlinks and linking pages, a link score such as Screaming Frog's Link Score, and the anchor texts used. Then join with Search Console clicks and impressions and with business value. The classic finding: high-value pages with few inlinks at high depth, while low-value pages like tags, archives and filters soak up most of the internal links.
Finding orphans
Orphans deserve their own routine. Crawl from the home page. Separately, collect URLs from XML sitemaps, Search Console exports, analytics landing pages and server logs. Anything in the second list but not in the first is an orphan. For each one, decide: link it, redirect it, or remove it. Orphans with traffic or backlinks are usually the most valuable to rescue.
Hands-on: internal PageRank
Now the hands-on. The lesson text has a script using networkx. Export all internal links from your crawler, keep followed hyperlinks, and build a directed graph. Compute PageRank on it, which gives each page a relative internal importance score. Join that with crawl depth, inlinks and Search Console impressions. Then list pages with lots of impressions but weak internal PageRank. That's your prioritised list of linking opportunities, where contextual links from relevant, strong pages are most likely to help.
Example 1: agency blog (illustrative)
Worked example one, with illustrative details. An agency has four hundred blog posts. Its top-converting guide, on Google Business Profile optimisation, sits at depth seven, linked only from paginated blog archives. The fix: create a local SEO hub linked from the main navigation, link the hub to the twenty-five most relevant posts with descriptive anchors and have each link back, add contextual links from the three service pages that mention Business Profiles, and merge weak overlapping posts with redirects. After recrawling, the guide sits at depth two with dozens of relevant links.
Example 2: Jeddah 'new arrivals' trap (illustrative)
Worked example two, also illustrative. A Jeddah fashion retailer links every product from a new arrivals carousel on the home page, until the item rotates out after two weeks and loses almost all its internal links. Evergreen best-sellers end up six or more clicks deep. Fixes: permanent category and best-sellers modules with real links, related products linked by attributes like fabric and occasion, breadcrumbs with markup, and a quarterly check of depth and internal PageRank for the top two hundred revenue products.
Typical structures
Let's look at typical structures. For a service business, it's usually home, then a services hub, then individual services, with a separate guides hub and case studies. For e-commerce, it's home, then department, category, subcategory and product, supported by cross-links like related products and shop by brand, rendered as real links. For publishers, it's topic hubs plus date archives. The pattern is the same: a clear hierarchy for discovery, plus contextual cross-links that follow how users actually think about the topic.
The internal linking checklist
Here's an internal linking audit checklist to finish the lesson. Does the home page link to top categories and hubs? Are there any internal links to redirects, errors or non-canonical URLs? Is any important page deeper than sensible for the site's size? Does every indexable page have links from other indexable pages? Is anchor text descriptive and varied? Are breadcrumbs present and marked up? And do pagination and facet links behave as designed? Seven questions, and you'll catch nearly every architecture issue.
Watch me do it: linking opportunities
Watch me do it. I'll find linking opportunities for an agency blog in thirty minutes. Step one: I export all inlinks and the internal HTML report from my crawler, and a page-level Search Console export for the last three months. Step two: I run the networkx script. It builds the link graph and computes internal PageRank for about six hundred pages. Step three: I look at the output list of high-impression, low-PageRank pages. Top of the list: a guide to Google Business Profile, with lots of impressions, crawl depth seven, and three inlinks, all from paginated archives. Step four: I find strong, relevant pages to link from. I sort by internal PageRank and filter by topic. The local SEO service page, the reviews guide and the citations guide are all strong and relevant. Step five: I write the links. In each, I add a sentence where the guide genuinely helps the reader, with a descriptive anchor, like step-by-step Google Business Profile guide. Step six: I add the guide to the local SEO hub. Step seven: I recrawl. The guide is now at depth two, with eleven inlinks and a much higher internal PageRank. I note the date and check impressions and crawl frequency in a month.
Mistakes and measures
Common mistakes. Adding internal nofollow to sculpt PageRank, which doesn't redistribute value. Relying on XML sitemaps instead of links to surface important pages. Automated related links that connect pages randomly rather than by topic. And treating footers as a dumping ground. Measure success with crawl depth and internal PageRank of your priority pages, the orphan count, and crawl frequency of those pages in your logs.
Recap and try this now
Recap. Internal links drive discovery, signal importance and provide context. Keep priority pages shallow, build hubs, use descriptive anchors, eliminate orphans, and measure with depth, link scores and your own internal PageRank. Try this now. Run the networkx script on your latest crawl, and pick the top five high-impression, low-PageRank pages. For each, add three contextual links from relevant, strong pages this week.
Key takeaways
- Internal links drive discovery, pass signals and provide anchor-text context.
- Keep priority pages shallow, group content into hubs and spokes, and avoid orphans.
- Link to final canonical URLs; remove internal nofollow and links to redirects or errors.
- Find orphans by comparing a link-based crawl with sitemaps, logs, analytics and Search Console URLs.
Try it
Crawl a site, sort indexable pages by business value, and identify three high-value pages with high depth or few inlinks. Propose specific links to add.