Programmatic SEO and AI Content at Scale — Without Getting Penalized · Building page systems: templates, links and structured data · lesson 6 of 14 · 8 min
Internal linking at scale: hubs, facets and related links
Why internal links make or break page systems
Search engines discover, prioritize and understand pages largely through links. A programmatic system with thousands of pages and weak internal linking produces orphan pages (no internal links pointing to them) that are rarely crawled or indexed. Good linking also helps users navigate — which is the main point.
The architecture: hubs and spokes
Home
└── Hub: /tutors/ (all subjects, all cities)
├── Subject hub: /tutors/physics/ (cities for physics)
│ ├── /tutors/physics/lahore/
│ └── /tutors/physics/karachi/
└── City hub: /tutors/in/lahore/ (subjects in Lahore)
├── /tutors/physics/lahore/
└── /tutors/math/lahore/
- Hubs list and summarize children (with their own useful data: counts, price ranges).
- Spokes link up to hubs (breadcrumbs), across to siblings (nearby cities, related subjects) and down to entities (individual tutor profiles).
- Keep important pages within a few clicks of the home page.
Faceted navigation without crawl traps
Filters (price, rating, language, online/in-person) can generate millions of URL combinations. Strategy:
- Decide which facet combinations deserve indexable landing pages (those matching real query patterns with demand, e.g., "female physics tutors in Lahore" if that is searched).
- Keep other filter combinations crawlable only if useful, otherwise prevent crawling of parameter combinations (e.g., via robots.txt disallow for certain parameters) and point canonicals to the base page.
- Avoid infinite spaces: calendars, sort orders and session IDs in URLs.
- Google's guidance on faceted navigation recommends preventing crawling of facet URLs that don't need indexing and using consistent URL parameter ordering.
Related-link algorithms
"Related" modules should be relevant, not random. Useful signals:
- Geography: nearest areas by distance.
- Taxonomy: parent/sibling categories.
- Behavior: co-viewed or co-booked pages (aggregated).
- Similarity: content or attribute similarity.
import math
def haversine_km(a, b):
lat1, lon1, lat2, lon2 = map(math.radians, [a[0], a[1], b[0], b[1]])
d = math.sin((lat2 - lat1) / 2) ** 2 + math.cos(lat1) * math.cos(lat2) * math.sin((lon2 - lon1) / 2) ** 2
return 6371 * 2 * math.asin(math.sqrt(d))
areas = { # area -> (lat, lon, published?)
"dubai-marina": (25.0805, 55.1403, True), "jlt": (25.0693, 55.1413, True),
"al-barsha": (25.1115, 55.1990, True), "business-bay": (25.1850, 55.2650, False),
}
def nearby(area, k=3):
here = areas[area][:2]
cands = [(haversine_km(here, v[:2]), a) for a, v in areas.items() if a != area and v[2]]
return [a for _, a in sorted(cands)[:k]]
print(nearby("dubai-marina"))
Only link to published, indexable pages from related modules; linking to noindexed or empty pages wastes crawl and frustrates users.
Anchor text and placement
- Use descriptive anchors ("Physics tutors in Karachi"), varied naturally.
- Contextual links inside content carry more meaning than footer link dumps.
- Limit very large link blocks; paginate hubs.
Detecting orphans and depth problems
Crawl your site with a crawler (e.g., Screaming Frog, Sitebulb, or an open-source crawler) and compare with your sitemap and database:
- URLs in sitemap but not found by crawl = orphans.
- Crawl depth distribution per page type.
- Internal inlinks per page — pages with very few inlinks are at risk.
Worked example: a Karachi-based B2B software directory
The directory has 3,000 "[software] alternatives" and "[software A] vs [software B]" pages. A crawl found 40% of vs-pages were orphans because they were linked only from sitemaps. Fix: each product page links to its top 5 comparison pages by search demand and co-view data; category hubs list top comparisons; breadcrumbs added. Indexing coverage of vs-pages rose after the change (monitored in Search Console's page indexing report).
Pagination and hub quality
Large hubs (for example, "all physics tutors in Pakistan") need pagination. Use plain, crawlable links to subsequent pages (?page=2 or /page/2/), give each paginated page a self-referencing canonical, and avoid pointing all paginated pages' canonicals to page one — that hides the items on later pages. Make page one of each hub genuinely useful on its own: a summary of the data (counts, price ranges, top areas), the most relevant items first, and links into sub-hubs. A hub that is only a list of links is itself a thin page.
Pitfalls
- Relying on XML sitemaps instead of internal links.
- Random "related" links.
- Letting facets create millions of crawlable URLs.
- Linking to pages that are noindexed or below threshold.
How to measure success
Zero orphan pages among indexable URLs, click depth within target, crawl stats showing efficient crawling of important page types, and improved indexing rates after linking changes.
Video lecture: Internal linking at scale: hubs, facets and related links
Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.
- Internal linking at scale
- Why it matters
- Hubs and spokes
- Faceted navigation
- Related links
- Simple example: Riyadh car rental (illustrative)
- Anchors and placement
- Finding problems
- Pagination and hub quality
- Example: Karachi software directory (illustrative)
- Pitfalls and success
- Mistakes + try this now
- Watch me do it: fixing orphans (illustrative)
- Recap and next step
Lecture transcript
Internal linking at scale
You can publish ten thousand excellent pages and still get almost none of them indexed, if nothing links to them. Internal links are how search engines discover pages, decide which matter, and understand how they relate. And they're how users move around. In this lecture you'll learn hub and spoke architecture, how to handle faceted navigation without creating crawl traps, how to build related links that are actually relevant, and how to find orphan pages.
Why it matters
Why does this matter? Because pages that aren't linked are, for practical purposes, pages that don't exist. Here's an analogy. Think of a city's road network. A beautiful new neighborhood with no roads leading to it will stay empty, no matter how nice the houses are. Internal links are your roads. Hubs are the main highways, breadcrumbs are the signposts back to the center, and related links are the side streets that help people explore. Build the roads and both visitors and crawlers arrive.
Hubs and spokes
Start with the architecture. A top hub lists everything in the system. Subject hubs list cities for a subject. City hubs list subjects in a city. Each spoke page, like physics tutors in Lahore, links up to its hubs through breadcrumbs, across to siblings like nearby cities and related subjects, and down to individual entities like tutor profiles. Hubs shouldn't be bare link lists; give them their own data, like counts and price ranges. And keep important pages within a few clicks of the home page.
Faceted navigation
Filters are where page systems quietly explode. Price, rating, language, online or in-person, sort order: combine them and you get millions of URLs. So decide which combinations deserve indexable landing pages, the ones matching real query patterns with demand. Prevent crawling of the rest, point canonicals to the base page, keep parameter order consistent, and avoid infinite spaces like calendars, sort orders and session IDs in URLs. Google's own faceted navigation guidance says much the same.
Related links
Related link modules should be relevant, not random. Use geography, like the nearest areas by distance. Taxonomy, like parent and sibling categories. Behavior, like pages that are often viewed or booked together, in aggregate. And similarity of content or attributes. The lesson includes a small Python function that finds the nearest published areas using the haversine distance formula. Note the word published. Only link to indexable pages with real content.
Simple example: Riyadh car rental (illustrative)
Here's a simple worked example. A Riyadh car rental comparison site has pages for car rental in each district. Before: pages linked only from the XML sitemap. After: a Riyadh hub lists all districts with price ranges; each district page has breadcrumbs back to Riyadh and to car rental in Saudi Arabia; and a related module shows the three nearest districts with available cars, calculated by distance. Pages for districts with no cars available aren't linked from the related module. Within weeks, crawl stats show more frequent crawling of district pages, and more of them get indexed.
Anchors and placement
A few details on anchors and placement. Use descriptive anchor text, like physics tutors in Karachi, varied naturally rather than stuffed. Links inside the content, in context, carry more meaning than a dump of links in the footer. And avoid enormous link blocks; paginate hubs instead.
Finding problems
How do you find problems? Crawl your own site with a crawler, and compare what it finds with your sitemap and your database. URLs in the sitemap that the crawl never reached are orphans. Check the crawl depth distribution by page type. And look at internal links per page, because pages with very few are at risk of being ignored.
Pagination and hub quality
Big hubs need pagination, so handle it cleanly. Use plain, crawlable links to the next pages, give each paginated page its own self-referencing canonical, and don't point every page's canonical back to page one, because that hides everything listed on later pages. And make page one of every hub genuinely useful on its own, with a data summary like counts, price ranges and top areas, the most relevant items first, and links into sub-hubs. A hub that's only a list of links is itself a thin page.
Example: Karachi software directory (illustrative)
Here's an illustrative example. A Karachi-based software directory had three thousand alternatives and versus pages. A crawl showed about forty percent of the versus pages were orphans, linked only from sitemaps. The fix: every product page links to its top five comparisons by demand and co-views, category hubs list top comparisons, and breadcrumbs were added. After the change, indexing coverage of those pages rose, tracked in Search Console's page indexing report.
Pitfalls and success
Pitfalls to avoid. Relying on XML sitemaps instead of links. Random related links. Letting facets create millions of crawlable URLs. And linking to pages that are noindexed or below threshold. Success looks like zero orphans among indexable pages, click depth within target, efficient crawling of important page types in Search Console's crawl stats, and better indexing after linking changes.
Mistakes + try this now
Common linking mistakes. Relying on sitemaps instead of real links. Random related links that confuse users. Letting filters create millions of crawlable URLs. Linking to noindexed or empty pages. And hubs that are nothing but long lists of links. Try this now: crawl your site with any crawler, export the list of URLs, and compare it to your sitemap. Count the URLs in your sitemap that the crawler never found. Each one is an orphan waiting for a link.
Watch me do it: fixing orphans (illustrative)
Watch me do it. Let's fix orphans on an illustrative Lahore restaurant directory. Step one, I crawl the site with a crawler starting from the home page and export all found URLs. Step two, I export the sitemap URLs and the database's list of published pages. Step three, a quick comparison in Python: two thousand four hundred pages in the database, two thousand three hundred in the sitemap, but only one thousand five hundred found by the crawl. So about eight hundred pages are orphans. Step four, I look at which kind: almost all are cuisine in area pages for smaller areas, which are only reachable through the site search. Step five, the fix: every area hub lists its cuisines with counts; every cuisine page links to the three nearest areas that have that cuisine; breadcrumbs link cuisine in area pages up to both the area hub and the cuisine hub. Step six, recrawl after deployment: two thousand three hundred and eighty pages found; the remaining twenty are below threshold and correctly noindexed and unlinked. Step seven, Search Console's crawl stats over the next month confirm more frequent crawling of those pages.
Recap and next step
Recap. Internal links drive discovery, priority and navigation. Build hubs and spokes with breadcrumbs and sibling links. Control facets. Make related links relevant and point only at published pages. Crawl regularly. Your next step: draw the hub and spoke structure for your page system, and write the rules for your related-link module: which signals, how many links, and which pages are eligible.
Key takeaways
- Internal links drive discovery, prioritization and user navigation; sitemaps aren't a substitute.
- Use hubs and spokes with breadcrumbs, sibling links and entity links; keep key pages shallow.
- Control faceted navigation: index only combinations matching real demand; prevent crawl traps.
- Build relevant related-link modules from geography, taxonomy, behavior and similarity — only to published pages.
- Crawl regularly to find orphans and depth problems.
Try it
Draw the hub-and-spoke structure for your page system and define the related-link rules (signals, number of links, eligibility).