Technical SEO MasteryStatus codes, redirects, canonicals and duplicate URLs · Lesson 8 of 18

Pagination and faceted navigation

Article · 14 min · 9 min lecture

Video lecture

Pagination and faceted navigation

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Pagination and faceted navigation

  • Pagination done right
  • Decide, then implement facets
  • Measure which parameters eat the crawl

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why these two patterns cause most large-site problems

Category listings split across pages, and filters that combine colour, size, brand, price and sort order, are essential for users. For crawlers, they can create an almost infinite number of URLs with overlapping content. Handling them well is a hallmark of senior technical SEO.

Pagination: what Google recommends now

Google stopped using rel="next" and rel="prev" as an indexing signal in 2019 (they remain harmless and some other engines and accessibility tools may use them). Current best practice:

  1. Give each page a unique URL, e.g. /shoes/?page=2 or /shoes/page/2/.
  2. Link pages sequentially with real anchor links (next, previous, and ideally numbered links so deep pages are fewer clicks away).
  3. Self-canonicalise each page. Do not canonicalise everything to page 1 — items on page 5 may then never be discovered through that path.
  4. Keep page 1 at the clean URL and redirect ?page=1 to it (or canonicalise it there).
  5. Do not noindex paginated pages by default. It can reduce crawling of the links on them over time. Consider it only if you have another strong discovery path (for example complete sitemaps and dense internal linking) and a real reason.
  6. View-all pages are fine if they load fast; then paginated pages can canonicalise to the view-all.
<!-- On /shoes/?page=2 -->
<link rel="canonical" href="https://www.example.com/shoes/?page=2">
<nav aria-label="Pagination">
  <a href="/shoes/">1</a> <a href="/shoes/?page=2" aria-current="page">2</a>
  <a href="/shoes/?page=3">3</a> <a href="/shoes/?page=3">Next</a>
</nav>

Infinite scroll and "Load more": Google does not scroll or click. Back infinite scroll with paginated URLs that are also linked as anchors, and update the URL with the History API as users scroll.

Faceted navigation: decide, then implement

Google published dedicated faceted navigation guidance (December 2024). The first step is a business decision per facet:

Facet typeSearch demand?Treatment
Brand, key colour or type (/trainers/nike/)Often yesIndexable, clean static URL, unique title/H1/intro, in sitemap
Size, price range, ratingUsually lowCrawlable for users only, or blocked
Sort order, view mode, items-per-pageNoneBlock or keep out of crawl paths
Multi-select combinations (3+ facets)RarelyBlock

Implementation options for facets you do not want in search:

  1. robots.txt disallow for parameter patterns — the most effective way to stop crawling of large facet spaces:
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*price=
Disallow: /*?*size=*&
  1. URL fragments (/shoes#size=9) — Google generally ignores fragments, so filter states are not crawled as separate URLs.
  2. rel=canonical to the unfiltered category — reduces indexing of variants, but Google still crawls them, so it does little for crawl efficiency on huge sites.
  3. noindex — keeps variants out of the index but they are still crawled.
  4. nofollow on filter links — a hint only; not a reliable crawl control.

For facets you do want indexed:

  • Use a clean, consistent path or parameter order (?brand=nike&colour=red, never also ?colour=red&brand=nike).
  • Return 404 for filter combinations with no results rather than an empty 200 page.
  • Generate unique titles, headings and ideally a short intro.
  • Include them in sitemaps and link to them from relevant categories.

Worked example 1: a fashion marketplace

A Karachi-based marketplace has 3,000 categories and 14 filter types. Logs show the large majority of Googlebot hits going to parameter URLs, while new products take weeks to be crawled. The team:

  1. Pulls keyword demand for each facet type; brand and colour for key categories have demand, others do not.
  2. Creates static indexable URLs for demand-backed brand and colour pages (a few thousand pages, each with a unique intro).
  3. Moves all other filters to parameters disallowed in robots.txt, with consistent ordering.
  4. Returns 404 for zero-result combinations.
  5. Re-measures in logs after several weeks: share of bot hits on parameter URLs, crawl frequency of new product pages, and indexed count of valuable facet pages.

(The approach is standard; how quickly and how much crawl behaviour shifts varies by site — measure rather than promise.)

Hands-on: classify facet URLs from your logs or crawl

Before deciding treatments, measure how facets behave. This regex-driven classifier counts URLs by parameter combination so you can see which facets generate the most URLs and crawl:

import re
from collections import Counter
from urllib.parse import urlsplit, parse_qsl

LOW_VALUE = {"sort", "view", "per_page", "price", "size", "sessionid", "utm_source", "utm_medium"}
counts, crawlhits = Counter(), Counter()
for line in open("googlebot_urls.txt"):                  # one URL (or path) per line from verified logs
    q = urlsplit(line.strip()).query
    keys = tuple(sorted(k for k, _ in parse_qsl(q, keep_blank_values=True)))
    counts[keys] += 1
    if LOW_VALUE & set(keys):
        crawlhits["low_value"] += 1
    elif len(keys) >= 3:
        crawlhits["multi_facet"] += 1
    elif keys:
        crawlhits["single_facet"] += 1
    else:
        crawlhits["clean"] += 1

for keys, n in counts.most_common(15):
    print(n, "&".join(keys) or "(no parameters)")
print(crawlhits)

Useful robots.txt patterns once decisions are made (test before shipping):

User-agent: *
Disallow: /*?*sort=
Disallow: /*?*view=
Disallow: /*?*per_page=
Disallow: /*?*sessionid=

Worked example 2: a UK DIY retailer's "Load more"

A UK DIY retailer (illustrative) replaced numbered pagination with a JavaScript "Load more" button. Products beyond the first 36 in each category lost their internal links from category pages, and logs showed many long-tail products uncrawled for weeks. The fix kept the button for users but backed it with real paginated URLs (?page=2, ?page=3) rendered as anchor links, each self-canonical, with the URL updated via the History API as users load more.

Common mistakes

  • Blocking in robots.txt URLs that are already indexed and noindexed — Google can no longer see the noindex. Sequence your changes.
  • Canonicalising all pages of a category to page 1.
  • Allowing sort parameters to be linked from every listing.
  • Letting internal search result pages be crawled and indexed — Google's guidelines have long advised against indexing auto-generated search result pages.

Key takeaways

  • Paginated pages need unique URLs, crawlable sequential links and self-referencing canonicals; rel=next/prev is no longer a Google signal.
  • Back infinite scroll and Load more with linked paginated URLs.
  • Decide per facet whether it has search demand; index only demand-backed facets with clean, consistent URLs.
  • robots.txt or fragments control facet crawling most effectively; canonical and noindex still allow crawling.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which method most effectively stops Google crawling millions of sort and price filter URLs?
  2. A filter combination returns zero products. What should the URL return?
  3. A site uses infinite scroll with no paginated URLs. What is the SEO risk?

Put it into practice

For one category on a site you know, classify every filter as index / crawl-only / block, and draft the robots.txt rules and URL patterns.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.