Technical SEO MasteryStatus codes, redirects, canonicals and duplicate URLs · Lesson 8 of 18
Pagination and faceted navigation
Video lecture
Pagination and faceted navigation
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Pagination and faceted navigation
On large sites, two everyday features cause more crawl and indexing problems than almost anything else: pagination and faceted navigation. Users love filters for colour, size, brand and price. Crawlers can drown in them. In this lecture you'll learn how Google recommends handling pagination today, how to decide which facets deserve to be indexed, the implementation options and their trade-offs, and a small script that shows you which parameters are eating your crawl.
0:32 Pagination today
Pagination first. Google stopped using rel next and rel prev as an indexing signal back in twenty nineteen. They're harmless, but they don't do what people think. Today's best practice is simple. Give each page a unique URL. Link pages sequentially with real anchor links, ideally with numbered links so deep pages are fewer clicks away. Self-canonicalise each page, rather than pointing everything at page one. Keep page one at the clean URL. And don't noindex paginated pages by default, because that can reduce crawling of the links on them.
1:11 Infinite scroll and 'Load more'
What about infinite scroll and load more buttons? Google doesn't scroll or click. So if products beyond the first batch only appear after a click, Google may never see them, and they lose their internal links. Back infinite scroll with real paginated URLs that are also linked as anchors, and update the URL with the History API as people scroll. Users keep the smooth experience. Crawlers get a path to every item.
1:42 Facets: decide first
Now facets. Google published dedicated faceted navigation guidance in December twenty twenty-four, and the first step is a business decision per facet. Brand, key colour or type, like Nike trainers, often has real search demand, so it can be indexable, on a clean static URL, with a unique title, heading and intro, and included in sitemaps. Size, price range and rating usually have low demand. Sort order, view mode and items per page have none. And combinations of three or more facets rarely do. Decide first, implement second.
2:20 Implementation options
For facets you don't want in search, you have five options, with different strengths. robots.txt disallow for parameter patterns is the most effective way to stop crawling of huge facet spaces. URL fragments, the part after the hash, aren't crawled as separate URLs. rel canonical to the unfiltered category reduces indexing but Google still crawls the variants. Noindex keeps them out of the index but they're still crawled. And nofollow on filter links is only a hint. For crawl efficiency on big sites, robots.txt and fragments do the heavy lifting.
2:59 Indexable facets + sequencing
For facets you do want indexed: use a clean, consistent URL or parameter order, never both brand then colour and colour then brand. Return four-oh-four for combinations with no results rather than an empty two-hundred page. Generate unique titles and headings, ideally with a short intro. And include them in sitemaps and link to them from relevant categories. And one sequencing warning: don't block in robots.txt URLs that are already indexed and noindexed, because Google can no longer see the noindex. Noindex first, wait, then block.
3:36 Hands-on: facet classifier
Hands-on. Before you decide treatments, measure. The lesson text has a short Python script that reads URLs from your verified Googlebot logs or your crawl, groups them by parameter combination, and counts how many fall into low-value parameters, multi-facet combinations, single facets and clean URLs. The top fifteen combinations usually reveal the culprits within seconds. Then use the robots.txt patterns in the lesson, and test them before shipping.
4:06 Example 1: Karachi marketplace (illustrative)
Worked example one, with illustrative details. A Karachi-based marketplace has three thousand categories and fourteen filter types. Logs show most Googlebot hits going to parameter URLs, while new products take weeks to be crawled. The team checks keyword demand per facet: brand and colour for key categories have demand. They create static indexable URLs for those, a few thousand pages with unique intros. Everything else moves to parameters disallowed in robots.txt with consistent ordering. Zero-result combinations return four-oh-four. Then they re-measure in logs after several weeks.
4:43 Example 2: UK DIY 'Load more' (illustrative)
Worked example two, also illustrative. A UK DIY retailer replaced numbered pagination with a JavaScript load more button. Products beyond the first thirty-six in each category lost their internal links, and logs showed many long-tail products uncrawled for weeks. The fix kept the button for users, but backed it with real paginated URLs rendered as anchor links, each self-canonical, with the URL updated through the History API as users load more. Within a few weeks, the missing products started appearing in the logs again.
5:20 View-all pages
A note on view-all pages. If you have a view-all version of a category that loads quickly, it's fine to let paginated pages point their canonical at the view-all page, because it genuinely contains everything. But if view-all is slow, heavy, or would push the HTML past sensible sizes, don't do it. Remember Googlebot only fetches the first two megabytes of HTML. For most large catalogues, well-linked, self-canonical paginated pages are the safer default.
5:52 Internal search pages
Internal search result pages deserve a special mention, because they combine every facet problem at once. Every query a user types can become a new URL. If those URLs are linked, for example from a popular searches widget, crawlers will follow them, and you'll get thin, near-infinite pages in the index. Google's guidelines have long advised against letting auto-generated search result pages be indexed. The usual treatment is to disallow the search path in robots.txt, remove links to search result URLs, and, if you want to capture popular queries, build proper category or guide pages for them instead.
6:35 Watch me do it: facet decision prep
Watch me do it. I'll run a facet decision workshop's prep for a fashion retailer. Step one: I extract a week of verified Googlebot URLs from the logs into a text file and run the facet classifier. The top combination is colour plus size plus sort, with thousands of hits. The counts show low-value parameters take more than half of all facet hits. Step two: I pull keyword demand for brand and colour facets in the top twenty categories. Colour pages for dresses and abayas have real demand; size and price don't. Step three: I build the decision table. Brand and key colour pages become indexable, static URLs with unique intros. Size, price and sort stay available for users but aren't crawlable links. Step four: I check which parameter URLs are already indexed with a site search and Search Console's Page indexing. A few hundred sort URLs are indexed. Because I want them gone, the sequence is noindex first, wait, then disallow. Step five: I draft the robots.txt lines and test them against a list of URLs that must stay crawlable, like the approved colour pages. All pass. The prep takes two hours, and the meeting takes thirty minutes, because the data answers most questions.
8:04 Mistakes and measures
Common mistakes. Blocking URLs in robots.txt that are already indexed and noindexed. Canonicalising every page of a category to page one. Letting sort parameters be linked from every listing. And letting internal search result pages be crawled and indexed, which Google's guidelines have long advised against. Measure success in the logs: the share of Googlebot hits on low-value parameters going down, the crawl frequency of products going up, and the indexed count of your approved facet pages.
8:37 Recap and try this now
Recap. Paginate with unique, linked, self-canonical URLs. Back infinite scroll with real pages. For facets, decide by demand, then implement: indexable clean pages for the valuable few, and robots.txt or fragments for the rest. Sequence noindex before blocking. Try this now. Run the facet classifier on a week of verified Googlebot URLs and write down the three parameter combinations taking the most crawl. That's your facet decision meeting agenda.
Why these two patterns cause most large-site problems
Category listings split across pages, and filters that combine colour, size, brand, price and sort order, are essential for users. For crawlers, they can create an almost infinite number of URLs with overlapping content. Handling them well is a hallmark of senior technical SEO.
Pagination: what Google recommends now
Google stopped using rel="next" and rel="prev" as an indexing signal in 2019 (they remain harmless and some other engines and accessibility tools may use them). Current best practice:
- Give each page a unique URL, e.g.
/shoes/?page=2or/shoes/page/2/. - Link pages sequentially with real anchor links (next, previous, and ideally numbered links so deep pages are fewer clicks away).
- Self-canonicalise each page. Do not canonicalise everything to page 1 — items on page 5 may then never be discovered through that path.
- Keep page 1 at the clean URL and redirect
?page=1to it (or canonicalise it there). - Do not noindex paginated pages by default. It can reduce crawling of the links on them over time. Consider it only if you have another strong discovery path (for example complete sitemaps and dense internal linking) and a real reason.
- View-all pages are fine if they load fast; then paginated pages can canonicalise to the view-all.
<!-- On /shoes/?page=2 -->
<link rel="canonical" href="https://www.example.com/shoes/?page=2">
<nav aria-label="Pagination">
<a href="/shoes/">1</a> <a href="/shoes/?page=2" aria-current="page">2</a>
<a href="/shoes/?page=3">3</a> <a href="/shoes/?page=3">Next</a>
</nav>Infinite scroll and "Load more": Google does not scroll or click. Back infinite scroll with paginated URLs that are also linked as anchors, and update the URL with the History API as users scroll.
Faceted navigation: decide, then implement
Google published dedicated faceted navigation guidance (December 2024). The first step is a business decision per facet:
| Facet type | Search demand? | Treatment |
|---|---|---|
Brand, key colour or type (/trainers/nike/) | Often yes | Indexable, clean static URL, unique title/H1/intro, in sitemap |
| Size, price range, rating | Usually low | Crawlable for users only, or blocked |
| Sort order, view mode, items-per-page | None | Block or keep out of crawl paths |
| Multi-select combinations (3+ facets) | Rarely | Block |
Implementation options for facets you do not want in search:
- robots.txt disallow for parameter patterns — the most effective way to stop crawling of large facet spaces:
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*price=
Disallow: /*?*size=*&- URL fragments (
/shoes#size=9) — Google generally ignores fragments, so filter states are not crawled as separate URLs. - rel=canonical to the unfiltered category — reduces indexing of variants, but Google still crawls them, so it does little for crawl efficiency on huge sites.
- noindex — keeps variants out of the index but they are still crawled.
- nofollow on filter links — a hint only; not a reliable crawl control.
For facets you do want indexed:
- Use a clean, consistent path or parameter order (
?brand=nike&colour=red, never also?colour=red&brand=nike). - Return 404 for filter combinations with no results rather than an empty 200 page.
- Generate unique titles, headings and ideally a short intro.
- Include them in sitemaps and link to them from relevant categories.
Worked example 1: a fashion marketplace
A Karachi-based marketplace has 3,000 categories and 14 filter types. Logs show the large majority of Googlebot hits going to parameter URLs, while new products take weeks to be crawled. The team:
- Pulls keyword demand for each facet type; brand and colour for key categories have demand, others do not.
- Creates static indexable URLs for demand-backed brand and colour pages (a few thousand pages, each with a unique intro).
- Moves all other filters to parameters disallowed in robots.txt, with consistent ordering.
- Returns 404 for zero-result combinations.
- Re-measures in logs after several weeks: share of bot hits on parameter URLs, crawl frequency of new product pages, and indexed count of valuable facet pages.
(The approach is standard; how quickly and how much crawl behaviour shifts varies by site — measure rather than promise.)
Hands-on: classify facet URLs from your logs or crawl
Before deciding treatments, measure how facets behave. This regex-driven classifier counts URLs by parameter combination so you can see which facets generate the most URLs and crawl:
import re
from collections import Counter
from urllib.parse import urlsplit, parse_qsl
LOW_VALUE = {"sort", "view", "per_page", "price", "size", "sessionid", "utm_source", "utm_medium"}
counts, crawlhits = Counter(), Counter()
for line in open("googlebot_urls.txt"): # one URL (or path) per line from verified logs
q = urlsplit(line.strip()).query
keys = tuple(sorted(k for k, _ in parse_qsl(q, keep_blank_values=True)))
counts[keys] += 1
if LOW_VALUE & set(keys):
crawlhits["low_value"] += 1
elif len(keys) >= 3:
crawlhits["multi_facet"] += 1
elif keys:
crawlhits["single_facet"] += 1
else:
crawlhits["clean"] += 1
for keys, n in counts.most_common(15):
print(n, "&".join(keys) or "(no parameters)")
print(crawlhits)Useful robots.txt patterns once decisions are made (test before shipping):
User-agent: *
Disallow: /*?*sort=
Disallow: /*?*view=
Disallow: /*?*per_page=
Disallow: /*?*sessionid=Worked example 2: a UK DIY retailer's "Load more"
A UK DIY retailer (illustrative) replaced numbered pagination with a JavaScript "Load more" button. Products beyond the first 36 in each category lost their internal links from category pages, and logs showed many long-tail products uncrawled for weeks. The fix kept the button for users but backed it with real paginated URLs (?page=2, ?page=3) rendered as anchor links, each self-canonical, with the URL updated via the History API as users load more.
Common mistakes
- Blocking in robots.txt URLs that are already indexed and noindexed — Google can no longer see the noindex. Sequence your changes.
- Canonicalising all pages of a category to page 1.
- Allowing sort parameters to be linked from every listing.
- Letting internal search result pages be crawled and indexed — Google's guidelines have long advised against indexing auto-generated search result pages.
Key takeaways
- Paginated pages need unique URLs, crawlable sequential links and self-referencing canonicals; rel=next/prev is no longer a Google signal.
- Back infinite scroll and Load more with linked paginated URLs.
- Decide per facet whether it has search demand; index only demand-backed facets with clean, consistent URLs.
- robots.txt or fragments control facet crawling most effectively; canonical and noindex still allow crawling.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
For one category on a site you know, classify every filter as index / crawl-only / block, and draft the robots.txt rules and URL patterns.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.