Technical SEO MasteryStatus codes, redirects, canonicals and duplicate URLs · Lesson 7 of 18
Canonicalisation and duplicate content
Video lecture
Canonicalisation and duplicate content
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Canonicalisation and duplicate content
Here's a myth to retire on day one. Duplicate content is not, by itself, a penalty. Google has been consistent about that. But duplicates still cost you. Signals get split, Google may show the wrong URL, and crawl resources are wasted. In this lecture you'll learn where duplicates come from, which canonical signals Google weighs, how to implement rel canonical correctly, how to audit canonicals at scale with pandas, and when canonical tags simply can't fix the problem.
0:34 Why duplicates cost you
The real costs of duplicates are practical. Links and other signals get split across URL variants. Google may choose a different canonical than you want, so the wrong URL shows in results. And crawl resources get wasted fetching the same content again and again. Canonicalisation is simply telling search engines which URL in a cluster of duplicates is the representative one. Deliberately deceptive or scraped duplication is a different matter, but ordinary site duplication is a housekeeping job.
1:08 Where duplicates come from
Where do duplicates come from? Protocol and host variants, like http and https, www and non-www. Trailing slashes and uppercase letters. Parameters for tracking, sorting and sessions. Index files like index dot html. Faceted paths reachable in different orders. Print or alternate views. Products that live in several categories. Syndicated articles. And near-duplicates, like city pages where only the town name changes. Most sites have at least five of these.
1:38 Canonical signals, strongest first
Google combines canonical signals, and none is an absolute directive. From strongest to weakest: redirects, where the source effectively disappears; rel canonical in the HTML head or HTTP header, a strong hint; internal linking consistency, meaning which variant you link to; sitemap inclusion, a weak hint; and preferences such as HTTPS over HTTP and cleaner URLs. When signals conflict, say the canonical says A but internal links and the sitemap say B, Google may pick B. Consistency is the whole game.
2:13 rel=canonical rules
Implementing rel canonical correctly. Use absolute URLs with the right protocol and host. Put it in the head, because a canonical in the body is ignored, and watch out for invalid elements that close the head early. Use one canonical per page. Self-reference on canonical pages. Make sure the target is indexable: two hundred, not noindexed, not blocked, not redirected. Don't canonicalise paginated pages to page one; they're not duplicates. Don't combine noindex with a canonical pointing elsewhere. And for syndication, ask partners for a cross-domain canonical plus a clear link back, knowing Google may not always honour it.
2:56 What did Google choose?
How do you check what Google chose? URL Inspection shows the user-declared canonical and the Google-selected canonical. In the Page indexing report, look at duplicate, Google chose different canonical than user, and duplicate without user-selected canonical. Sample URLs from each and ask one question: which signal is pulling Google the other way? It's usually internal links, sitemaps or redirects disagreeing with your tag.
3:23 Hands-on: canonical audit script
Now the hands-on. Export your crawl with address, status, indexability and canonical. The pandas script in the lesson text flags six problems: missing canonicals on two-hundred pages, relative canonicals, canonicals on the wrong host, canonicals pointing at non-two-hundred URLs, canonicals pointing at non-indexable URLs, and canonical chains, where A points to B and B points to C. It prints counts and writes the flagged rows to a CSV for tickets. For JavaScript sites, run it on both raw and rendered canonicals.
3:58 Example 1: every canonical → home page
Worked example one, simple. A Sharjah furniture store's CMS sets every page's canonical to the home page after a theme update. Within weeks, product pages start dropping out of the index. The audit script flags thousands of canonical-to-home rows immediately. The fix is a template change to self-referencing canonicals, followed by monitoring Page indexing as products return. One bad default, huge impact, which is why canonical checks belong in release tests.
4:29 Example 2: near-duplicate city pages (illustrative)
Worked example two, with illustrative details. An Islamabad travel agency has a hundred and forty Umrah packages from city pages that differ only by the city name. Google picks a handful as canonical and ignores the rest. Canonical tags can't fix this, because the pages are near-duplicates meant to rank separately. The agency keeps pages only for cities it genuinely departs from, with real schedules, local contacts and prices in rupees, consolidates the rest into one page with a city selector, and redirects the retired URLs. That also avoids doorway-style patterns that Google's spam policies target.
5:11 A reusable canonical audit
Let's walk through a canonical audit procedure you can reuse on any site. First, crawl with canonical extraction on, both raw and rendered if it's a JavaScript site. Second, flag missing canonicals, canonicals to non-two-hundred or noindexed URLs, chains, relative or wrong-host canonicals, and raw-versus-rendered mismatches. Third, compare canonical targets with internal link targets: are you linking to the non-canonical variant? Fourth, compare with the sitemap, where every URL should be self-canonical. And fifth, enforce host and protocol with sitewide redirects, and strip tracking parameters from internal links. Five steps, and most canonical problems surface in an afternoon.
5:54 Consolidate or differentiate?
One more concept worth separating clearly: consolidation versus differentiation. If two URLs show the same thing, consolidate them with redirects or canonicals. If two URLs are meant to rank for different searches, they must be genuinely different, with unique, useful content. Canonical tags are a consolidation tool. They can't make near-identical pages rank separately, and trying to scale hundreds of barely different pages can drift into doorway territory under Google's spam policies. So for every duplicate cluster, ask: should these be one page, or should they be truly different pages?
6:33 Watch me do it: canonical audit
Watch me do it. I'll run a canonical audit on a fresh crawl. Step one: I export internal HTML pages with address, status, indexability and the canonical link element, then run the pandas script. It prints counts: missing canonicals, twelve. Relative canonicals, zero. Wrong host, three hundred and forty. Canonical to non-two-hundred, fifty-one. Chains, eight. Step three: I open the wrong-host rows. They all point to the staging host. A template in the blog section still has a hard-coded staging domain. That's the big one. Step four: the non-two-hundred rows point at URLs that now redirect after a category rename. The fix is updating the canonical to the final URL. Step five: I pick five flagged URLs and check URL Inspection. For the blog posts, Google chose its own canonical, ignoring the staging URL, so some damage is contained, but signals are muddled. Step six: I compare internal links with canonical targets. The main navigation links to the uppercase version of two category URLs. Step seven: I write three tickets: fix the blog template host, update canonicals after renames, and fix the navigation links. And I add a release check that fails if any canonical contains the staging host.
8:00 Mistakes and measures
Common mistakes. Every page canonicalised to the home page. Canonicals pointing to URLs that redirect. Parameter variants linked internally, like sort orders in navigation, while the canonical says otherwise. And relying on canonical tags to fix a problem a redirect should solve, such as http versus https. Measure success by the share of sampled URLs where Google's selected canonical matches yours, and by shrinking duplicate buckets in Page indexing.
8:30 Recap and try this now
Recap. Duplicates aren't a penalty, but they split signals and confuse selection. Redirects are the strongest signal, then rel canonical, then links and sitemaps. Keep every signal consistent, and don't expect canonicals to fix near-duplicate pages meant to rank separately. Try this now. Run the canonical audit script on your latest crawl, then use URL Inspection on five flagged URLs to see whether Google agrees with you.
Duplicate content is a consolidation problem, not a penalty
Google's position is consistent: duplicate content on your own site is normal and not a penalty in itself (deliberately deceptive, scraped or manipulative duplication is different). The real costs are practical:
- Signals such as links get split across URL variants.
- Google may choose a different canonical than you want — the wrong one shows in results.
- Crawl resources are wasted on duplicates.
Canonicalisation is the process of telling search engines which URL of a duplicate cluster is the representative one.
Where duplicates come from
| Source | Example |
|---|---|
| Protocol and host | http://, https://, www., non-www |
| Trailing slash and case | /Shoes/ vs /shoes |
| Parameters | ?utm_source=, ?sort=price, ?sessionid= |
| Index files | /index.html vs / |
| Faceted paths | /shoes/red/ and /red/shoes/ |
| Print/AMP/alternate views | /page?print=1 |
| Product in multiple categories | /men/trainers/x and /sale/x |
| Syndication | Your article republished on a partner site |
| Near-duplicates | City pages with only the town name changed |
The canonical signals, strongest first
Google combines signals; none is an absolute directive.
- Redirects (301/308) — strongest; the source effectively disappears.
- rel="canonical" in HTML head or HTTP header — strong hint.
- Internal linking consistency — which variant you link to.
- Sitemap inclusion — weak hint.
- HTTPS over HTTP, and other preferences such as cleaner URLs.
When signals conflict (canonical says A, internal links and sitemap say B), Google may pick B. Consistency is the whole game.
Implementing rel=canonical correctly
<link rel="canonical" href="https://www.example.com/shoes/red-trainers/">For non-HTML files:
Link: <https://www.example.com/whitepaper.pdf>; rel="canonical"Rules:
- Use absolute URLs with the correct protocol and host.
- Place it in the
head; a canonical in the body is ignored. Invalid elements in the head (for example an unclosed tag or injecteddiv) can end the head early, pushing the canonical into the body. - Use one canonical per page. Multiple conflicting canonicals are typically ignored.
- Self-reference on canonical pages — it protects against parameter variants.
- The canonical target must be indexable: 200 status, not noindexed, not blocked, not redirected.
- Do not canonicalise paginated pages 2, 3, 4 to page 1 — they are not duplicates.
- Do not combine
noindexwith a canonical pointing elsewhere — it sends mixed signals. - For cross-domain syndication, ask partners to add a cross-domain canonical to your original. Google notes this may not always be honoured for syndicated content, so also ask for a clear link back to the original, or for noindex on their copy if appropriate.
Checking what Google chose
URL Inspection shows the user-declared canonical and Google-selected canonical. In the Page indexing report, look at "Duplicate, Google chose different canonical than user" and "Duplicate without user-selected canonical". Sample URLs from each and ask: which signal is pulling Google the other way?
A canonical audit procedure
- Crawl the site with canonical extraction on (both raw and rendered for JS sites).
- Flag: missing canonicals on indexable pages; canonicals to non-200 URLs; canonicals to noindexed URLs; canonical chains (A→B, B→C); relative or wrong-host canonicals; mismatches between raw and rendered HTML.
- Compare canonical targets with internal link targets — are you linking to the non-canonical variant?
- Compare with sitemap URLs — every sitemap URL should be self-canonical.
- Enforce host and protocol with sitewide redirects; strip tracking parameters from internal links.
Near-duplicates and thin variants
Canonical tags cannot fix pages that are meant to rank separately but are nearly identical — for example 200 "SEO services in [city]" pages with swapped place names. Either make each genuinely useful (local proof, staff, case studies, pricing, directions) or consolidate. Large volumes of templated doorway-style pages can fall under Google's spam policies on doorway abuse and scaled content abuse.
Hands-on: a canonical audit on a crawl export (pandas)
Export from your crawler (Screaming Frog, Sitebulb or similar) a CSV with at least: Address, Status Code, Indexability, Canonical Link Element 1, and the rendered canonical if you crawled with JavaScript. Then:
import pandas as pd
df = pd.read_csv("internal_html.csv")
df = df.rename(columns={"Address": "url", "Status Code": "status",
"Canonical Link Element 1": "canonical"})
status = dict(zip(df.url, df.status))
idx = dict(zip(df.url, df.Indexability))
df["canon_missing"] = df.canonical.isna() & (df.status == 200)
df["canon_relative"] = df.canonical.fillna("").str.match(r"^/")
df["canon_wrong_host"]= df.canonical.fillna("").str.match(r"^https?://") & \
~df.canonical.fillna("").str.startswith("https://www.example.com/")
df["canon_target_status"] = df.canonical.map(status)
df["canon_to_non200"] = df.canon_target_status.notna() & (df.canon_target_status != 200)
df["canon_to_nonindexable"] = df.canonical.map(idx).eq("Non-Indexable") & (df.canonical != df.url)
df["canon_chain"] = df.canonical.map(dict(zip(df.url, df.canonical))).ne(df.canonical) & \
df.canonical.ne(df.url) & df.canonical.isin(df.url)
issues = ["canon_missing", "canon_relative", "canon_wrong_host", "canon_to_non200",
"canon_to_nonindexable", "canon_chain"]
print(df[issues].sum())
df[df[issues].any(axis=1)][["url", "canonical"] + issues].to_csv("canonical_issues.csv", index=False)Then compare with Search Console: URL Inspection on a sample of flagged URLs, and the Page indexing buckets "Duplicate, Google chose different canonical than user" and "Duplicate without user-selected canonical".
Worked example 2: an Islamabad travel agency's city pages
An Islamabad travel agency (illustrative) has 140 "Umrah packages from [city]" pages that differ only by the city name. Google picks a handful as canonical and ignores the rest. Canonical tags can't fix this — the pages are near-duplicates intended to rank separately. The agency keeps pages only for cities it genuinely departs from (with real departure schedules, local contact numbers and prices in PKR), consolidates the rest into one page with a city selector, and 301s the retired URLs. That reduces duplication and avoids doorway-style patterns that Google's spam policies target.
Common mistakes
- Every page canonicalised to the home page (a CMS misconfiguration that can wipe out indexing).
- Canonical tags pointing to URLs that redirect.
- Parameter variants linked internally (sort orders in navigation) while the canonical says otherwise.
- Relying on canonical to fix a problem a redirect should solve (e.g. http vs https).
Key takeaways
- Duplicate content is a signal-splitting and canonical-selection problem, not an automatic penalty.
- Redirects are the strongest consolidation signal; rel=canonical is a strong hint that needs consistent supporting signals.
- Canonical targets must be absolute, indexable, 200-status URLs; paginated pages should self-canonicalise.
- Use URL Inspection to compare the user-declared and Google-selected canonical.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick one template, list every URL variant that returns 200 for the same content, and write the redirect and canonical rules that consolidate them.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.