Technical SEO MasteryStatus codes, redirects, canonicals and duplicate URLs · Lesson 7 of 18

Canonicalisation and duplicate content

Article · 13 min · 9 min lecture

Video lecture

Canonicalisation and duplicate content

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Canonicalisation and duplicate content

  • Not a penalty, but a real cost
  • Canonical signals and rules
  • Auditing at scale

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Duplicate content is a consolidation problem, not a penalty

Google's position is consistent: duplicate content on your own site is normal and not a penalty in itself (deliberately deceptive, scraped or manipulative duplication is different). The real costs are practical:

  • Signals such as links get split across URL variants.
  • Google may choose a different canonical than you want — the wrong one shows in results.
  • Crawl resources are wasted on duplicates.

Canonicalisation is the process of telling search engines which URL of a duplicate cluster is the representative one.

Where duplicates come from

SourceExample
Protocol and hosthttp://, https://, www., non-www
Trailing slash and case/Shoes/ vs /shoes
Parameters?utm_source=, ?sort=price, ?sessionid=
Index files/index.html vs /
Faceted paths/shoes/red/ and /red/shoes/
Print/AMP/alternate views/page?print=1
Product in multiple categories/men/trainers/x and /sale/x
SyndicationYour article republished on a partner site
Near-duplicatesCity pages with only the town name changed

The canonical signals, strongest first

Google combines signals; none is an absolute directive.

  1. Redirects (301/308) — strongest; the source effectively disappears.
  2. rel="canonical" in HTML head or HTTP header — strong hint.
  3. Internal linking consistency — which variant you link to.
  4. Sitemap inclusion — weak hint.
  5. HTTPS over HTTP, and other preferences such as cleaner URLs.

When signals conflict (canonical says A, internal links and sitemap say B), Google may pick B. Consistency is the whole game.

Implementing rel=canonical correctly

<link rel="canonical" href="https://www.example.com/shoes/red-trainers/">

For non-HTML files:

Link: <https://www.example.com/whitepaper.pdf>; rel="canonical"

Rules:

  • Use absolute URLs with the correct protocol and host.
  • Place it in the head; a canonical in the body is ignored. Invalid elements in the head (for example an unclosed tag or injected div) can end the head early, pushing the canonical into the body.
  • Use one canonical per page. Multiple conflicting canonicals are typically ignored.
  • Self-reference on canonical pages — it protects against parameter variants.
  • The canonical target must be indexable: 200 status, not noindexed, not blocked, not redirected.
  • Do not canonicalise paginated pages 2, 3, 4 to page 1 — they are not duplicates.
  • Do not combine noindex with a canonical pointing elsewhere — it sends mixed signals.
  • For cross-domain syndication, ask partners to add a cross-domain canonical to your original. Google notes this may not always be honoured for syndicated content, so also ask for a clear link back to the original, or for noindex on their copy if appropriate.

Checking what Google chose

URL Inspection shows the user-declared canonical and Google-selected canonical. In the Page indexing report, look at "Duplicate, Google chose different canonical than user" and "Duplicate without user-selected canonical". Sample URLs from each and ask: which signal is pulling Google the other way?

A canonical audit procedure

  1. Crawl the site with canonical extraction on (both raw and rendered for JS sites).
  2. Flag: missing canonicals on indexable pages; canonicals to non-200 URLs; canonicals to noindexed URLs; canonical chains (A→B, B→C); relative or wrong-host canonicals; mismatches between raw and rendered HTML.
  3. Compare canonical targets with internal link targets — are you linking to the non-canonical variant?
  4. Compare with sitemap URLs — every sitemap URL should be self-canonical.
  5. Enforce host and protocol with sitewide redirects; strip tracking parameters from internal links.

Near-duplicates and thin variants

Canonical tags cannot fix pages that are meant to rank separately but are nearly identical — for example 200 "SEO services in [city]" pages with swapped place names. Either make each genuinely useful (local proof, staff, case studies, pricing, directions) or consolidate. Large volumes of templated doorway-style pages can fall under Google's spam policies on doorway abuse and scaled content abuse.

Hands-on: a canonical audit on a crawl export (pandas)

Export from your crawler (Screaming Frog, Sitebulb or similar) a CSV with at least: Address, Status Code, Indexability, Canonical Link Element 1, and the rendered canonical if you crawled with JavaScript. Then:

import pandas as pd
df = pd.read_csv("internal_html.csv")
df = df.rename(columns={"Address": "url", "Status Code": "status",
                        "Canonical Link Element 1": "canonical"})
status = dict(zip(df.url, df.status))
idx = dict(zip(df.url, df.Indexability))

df["canon_missing"]   = df.canonical.isna() & (df.status == 200)
df["canon_relative"]  = df.canonical.fillna("").str.match(r"^/")
df["canon_wrong_host"]= df.canonical.fillna("").str.match(r"^https?://") & \
                        ~df.canonical.fillna("").str.startswith("https://www.example.com/")
df["canon_target_status"] = df.canonical.map(status)
df["canon_to_non200"] = df.canon_target_status.notna() & (df.canon_target_status != 200)
df["canon_to_nonindexable"] = df.canonical.map(idx).eq("Non-Indexable") & (df.canonical != df.url)
df["canon_chain"] = df.canonical.map(dict(zip(df.url, df.canonical))).ne(df.canonical) & \
                    df.canonical.ne(df.url) & df.canonical.isin(df.url)

issues = ["canon_missing", "canon_relative", "canon_wrong_host", "canon_to_non200",
          "canon_to_nonindexable", "canon_chain"]
print(df[issues].sum())
df[df[issues].any(axis=1)][["url", "canonical"] + issues].to_csv("canonical_issues.csv", index=False)

Then compare with Search Console: URL Inspection on a sample of flagged URLs, and the Page indexing buckets "Duplicate, Google chose different canonical than user" and "Duplicate without user-selected canonical".

Worked example 2: an Islamabad travel agency's city pages

An Islamabad travel agency (illustrative) has 140 "Umrah packages from [city]" pages that differ only by the city name. Google picks a handful as canonical and ignores the rest. Canonical tags can't fix this — the pages are near-duplicates intended to rank separately. The agency keeps pages only for cities it genuinely departs from (with real departure schedules, local contact numbers and prices in PKR), consolidates the rest into one page with a city selector, and 301s the retired URLs. That reduces duplication and avoids doorway-style patterns that Google's spam policies target.

Common mistakes

  • Every page canonicalised to the home page (a CMS misconfiguration that can wipe out indexing).
  • Canonical tags pointing to URLs that redirect.
  • Parameter variants linked internally (sort orders in navigation) while the canonical says otherwise.
  • Relying on canonical to fix a problem a redirect should solve (e.g. http vs https).

Key takeaways

  • Duplicate content is a signal-splitting and canonical-selection problem, not an automatic penalty.
  • Redirects are the strongest consolidation signal; rel=canonical is a strong hint that needs consistent supporting signals.
  • Canonical targets must be absolute, indexable, 200-status URLs; paginated pages should self-canonicalise.
  • Use URL Inspection to compare the user-declared and Google-selected canonical.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Page A has rel=canonical to B, but the sitemap and all internal links point to A. What is likely?
  2. How should page 3 of a paginated category declare its canonical?
  3. Where must a rel=canonical link element sit to be honoured?

Put it into practice

Pick one template, list every URL variant that returns 200 for the same content, and write the redirect and canonical rules that consolidate them.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.