Technical SEO MasteryInternational SEO and structured data · Lesson 11 of 18

hreflang and international SEO

Article · 15 min · 8 min lecture

Video lecture

hreflang and international SEO

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

hreflang and international SEO

  • Choose a structure
  • Syntax and the rules that break
  • A reciprocity checker

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Two separate decisions

International SEO has two layers: site structure (where each version lives) and annotation (how you tell search engines which version is for whom).

Choosing a structure

StructureExampleProsCons
ccTLDexample.ae, example.pkClear country signal, local trustSeparate domains to build and maintain
Subdirectoryexample.com/en-gb/, example.com/ar-sa/Consolidated authority, simpler opsWeaker country signal alone
Subdomainuk.example.comCan host separatelyOperates somewhat like a separate site
Parametersexample.com?lang=arEasy to buildNot recommended

For most businesses expanding from one site, subdirectories are the pragmatic default. Search Console's old International Targeting (country targeting) report was retired in 2022, so geotargeting now relies on the ccTLD, hreflang, and local signals such as currency, addresses, language and local links.

Language and locale fundamentals

  • Translate content genuinely. Machine translation without review risks low quality; a region version with only currency changed may be seen as duplicate.
  • Mark the language in HTML: html lang="ar" dir="rtl" for Arabic.
  • Never auto-redirect by IP without an alternative. Googlebot crawls mostly from US IP addresses, so IP redirects can hide your local versions. Use a dismissible banner suggesting the local version instead.
  • Let users switch versions with crawlable links.

hreflang syntax

hreflang tells Google that URLs are equivalents in different languages or regions, so it can show the right one. Values are an ISO 639-1 language code, optionally followed by an ISO 3166-1 alpha-2 region code.

<link rel="alternate" hreflang="en-gb" href="https://www.example.com/en-gb/pricing/">
<link rel="alternate" hreflang="en-ae" href="https://www.example.com/en-ae/pricing/">
<link rel="alternate" hreflang="ar-ae" href="https://www.example.com/ar-ae/pricing/">
<link rel="alternate" hreflang="ar-sa" href="https://www.example.com/ar-sa/pricing/">
<link rel="alternate" hreflang="ur-pk" href="https://www.example.com/ur-pk/pricing/">
<link rel="alternate" hreflang="en" href="https://www.example.com/en/pricing/">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/pricing/">

Every version in the set must include this same complete list, including itself.

The rules that break most implementations

  1. Return links (reciprocity). If page A lists B, B must list A. Unconfirmed annotations may be ignored.
  2. Self-reference. Each page includes itself.
  3. Canonical alignment. hreflang URLs must be the canonical versions that return 200. Pointing to redirected or non-canonical URLs breaks the set. Each language version should self-canonicalise — never canonicalise /ar-ae/ to /en/.
  4. Valid codes. en-uk is invalid (the region code is gb). A region alone (hreflang="gb") is invalid. Language comes first.
  5. x-default for the fallback version (often a language selector or global English page).
  6. Consistency across methods. Use one method: HTML link tags, HTTP Link headers (for PDFs), or XML sitemaps. Mixing them invites conflicts.

hreflang in XML sitemaps

For large sites, sitemaps keep HTML lighter and are easier to generate centrally:

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:xhtml="http://www.w3.org/1999/xhtml">
  <url>
    <loc>https://www.example.com/en-gb/pricing/</loc>
    <xhtml:link rel="alternate" hreflang="en-gb" href="https://www.example.com/en-gb/pricing/"/>
    <xhtml:link rel="alternate" hreflang="ar-ae" href="https://www.example.com/ar-ae/pricing/"/>
    <xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/pricing/"/>
  </url>
  <!-- repeat a url block for each version, with the full set -->
</urlset>

hreflang audit procedure

  1. Crawl all language versions with hreflang extraction.
  2. Check for: missing return links, non-200 or non-canonical targets, invalid codes, missing self-reference, missing x-default, and pages whose canonical conflicts with hreflang.
  3. Check coverage: does every version of a page exist? Missing translations should be omitted from the set, not pointed at another language.
  4. In Search Console, filter Performance by country and check which URL version is getting impressions in each market.
  5. Search from the target market (VPN or a SERP localisation tool) for branded and key queries.

Worked example 1: UAE users seeing UK prices

A SaaS company runs /en/, /en-gb/, /ar/. UAE users see the UK page with GBP pricing. Audit shows /ar/ pages canonicalise to /en/ (a CMS default), so Google drops them from the hreflang cluster. Fixes: self-canonicals, add en-ae pages with AED pricing where the business actually serves the UAE, a complete reciprocal set, and x-default pointing to a selector page.

Hands-on: an hreflang reciprocity checker

Export from your crawler a CSV of hreflang annotations: one row per (source URL, hreflang value, target URL), plus each URL's status and canonical. Then:

import pandas as pd, re
h = pd.read_csv("hreflang.csv")                     # columns: source, lang, target
pages = pd.read_csv("internal_html.csv").rename(columns={"Address": "url", "Status Code": "status",
                                                        "Canonical Link Element 1": "canonical"})
status = dict(zip(pages.url, pages.status)); canon = dict(zip(pages.url, pages.canonical))
VALID = re.compile(r"^(x-default|[a-z]{2,3}(-[a-z]{4})?(-([a-z]{2}|\d{3}))?)$", re.I)

pairs = set(zip(h.source, h.target))
h["invalid_code"]   = ~h.lang.str.match(VALID)
h["uk_mistake"]     = h.lang.str.lower().str.endswith("-uk")
h["no_return_link"] = [(t, s) not in pairs for s, t in zip(h.source, h.target)]
h["target_not_200"] = h.target.map(status).ne(200)
h["target_not_canonical"] = h.target.map(canon).notna() & h.target.map(canon).ne(h.target)
self_ref = h[h.source == h.target].source.unique()
missing_self = set(h.source.unique()) - set(self_ref)

issues = ["invalid_code", "uk_mistake", "no_return_link", "target_not_200", "target_not_canonical"]
print(h[issues].sum()); print(len(missing_self), "pages missing a self-referencing hreflang")
h[h[issues].any(axis=1)].to_csv("hreflang_issues.csv", index=False)

(-uk passes a generic format check but is invalid for the United Kingdom — the ISO 3166-1 code is gb — hence the separate flag.)

Worked example 2: Arabic and English in Saudi Arabia

A Riyadh B2B software company (illustrative) runs /en-sa/ and /ar-sa/ plus a global /en/. Arabic pages rank poorly and English pages show to Arabic searchers. The checker finds ar pages listing en-sa but not themselves, and an x-default pointing at a URL that redirects. Fixes: complete self-referencing sets on both language versions, x-default pointing to the final global URL, html lang="ar" dir="rtl" on Arabic templates, and genuinely localised Arabic copy (not machine translation) for key pages. In Search Console they then filter Performance by country = Saudi Arabia and watch which URL version earns impressions.

Common mistakes

  • Using hreflang to target countries you do not actually serve with distinct content.
  • Relying on hreflang to fix poor translations.
  • Forgetting that hreflang is a hint; strong local relevance signals still matter.

Key takeaways

  • Pick a structure (ccTLD, subdirectory, subdomain) first; subdirectories are a pragmatic default.
  • hreflang values use ISO 639-1 language plus optional ISO 3166-1 region (en-gb, not en-uk).
  • Sets must be complete, self-referencing, reciprocal and point only at canonical 200 URLs.
  • Do not force IP-based redirects; Googlebot mostly crawls from US IPs.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which hreflang value is valid for English in the United Kingdom?
  2. The Arabic page lists English as an alternate, but the English page does not list Arabic. Result?
  3. Why is automatic IP-based redirection risky for international sites?

Put it into practice

Write a complete hreflang set (HTML tags) for a page available in English (UK), English (UAE), Arabic (Saudi Arabia) and Urdu (Pakistan), plus x-default.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.