Skip to content

Technical SEO Mastery · Site architecture, internal linking and XML sitemaps · lesson 10 of 18 · 12 min

XML sitemaps done properly

What sitemaps are (and are not)

An XML sitemap is a list of URLs you want search engines to know about, with optional metadata. It aids discovery, especially for large sites, new sites with few links, and pages that change often. It does not guarantee crawling or indexing, and it does not replace internal linking.

Hard limits and core rules

  • Up to 50,000 URLs or 50 MB uncompressed per sitemap file (gzip is allowed).
  • Use a sitemap index to list multiple sitemaps (an index can list up to 50,000 sitemaps).
  • Use absolute, canonical URLs that return 200 and are indexable.
  • UTF-8 encoding, entity-escape special characters (& becomes &).
  • A sitemap can include URLs only from its own host and path unless cross-submission is verified in Search Console.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemaps/products-1.xml.gz</loc>
    <lastmod>2026-03-14T08:00:00+00:00</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://www.example.com/sitemaps/guides.xml</loc>
  </sitemap>
</sitemapindex>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/guides/core-web-vitals/</loc>
    <lastmod>2026-02-27</lastmod>
  </url>
</urlset>

lastmod, priority and changefreq

  • lastmod: Google uses it if it is consistently and verifiably accurate — meaning it reflects the last significant change to the main content, not the time the sitemap was generated. Setting every URL to "today" teaches Google to ignore your lastmod.
  • priority and changefreq: Google ignores them. You can omit them.

Specialised sitemaps

Image sitemaps help discovery of images loaded in ways crawlers may miss:

<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
  <url>
    <loc>https://www.example.com/products/oud-perfume/</loc>
    <image:image><image:loc>https://cdn.example.com/img/oud-perfume.avif</image:loc></image:image>
  </url>
</urlset>

Google deprecated the image caption, title, geo_location and licence tags; image:loc is what matters.

Video sitemaps list video landing pages with video:thumbnail_loc, video:title, video:description and either video:content_loc or video:player_loc. They help Google find videos and understand metadata; the video must be on a page where it is prominent.

News sitemaps (for publishers eligible for Google News features) should include only articles published in the last two days, up to 1,000 URLs per sitemap, with news:publication, news:publication_date and news:title.

hreflang in sitemaps is an alternative to HTML tags for international sites (Module 5).

Submission and discovery

  1. Reference sitemaps in robots.txt: Sitemap: https://www.example.com/sitemap_index.xml.
  2. Submit in Search Console (Sitemaps report) and Bing Webmaster Tools.
  3. Note that Google's old sitemap "ping" endpoint was retired in 2023; rely on robots.txt, Search Console and accurate lastmod. Bing, Yandex, Seznam, Naver and Yep support IndexNow for instant change notifications (Lesson 2.3); Google does not.

Segmentation for diagnostics

Split sitemaps by template or section: products-*.xml, categories.xml, guides.xml, locations.xml. In Search Console you can then filter the Page indexing report by sitemap and see, for example, that most products are "Crawled – currently not indexed" while guides are fine. That is a powerful diagnostic that one giant file hides.

Sitemap hygiene audit

  1. Download all sitemaps (follow the index).
  2. Crawl the URL list: every URL should return 200, be self-canonical, indexable and not blocked by robots.txt.
  3. Compare with your site crawl: indexable canonical pages missing from sitemaps, and sitemap URLs missing from the crawl (orphans).
  4. Check lastmod accuracy against CMS update dates on a sample.
  5. Check Search Console for "Couldn't fetch" or parsing errors.
  6. Confirm generation is automatic — sitemaps built by hand drift out of date.

Large-site tip: sitemaps as a freshness feed

On sites with millions of URLs, keep a small "recent changes" sitemap containing only URLs created or substantially updated in the last few days, alongside the full segmented set. Because its lastmod values are genuine, it gives crawlers a compact, trustworthy list of what changed — useful for marketplaces, classifieds and publishers where new inventory must be discovered quickly. Rotate URLs out once they are older than your chosen window.

Hands-on: a sitemap hygiene checker

# pip install requests
import gzip, io, requests, xml.etree.ElementTree as ET
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}

def fetch_xml(url):
    r = requests.get(url, timeout=30); r.raise_for_status()
    data = gzip.decompress(r.content) if url.endswith(".gz") else r.content
    return ET.fromstring(data)

def urls_from(index_url):
    root = fetch_xml(index_url)
    if root.tag.endswith("sitemapindex"):
        for loc in root.findall("sm:sitemap/sm:loc", NS):
            yield from urls_from(loc.text.strip())
    else:
        for u in root.findall("sm:url", NS):
            lm = u.find("sm:lastmod", NS)
            yield u.find("sm:loc", NS).text.strip(), (lm.text if lm is not None else None)

problems, seen = [], 0
for url, lastmod in urls_from("https://www.example.com/sitemap_index.xml"):
    seen += 1
    r = requests.get(url, allow_redirects=False, timeout=20)
    canon = None
    if r.status_code == 200 and "text/html" in r.headers.get("Content-Type", ""):
        i = r.text.lower().find('rel="canonical"')
        if i != -1:
            j = r.text.find('href="', i); canon = r.text[j + 6:r.text.find('"', j + 6)] if j != -1 else None
    noindex = "noindex" in r.headers.get("X-Robots-Tag", "").lower() or 'content="noindex' in r.text.lower()
    if r.status_code != 200 or noindex or (canon and canon != url):
        problems.append((url, r.status_code, "noindex" if noindex else "", canon or ""))
print(f"{seen} URLs checked, {len(problems)} problems")
for p in problems[:50]:
    print(p)

For large sitemaps, sample (for example 2,000 URLs per segment) rather than fetching everything, respect your own server's capacity, and run from an allow-listed IP. A crawler's list mode does the same job with a UI.

Worked example 2: a Lahore property portal's lastmod

A Lahore property portal (illustrative) generated sitemaps nightly with every URL's lastmod set to the generation time. Google, as documented, only uses lastmod when it is consistently accurate, so the field carried no signal. The team changed the generator to use each listing's real "last significant change" timestamp (price, status or description changes — not view counters), split sitemaps into active listings, projects and guides, and added a small "changed in the last 48 hours" sitemap. Logs then showed updated listings recrawled sooner.

Common mistakes

  • Including redirected, 404, noindexed or non-canonical URLs ("dirty" sitemaps weaken trust in the file).
  • Letting a plugin generate sitemaps for tag archives, attachment pages or internal search.
  • Using lastmod as a "freshness hack".
  • Forgetting to update sitemaps after a migration — the new sitemap should list new URLs, and temporarily submitting the old URL list can help Google discover the redirects.

Video lecture: XML sitemaps done properly

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

  1. XML sitemaps done properly
  2. What sitemaps are
  3. Hard limits and rules
  4. lastmod, priority, changefreq
  5. Specialised sitemaps
  6. Submit and segment
  7. Hands-on: sitemap hygiene checker
  8. Example 1: Leeds recruitment agency
  9. Example 2: Lahore property portal (illustrative)
  10. Large sites: the freshness sitemap
  11. The hygiene routine
  12. Watch me do it: sitemap segment audit
  13. Mistakes and measures
  14. Recap and try this now

Lecture transcript

XML sitemaps done properly

XML sitemaps look like a solved problem. A plugin generates one, you submit it, done. But on real audits, sitemaps are full of redirects, noindexed pages, non-canonical URLs and fake lastmod dates, and that teaches search engines to trust them less. In this lecture you'll learn what sitemaps are for, the hard limits, how Google uses lastmod, specialised sitemaps, segmentation for diagnostics, and a Python checker that finds dirty URLs automatically.

What sitemaps are

A sitemap is a list of URLs you want search engines to know about, with optional metadata. It helps discovery, especially for large sites, new sites with few links, and pages that change often. It doesn't guarantee crawling or indexing, and it doesn't replace internal linking. Think of it as a list you hand to a delivery driver. Useful, especially in a big new estate. But if the roads don't connect to the houses, the list alone won't get everything delivered.

Hard limits and rules

The hard limits. Up to fifty thousand URLs or fifty megabytes uncompressed per file, and gzip is allowed. Use a sitemap index to list multiple sitemaps. Use absolute, canonical URLs that return two hundred and are indexable. Encode in UTF-8 and escape special characters, so an ampersand becomes amp. And a sitemap can only include URLs from its own host and path, unless cross-submission is verified in Search Console.

lastmod, priority, changefreq

Now lastmod, priority and changefreq. Google uses lastmod only if it's consistently and verifiably accurate, meaning it reflects the last significant change to the main content, not the time the sitemap was generated. Set every URL to today, and you teach Google to ignore your lastmod. Priority and changefreq? Google ignores them, so you can omit them. One field matters, and only if you tell the truth.

Specialised sitemaps

Specialised sitemaps. Image sitemaps help discovery of images loaded in ways crawlers might miss; Google now only needs image loc, having deprecated the caption, title, geo-location and licence tags. Video sitemaps describe video landing pages with thumbnail, title, description and a content or player location. News sitemaps, for eligible publishers, list only articles from the last two days, up to a thousand URLs. And hreflang can live in sitemaps too, which we'll cover in the international lesson.

Submit and segment

Submission and discovery. Reference your sitemap index in robots.txt. Submit it in Search Console and Bing Webmaster Tools. Google's old sitemap ping endpoint was retired in twenty twenty-three, so rely on robots.txt, Search Console and accurate lastmod. For Bing and several other engines, IndexNow gives you instant change notifications. And segment your sitemaps by template: products, categories, guides, locations. Then you can filter the Page indexing report by sitemap and see, for example, that products are mostly crawled, currently not indexed, while guides are fine. One giant file hides that.

Hands-on: sitemap hygiene checker

Hands-on. The lesson text includes a Python checker. It follows the sitemap index recursively, handles gzip, and for each URL fetches the page without following redirects. It flags anything that isn't a two hundred, anything with noindex in the header or meta robots, and anything whose canonical points elsewhere. For big sites, sample a couple of thousand URLs per segment and run it from an allow-listed IP. A crawler's list mode can do the same job with a user interface.

Example 1: Leeds recruitment agency

Worked example one, simple. A Leeds recruitment agency's sitemap plugin includes every expired job with a two-hundred page saying this role has closed, plus tag archives and attachment pages. The checker flags hundreds of problems. The fixes: exclude tags and attachments in the plugin settings, remove closed jobs from the sitemap, and give closed jobs a four-oh-four or a useful page linking to similar live roles. The sitemap shrinks to URLs worth crawling.

Example 2: Lahore property portal (illustrative)

Worked example two, with illustrative details. A Lahore property portal generated sitemaps nightly, with every URL's lastmod set to the generation time. So the field carried no signal. The team changed the generator to use each listing's real last significant change, meaning price, status or description, not view counters. They split sitemaps into active listings, projects and guides, and added a small changed in the last forty-eight hours sitemap. Logs then showed updated listings recrawled sooner.

Large sites: the freshness sitemap

For very large sites, here's a technique worth knowing: a freshness sitemap. Alongside your full, segmented sitemaps, keep a small sitemap that contains only URLs created or substantially updated in the last few days. Because its lastmod values are genuine, it gives crawlers a compact, trustworthy list of what changed. That's valuable for marketplaces, classifieds and publishers, where new inventory must be discovered quickly. Rotate URLs out once they're older than your chosen window, and keep the full sitemaps as the complete record.

The hygiene routine

Let me give you the sitemap hygiene audit as a routine. Download all sitemaps by following the index. Crawl the URL list: every URL should return two hundred, be self-canonical, indexable and not blocked by robots.txt. Compare with your site crawl to find indexable canonical pages missing from sitemaps, and sitemap URLs missing from the crawl, which are orphans. Check lastmod accuracy against CMS update dates on a sample. Check Search Console for couldn't fetch or parsing errors. And confirm generation is automatic, because hand-built sitemaps drift out of date.

Watch me do it: sitemap segment audit

Watch me do it. I'll audit a sitemap segment for a property portal. Step one: I run the hygiene checker on the listings sitemap index, sampling two thousand URLs. Output: two thousand checked, three hundred and ten problems. Step two: I open the problem rows. Two hundred and forty return three-oh-one: listings that sold and now redirect to the area page. Fifty are noindexed: draft listings that leaked into the sitemap. Twenty have a canonical pointing to another listing: duplicates created by agents re-posting. Step three: I check lastmod. Every URL has the same timestamp: two a.m. today. The generator writes the build time. Step four: I open Search Console's Sitemaps report and the Page indexing report filtered by this sitemap. Many URLs sit in crawled, currently not indexed, which is consistent with a sitemap Google has learned not to trust. Step five: I write the fix as one ticket with three rules: include only active, indexable, self-canonical listings; set lastmod from the listing's real last significant change; and add a small recent changes sitemap for the last forty-eight hours. After deploy, I re-run the checker: zero problems in the sample, and lastmod values finally vary.

Mistakes and measures

Common mistakes. Including redirected, four-oh-four, noindexed or non-canonical URLs, which weakens trust in the file. Letting a plugin generate sitemaps for tag archives, attachments or internal search. Using lastmod as a freshness hack. And forgetting to update sitemaps after a migration: the new sitemap should list new URLs, and temporarily keeping a list of old URLs can help Google discover the redirects. Measure success with the share of sitemap URLs that are clean, and the indexed-to-submitted ratio per segment.

Recap and try this now

Recap. Sitemaps aid discovery but don't guarantee indexing. Keep them clean: canonical, indexable, two-hundred URLs only. Make lastmod truthful or leave it out. Segment by template for diagnostics, and use IndexNow for engines that support it. Try this now. Run the hygiene checker on one sitemap segment, and fix the source of every problem it finds, rather than editing the file by hand.

Key takeaways

  • Sitemaps aid discovery; they do not guarantee indexing or replace internal links.
  • Limits: 50,000 URLs or 50 MB uncompressed per file; use a sitemap index beyond that.
  • Google uses lastmod only when consistently accurate and ignores priority and changefreq.
  • Segment sitemaps by template to diagnose indexing by section in Search Console.

Try it

Download a site's sitemaps, crawl the URL list, and report how many URLs are non-200, non-canonical or noindexed.