Programmatic SEO and AI Content at Scale — Without Getting PenalizedIndexing and crawl management at scale · Lesson 8 of 14

Indexing management: sitemaps, crawl budget and noindex strategy

Article · 8 min · 8 min lecture

Video lecture

Indexing management: sitemaps, crawl budget and noindex strategy

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Indexing management

  • Indexing as a quality signal
  • Crawl budget
  • Sitemaps
  • The right tool for each goal
  • Staged rollout

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Publishing is not indexing

Search engines do not index every page they find. They decide based on perceived value, duplication, site quality and resources. For page systems, indexing is the first real quality signal: if Google declines to index a large share of a page type ("Crawled – currently not indexed", "Discovered – currently not indexed"), it is telling you something about value or discovery.

Crawl budget: when it matters

Google's guidance says crawl budget is mainly a concern for very large sites (on the order of a million or more unique pages) or medium-large sites with content changing very frequently, and for sites with many "Discovered – currently not indexed" URLs. Crawl capacity depends on your server's health and responsiveness; crawl demand depends on popularity and freshness. You influence it by:

  • Removing low-value URLs from crawl paths (facets, duplicates, infinite spaces).
  • Fast, reliable servers (errors and slow responses reduce crawl rate).
  • Clear internal linking to important pages.
  • Accurate sitemaps with truthful lastmod.

XML sitemaps for page systems

  • Split sitemaps by page type (and by segment for large types) so you can monitor indexing per type in Search Console.
  • Respect limits: 50,000 URLs or 50 MB uncompressed per sitemap; use a sitemap index for more.
  • Include only canonical, indexable, 200-status URLs.
  • Use lastmod accurately (Google uses it when it is consistently reliable); Google ignores changefreq and priority.
from datetime import timezone
from xml.sax.saxutils import escape

def write_sitemap(path, rows):  # rows: iterable of (url, last_modified_datetime)
    with open(path, "w", encoding="utf-8") as f:
        f.write('<?xml version="1.0" encoding="UTF-8"?>\n')
        f.write('<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">\n')
        n = 0
        for url, lm in rows:
            f.write(f"  <url><loc>{escape(url)}</loc><lastmod>{lm.astimezone(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ')}</lastmod></url>\n")
            n += 1
            if n >= 50000:
                raise ValueError("split sitemap: 50,000 URL limit")
        f.write("</urlset>\n")

noindex, robots.txt and canonical: use the right tool

GoalToolWhy
Keep a page out of the index but let crawlers see it<meta name="robots" content="noindex"> or X-Robots-Tag headerCrawlers must be able to fetch the page to see noindex
Stop crawling of URL spaces (facets, parameters)robots.txt DisallowSaves crawl; but disallowed URLs can still be indexed without content if linked
Consolidate duplicates to a preferred URLrel="canonical"A hint; combine with consistent internal links
Remove permanently404/410Clear signal; update links and sitemaps
Move/merge301 redirectPasses signals to the target

Don't combine robots.txt Disallow with noindex on the same URL — Google can't see the noindex if it can't crawl.

A staged rollout strategy

  1. Pilot: publish 50–200 pages of one type; include them in a dedicated sitemap.
  2. Observe 4–8 weeks: indexing rate, impressions, clicks, engagement.
  3. Improve the template and thresholds based on what was and wasn't indexed.
  4. Expand in waves, keeping per-type sitemaps and monitoring.
  5. Hold back or noindex segments that underperform.

Monitoring tools

  • Search Console Page indexing report filtered by sitemap.
  • URL Inspection (and its API, with daily quotas) for samples.
  • Crawl stats report: response codes, response time, file types.
  • Server logs: which bots crawl which page types, how often (verify Googlebot via reverse DNS).
  • IndexNow: supported by Bing, Yandex and others for instant change notifications (Google does not use IndexNow).
  • The Indexing API is only for pages with JobPosting or BroadcastEvent (livestream) structured data — not general pages.

Worked example: a UK property data site

The site generated 120,000 "house prices in [street]" pages. After launch, most sat in "Discovered – currently not indexed". Analysis: thin pages for streets with one or two sales, weak internal links, and a slow database-backed render. Fixes: threshold of at least five sales in 5 years (others merged into postcode-district pages), faster cached rendering, links from district hubs, per-type sitemaps. The indexed share rose over the following months and traffic concentrated on the stronger pages.

Pitfalls

  • Submitting every URL in sitemaps, including noindexed or redirected ones.
  • Using robots.txt to "noindex".
  • Launching 100,000 pages at once.
  • Fake lastmod updates.

How to measure success

Indexed share per page type trending up, "Discovered/Crawled – currently not indexed" shrinking for qualified pages, healthy crawl stats (low error rates, stable response time), and indexed pages earning impressions.

Key takeaways

  • Indexing is the first quality signal for page systems — track it per page type.
  • Crawl budget matters mainly for very large or fast-changing sites; improve it by removing low-value URLs and speeding servers.
  • Split sitemaps by type, include only canonical indexable URLs, use honest lastmod.
  • Use noindex, robots.txt, canonical, 404/410 and 301 for their distinct purposes; don't block noindexed pages in robots.txt.
  • Roll out in stages and monitor with Search Console, crawl stats and server logs.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. You want a page crawled but not indexed. Which tool is correct?
  2. Which sitemap fields does Google ignore?
  3. For which pages is Google's Indexing API intended?

Put it into practice

Design the sitemap structure and rollout plan for your page system: sitemap files per type, pilot size, observation window, thresholds and the metrics that trigger expansion.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.