---
title: "Indexing management: sitemaps, crawl budget and noindex…"
description: "Publishing is not indexing Search engines do not index every page they find. They decide based on perceived value, duplication, site quality and…"
url: https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/indexing-and-crawl-management
updated: 2026-10-05
---

Programmatic SEO and AI Content at Scale — Without Getting Penalized · Indexing and crawl management at scale · lesson 8 of 14 · 8 min

# Indexing management: sitemaps, crawl budget and noindex strategy

## Publishing is not indexing

Search engines do not index every page they find. They decide based on perceived value, duplication, site quality and resources. For page systems, **indexing is the first real quality signal**: if Google declines to index a large share of a page type ("Crawled – currently not indexed", "Discovered – currently not indexed"), it is telling you something about value or discovery.

## Crawl budget: when it matters

Google's guidance says crawl budget is mainly a concern for **very large sites** (on the order of a million or more unique pages) or medium-large sites with content changing very frequently, and for sites with many "Discovered – currently not indexed" URLs. Crawl capacity depends on your server's health and responsiveness; crawl demand depends on popularity and freshness. You influence it by:

- Removing low-value URLs from crawl paths (facets, duplicates, infinite spaces).
- Fast, reliable servers (errors and slow responses reduce crawl rate).
- Clear internal linking to important pages.
- Accurate sitemaps with truthful `lastmod`.

## XML sitemaps for page systems

- Split sitemaps **by page type** (and by segment for large types) so you can monitor indexing per type in Search Console.
- Respect limits: 50,000 URLs or 50 MB uncompressed per sitemap; use a sitemap index for more.
- Include only canonical, indexable, 200-status URLs.
- Use `lastmod` accurately (Google uses it when it is consistently reliable); Google ignores `changefreq` and `priority`.

```python
from datetime import timezone
from xml.sax.saxutils import escape

def write_sitemap(path, rows):  # rows: iterable of (url, last_modified_datetime)
    with open(path, "w", encoding="utf-8") as f:
        f.write('<?xml version="1.0" encoding="UTF-8"?>\n')
        f.write('<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">\n')
        n = 0
        for url, lm in rows:
            f.write(f"  <url><loc>{escape(url)}</loc><lastmod>{lm.astimezone(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ')}</lastmod></url>\n")
            n += 1
            if n >= 50000:
                raise ValueError("split sitemap: 50,000 URL limit")
        f.write("</urlset>\n")
```

## noindex, robots.txt and canonical: use the right tool

| Goal | Tool | Why |
|---|---|---|
| Keep a page out of the index but let crawlers see it | `<meta name="robots" content="noindex">` or `X-Robots-Tag` header | Crawlers must be able to fetch the page to see noindex |
| Stop crawling of URL spaces (facets, parameters) | `robots.txt` Disallow | Saves crawl; but disallowed URLs can still be indexed without content if linked |
| Consolidate duplicates to a preferred URL | `rel="canonical"` | A hint; combine with consistent internal links |
| Remove permanently | 404/410 | Clear signal; update links and sitemaps |
| Move/merge | 301 redirect | Passes signals to the target |

**Don't combine robots.txt Disallow with noindex** on the same URL — Google can't see the noindex if it can't crawl.

## A staged rollout strategy

1. **Pilot**: publish 50–200 pages of one type; include them in a dedicated sitemap.
2. **Observe** 4–8 weeks: indexing rate, impressions, clicks, engagement.
3. **Improve** the template and thresholds based on what was and wasn't indexed.
4. **Expand** in waves, keeping per-type sitemaps and monitoring.
5. **Hold back** or noindex segments that underperform.

## Monitoring tools

- Search Console **Page indexing** report filtered by sitemap.
- **URL Inspection** (and its API, with daily quotas) for samples.
- **Crawl stats** report: response codes, response time, file types.
- **Server logs**: which bots crawl which page types, how often (verify Googlebot via reverse DNS).
- **IndexNow**: supported by Bing, Yandex and others for instant change notifications (Google does not use IndexNow).
- The **Indexing API** is only for pages with JobPosting or BroadcastEvent (livestream) structured data — not general pages.

## Worked example: a UK property data site

The site generated 120,000 "house prices in [street]" pages. After launch, most sat in "Discovered – currently not indexed". Analysis: thin pages for streets with one or two sales, weak internal links, and a slow database-backed render. Fixes: threshold of at least five sales in 5 years (others merged into postcode-district pages), faster cached rendering, links from district hubs, per-type sitemaps. The indexed share rose over the following months and traffic concentrated on the stronger pages.

## Pitfalls

- Submitting every URL in sitemaps, including noindexed or redirected ones.
- Using robots.txt to "noindex".
- Launching 100,000 pages at once.
- Fake `lastmod` updates.

## How to measure success

Indexed share per page type trending up, "Discovered/Crawled – currently not indexed" shrinking for qualified pages, healthy crawl stats (low error rates, stable response time), and indexed pages earning impressions.

## Video lecture: Indexing management: sitemaps, crawl budget and noindex strategy

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Indexing management
2. Why it matters
3. Indexing = quality signal
4. Crawl budget
5. Sitemaps
6. Simple example: Karachi furniture store (illustrative)
7. Right tool, right goal
8. Never combine
9. Staged rollout
10. Related tools and example
11. Mistakes + try this now
12. Quick self-check
13. Watch me do it: staged rollout (illustrative)
14. Recap and next step

## Lecture transcript

### Indexing management

Here's the moment of truth for every programmatic project. You publish ten thousand pages, open Search Console a month later, and see most of them sitting in discovered, currently not indexed. Publishing isn't indexing. In this lecture you'll learn why indexing is your first real quality signal, when crawl budget actually matters, how to structure sitemaps, which tool to use among noindex, robots, canonical and redirects, and how to roll out in stages.

### Why it matters

Why does this matter? Because Google's indexing decisions are the most honest feedback you'll get about a page system. Here's an analogy. Publishing pages is like submitting manuscripts to a library. The library doesn't shelve everything it receives. It shelves what it thinks readers will want. If most of your manuscripts come back unshelved, the answer isn't to send them again louder. It's to improve the manuscripts, or send fewer, better ones.

### Indexing = quality signal

Search engines don't index everything they find. They decide based on value, duplication, site quality and resources. So when Google declines to index a big share of a page type, showing crawled, currently not indexed, or discovered, currently not indexed, it's telling you something about value or discovery. Treat the indexed share per page type as your first quality metric.

### Crawl budget

When does crawl budget matter? Google says mainly for very large sites, on the order of a million or more pages, for medium-large sites whose content changes very often, and when lots of pages are stuck as discovered but not indexed. Crawl capacity depends on your server's health and speed. Crawl demand depends on popularity and freshness. You help by removing low-value URLs from crawl paths, keeping servers fast and error-free, linking clearly to important pages, and keeping sitemaps accurate.

### Sitemaps

Sitemaps for page systems. Split them by page type, and by segment for large types, so you can see indexing per type in Search Console. Respect the limits: fifty thousand URLs or fifty megabytes uncompressed per file, with a sitemap index for more. Include only canonical, indexable URLs that return a normal success status. Use last modified accurately, because Google uses it when it's consistently reliable, and ignores change frequency and priority.

### Simple example: Karachi furniture store (illustrative)

Here's a simple worked example of choosing the right directive. A Karachi furniture store has three kinds of problem URLs. Filter combinations like sofas sorted by price with color gray: these don't need crawling, so they're disallowed in robots dot t x t for those parameters. Product pages for discontinued items with no replacement: these return four ten, and are removed from sitemaps and internal links. And category pages for areas with no stock yet: these stay live for users with a notify-me module, but carry noindex, and are left crawlable so Google can see the noindex. Three problems, three different tools.

### Right tool, right goal

Now the right tool for each goal. To keep a page out of the index but crawlable, use a noindex meta tag or header. To stop crawling of URL spaces like facets, use robots dot t x t disallow, knowing that blocked URLs can still appear in the index without content if linked. To consolidate duplicates, use canonical, which is a hint. To remove permanently, return four oh four or four ten. To move or merge, use a three oh one redirect.

### Never combine

One combination to never use: disallowing a URL in robots dot t x t and also putting noindex on it. If Google can't crawl the page, it can't see the noindex. So the page may stay in the index. Pick one approach per URL, based on the goal.

### Staged rollout

Roll out in stages. Pilot fifty to two hundred pages of one type in their own sitemap. Observe for four to eight weeks: indexing rate, impressions, clicks, engagement. Improve the template and thresholds based on what was and wasn't indexed. Expand in waves, keeping per-type sitemaps. And hold back or noindex segments that underperform. Monitor with Search Console's page indexing and crawl stats reports, URL inspection samples, and your server logs.

### Related tools and example

A few related tools. IndexNow lets you notify Bing, Yandex and other participating engines of changes instantly, though Google doesn't use it. And Google's Indexing API is only for pages with job posting or livestream broadcast event markup, not general pages. Here's an illustrative example: a UK property data site launched a hundred and twenty thousand street pages, most stuck as discovered but not indexed. They set a threshold of five sales in five years, merged thinner streets into district pages, sped up rendering, added hub links and per-type sitemaps. The indexed share climbed over the following months.

### Mistakes + try this now

Common indexing mistakes. Submitting noindexed, redirected or error pages in sitemaps. Using robots dot t x t to try to remove pages from the index. Combining a robots block with noindex. Launching a hundred thousand pages at once. And faking last modified dates. Try this now: open Search Console's page indexing report, filter by one of your sitemaps, and write down the top two reasons pages aren't indexed. Then decide for each reason whether the fix is quality, discovery or duplication.

### Quick self-check

Quick self-check. You've noindexed two thousand thin pages, but they're still listed in your XML sitemap. Is that a problem? Pause. Yes, a small but real one. Your sitemap is telling Google these URLs are important and canonical, while the pages themselves say don't index me. That mixed signal wastes crawling and muddies your indexing reports. Keep sitemaps to canonical, indexable pages only, and generate them from the same rules that decide noindex.

### Watch me do it: staged rollout (illustrative)

Watch me do it. Let's plan the rollout for an illustrative UK tradesperson marketplace launching plumber, electrician and roofer in town pages. Step one, sitemaps: one sitemap per trade, each generated from the same rules that set indexability, so only pages meeting the threshold appear. Step two, pilot: plumbers only, in the one hundred towns with the most verified tradespeople, about one hundred pages. Step three, observation: six weeks. I set targets before launch: at least seventy percent indexed, and impressions for at least half of the indexed pages. Step four, week three check: the plumbers sitemap shows sixty-two percent indexed; the rest are mostly crawled, currently not indexed. I compare them with indexed ones: the unindexed pages have fewer reviews and shorter lists. Step five, I raise the threshold slightly and add a computed insight module, average response time. Step six, week six: seventy-eight percent indexed. Step seven, expansion: electricians in the same hundred towns, then more towns in waves. Each wave has its own sitemap segment, so if something goes wrong, we can see exactly where.

### Recap and next step

Recap. Publishing isn't indexing, and indexing is your first quality signal. Crawl budget matters for very large or fast-changing sites. Split sitemaps by type with honest dates. Use each directive for its purpose, and never combine robots blocking with noindex. Roll out in stages. Your next step: design your sitemap structure and rollout plan, including pilot size, observation window, thresholds and the metrics that trigger expansion.

## Key takeaways

- Indexing is the first quality signal for page systems — track it per page type.
- Crawl budget matters mainly for very large or fast-changing sites; improve it by removing low-value URLs and speeding servers.
- Split sitemaps by type, include only canonical indexable URLs, use honest lastmod.
- Use noindex, robots.txt, canonical, 404/410 and 301 for their distinct purposes; don't block noindexed pages in robots.txt.
- Roll out in stages and monitor with Search Console, crawl stats and server logs.

## Try it

Design the sitemap structure and rollout plan for your page system: sitemap files per type, pilot size, observation window, thresholds and the metrics that trigger expansion.

- [Previous: Structured data at scale: JSON-LD generation and validation](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/structured-data-at-scale)
- [Next: AI-assisted content workflows with grounding and human review](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/ai-content-workflows-at-scale)
- [All lessons of Programmatic SEO and AI Content at Scale — Without Getting Penalized](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale)
