Programmatic SEO and AI Content at Scale — Without Getting PenalizedIndexing and crawl management at scale · Lesson 8 of 14
Indexing management: sitemaps, crawl budget and noindex strategy
Video lecture
Indexing management: sitemaps, crawl budget and noindex strategy
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Indexing management
Here's the moment of truth for every programmatic project. You publish ten thousand pages, open Search Console a month later, and see most of them sitting in discovered, currently not indexed. Publishing isn't indexing. In this lecture you'll learn why indexing is your first real quality signal, when crawl budget actually matters, how to structure sitemaps, which tool to use among noindex, robots, canonical and redirects, and how to roll out in stages.
0:32 Why it matters
Why does this matter? Because Google's indexing decisions are the most honest feedback you'll get about a page system. Here's an analogy. Publishing pages is like submitting manuscripts to a library. The library doesn't shelve everything it receives. It shelves what it thinks readers will want. If most of your manuscripts come back unshelved, the answer isn't to send them again louder. It's to improve the manuscripts, or send fewer, better ones.
1:03 Indexing = quality signal
Search engines don't index everything they find. They decide based on value, duplication, site quality and resources. So when Google declines to index a big share of a page type, showing crawled, currently not indexed, or discovered, currently not indexed, it's telling you something about value or discovery. Treat the indexed share per page type as your first quality metric.
1:29 Crawl budget
When does crawl budget matter? Google says mainly for very large sites, on the order of a million or more pages, for medium-large sites whose content changes very often, and when lots of pages are stuck as discovered but not indexed. Crawl capacity depends on your server's health and speed. Crawl demand depends on popularity and freshness. You help by removing low-value URLs from crawl paths, keeping servers fast and error-free, linking clearly to important pages, and keeping sitemaps accurate.
2:04 Sitemaps
Sitemaps for page systems. Split them by page type, and by segment for large types, so you can see indexing per type in Search Console. Respect the limits: fifty thousand URLs or fifty megabytes uncompressed per file, with a sitemap index for more. Include only canonical, indexable URLs that return a normal success status. Use last modified accurately, because Google uses it when it's consistently reliable, and ignores change frequency and priority.
2:35 Simple example: Karachi furniture store (illustrative)
Here's a simple worked example of choosing the right directive. A Karachi furniture store has three kinds of problem URLs. Filter combinations like sofas sorted by price with color gray: these don't need crawling, so they're disallowed in robots dot t x t for those parameters. Product pages for discontinued items with no replacement: these return four ten, and are removed from sitemaps and internal links. And category pages for areas with no stock yet: these stay live for users with a notify-me module, but carry noindex, and are left crawlable so Google can see the noindex. Three problems, three different tools.
3:19 Right tool, right goal
Now the right tool for each goal. To keep a page out of the index but crawlable, use a noindex meta tag or header. To stop crawling of URL spaces like facets, use robots dot t x t disallow, knowing that blocked URLs can still appear in the index without content if linked. To consolidate duplicates, use canonical, which is a hint. To remove permanently, return four oh four or four ten. To move or merge, use a three oh one redirect.
3:55 Never combine
One combination to never use: disallowing a URL in robots dot t x t and also putting noindex on it. If Google can't crawl the page, it can't see the noindex. So the page may stay in the index. Pick one approach per URL, based on the goal.
4:16 Staged rollout
Roll out in stages. Pilot fifty to two hundred pages of one type in their own sitemap. Observe for four to eight weeks: indexing rate, impressions, clicks, engagement. Improve the template and thresholds based on what was and wasn't indexed. Expand in waves, keeping per-type sitemaps. And hold back or noindex segments that underperform. Monitor with Search Console's page indexing and crawl stats reports, URL inspection samples, and your server logs.
4:47 Related tools and example
A few related tools. IndexNow lets you notify Bing, Yandex and other participating engines of changes instantly, though Google doesn't use it. And Google's Indexing API is only for pages with job posting or livestream broadcast event markup, not general pages. Here's an illustrative example: a UK property data site launched a hundred and twenty thousand street pages, most stuck as discovered but not indexed. They set a threshold of five sales in five years, merged thinner streets into district pages, sped up rendering, added hub links and per-type sitemaps. The indexed share climbed over the following months.
5:30 Mistakes + try this now
Common indexing mistakes. Submitting noindexed, redirected or error pages in sitemaps. Using robots dot t x t to try to remove pages from the index. Combining a robots block with noindex. Launching a hundred thousand pages at once. And faking last modified dates. Try this now: open Search Console's page indexing report, filter by one of your sitemaps, and write down the top two reasons pages aren't indexed. Then decide for each reason whether the fix is quality, discovery or duplication.
6:05 Quick self-check
Quick self-check. You've noindexed two thousand thin pages, but they're still listed in your XML sitemap. Is that a problem? Pause. Yes, a small but real one. Your sitemap is telling Google these URLs are important and canonical, while the pages themselves say don't index me. That mixed signal wastes crawling and muddies your indexing reports. Keep sitemaps to canonical, indexable pages only, and generate them from the same rules that decide noindex.
6:37 Watch me do it: staged rollout (illustrative)
Watch me do it. Let's plan the rollout for an illustrative UK tradesperson marketplace launching plumber, electrician and roofer in town pages. Step one, sitemaps: one sitemap per trade, each generated from the same rules that set indexability, so only pages meeting the threshold appear. Step two, pilot: plumbers only, in the one hundred towns with the most verified tradespeople, about one hundred pages. Step three, observation: six weeks. I set targets before launch: at least seventy percent indexed, and impressions for at least half of the indexed pages. Step four, week three check: the plumbers sitemap shows sixty-two percent indexed; the rest are mostly crawled, currently not indexed. I compare them with indexed ones: the unindexed pages have fewer reviews and shorter lists. Step five, I raise the threshold slightly and add a computed insight module, average response time. Step six, week six: seventy-eight percent indexed. Step seven, expansion: electricians in the same hundred towns, then more towns in waves. Each wave has its own sitemap segment, so if something goes wrong, we can see exactly where.
7:54 Recap and next step
Recap. Publishing isn't indexing, and indexing is your first quality signal. Crawl budget matters for very large or fast-changing sites. Split sitemaps by type with honest dates. Use each directive for its purpose, and never combine robots blocking with noindex. Roll out in stages. Your next step: design your sitemap structure and rollout plan, including pilot size, observation window, thresholds and the metrics that trigger expansion.
Publishing is not indexing
Search engines do not index every page they find. They decide based on perceived value, duplication, site quality and resources. For page systems, indexing is the first real quality signal: if Google declines to index a large share of a page type ("Crawled – currently not indexed", "Discovered – currently not indexed"), it is telling you something about value or discovery.
Crawl budget: when it matters
Google's guidance says crawl budget is mainly a concern for very large sites (on the order of a million or more unique pages) or medium-large sites with content changing very frequently, and for sites with many "Discovered – currently not indexed" URLs. Crawl capacity depends on your server's health and responsiveness; crawl demand depends on popularity and freshness. You influence it by:
- Removing low-value URLs from crawl paths (facets, duplicates, infinite spaces).
- Fast, reliable servers (errors and slow responses reduce crawl rate).
- Clear internal linking to important pages.
- Accurate sitemaps with truthful
lastmod.
XML sitemaps for page systems
- Split sitemaps by page type (and by segment for large types) so you can monitor indexing per type in Search Console.
- Respect limits: 50,000 URLs or 50 MB uncompressed per sitemap; use a sitemap index for more.
- Include only canonical, indexable, 200-status URLs.
- Use
lastmodaccurately (Google uses it when it is consistently reliable); Google ignoreschangefreqandpriority.
from datetime import timezone
from xml.sax.saxutils import escape
def write_sitemap(path, rows): # rows: iterable of (url, last_modified_datetime)
with open(path, "w", encoding="utf-8") as f:
f.write('<?xml version="1.0" encoding="UTF-8"?>\n')
f.write('<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">\n')
n = 0
for url, lm in rows:
f.write(f" <url><loc>{escape(url)}</loc><lastmod>{lm.astimezone(timezone.utc).strftime('%Y-%m-%dT%H:%M:%SZ')}</lastmod></url>\n")
n += 1
if n >= 50000:
raise ValueError("split sitemap: 50,000 URL limit")
f.write("</urlset>\n")noindex, robots.txt and canonical: use the right tool
| Goal | Tool | Why |
|---|---|---|
| Keep a page out of the index but let crawlers see it | <meta name="robots" content="noindex"> or X-Robots-Tag header | Crawlers must be able to fetch the page to see noindex |
| Stop crawling of URL spaces (facets, parameters) | robots.txt Disallow | Saves crawl; but disallowed URLs can still be indexed without content if linked |
| Consolidate duplicates to a preferred URL | rel="canonical" | A hint; combine with consistent internal links |
| Remove permanently | 404/410 | Clear signal; update links and sitemaps |
| Move/merge | 301 redirect | Passes signals to the target |
Don't combine robots.txt Disallow with noindex on the same URL — Google can't see the noindex if it can't crawl.
A staged rollout strategy
- Pilot: publish 50–200 pages of one type; include them in a dedicated sitemap.
- Observe 4–8 weeks: indexing rate, impressions, clicks, engagement.
- Improve the template and thresholds based on what was and wasn't indexed.
- Expand in waves, keeping per-type sitemaps and monitoring.
- Hold back or noindex segments that underperform.
Monitoring tools
- Search Console Page indexing report filtered by sitemap.
- URL Inspection (and its API, with daily quotas) for samples.
- Crawl stats report: response codes, response time, file types.
- Server logs: which bots crawl which page types, how often (verify Googlebot via reverse DNS).
- IndexNow: supported by Bing, Yandex and others for instant change notifications (Google does not use IndexNow).
- The Indexing API is only for pages with JobPosting or BroadcastEvent (livestream) structured data — not general pages.
Worked example: a UK property data site
The site generated 120,000 "house prices in [street]" pages. After launch, most sat in "Discovered – currently not indexed". Analysis: thin pages for streets with one or two sales, weak internal links, and a slow database-backed render. Fixes: threshold of at least five sales in 5 years (others merged into postcode-district pages), faster cached rendering, links from district hubs, per-type sitemaps. The indexed share rose over the following months and traffic concentrated on the stronger pages.
Pitfalls
- Submitting every URL in sitemaps, including noindexed or redirected ones.
- Using robots.txt to "noindex".
- Launching 100,000 pages at once.
- Fake
lastmodupdates.
How to measure success
Indexed share per page type trending up, "Discovered/Crawled – currently not indexed" shrinking for qualified pages, healthy crawl stats (low error rates, stable response time), and indexed pages earning impressions.
Key takeaways
- Indexing is the first quality signal for page systems — track it per page type.
- Crawl budget matters mainly for very large or fast-changing sites; improve it by removing low-value URLs and speeding servers.
- Split sitemaps by type, include only canonical indexable URLs, use honest lastmod.
- Use noindex, robots.txt, canonical, 404/410 and 301 for their distinct purposes; don't block noindexed pages in robots.txt.
- Roll out in stages and monitor with Search Console, crawl stats and server logs.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Design the sitemap structure and rollout plan for your page system: sitemap files per type, pilot size, observation window, thresholds and the metrics that trigger expansion.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.