Technical SEO MasteryHow search engines crawl, render and index · Lesson 1 of 18
The crawl, render and index pipeline
Video lecture
The crawl, render and index pipeline
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The crawl, render and index pipeline
When a page doesn't rank, most people ask, what keywords are we missing? An expert asks a different question first. Which stage did this URL fail at? Every technical SEO problem is a failure somewhere in a pipeline: discovery, crawling, rendering, indexing and serving. In this lecture you'll learn that pipeline properly, using Google's documented behaviour, and you'll leave with a diagnostic routine, and a small Python script, that tells you exactly where a URL is getting stuck.
0:34 The parcel analogy
Think of it like a parcel being delivered. First, the courier has to know the address exists. That's discovery. Then they have to be allowed through the gate and pick it up. That's crawling. Then they unpack it and assemble it. That's rendering. Then someone decides whether it's worth putting on the shelf. That's indexing. And finally, a customer has to choose it from the shelf. That's serving, or ranking. If the parcel never left the depot, arguing about the shelf display is pointless.
1:11 Stage 1: Discovery
Stage one, discovery. Search engines learn about URLs from links on pages they've already crawled, which is still the main route, from XML sitemaps, from redirect and canonical targets, and, for engines that support it, from IndexNow. Bing, Yandex, Seznam, Naver and Yep use IndexNow; Google doesn't. A page with no internal links and no sitemap entry is an orphan, and orphans are discovered slowly or not at all.
1:41 Stage 2: Crawling
Stage two, crawling. Googlebot checks robots.txt, requests the URL, and records the status code, headers and raw HTML. Three facts matter. robots.txt controls crawling, not indexing, so a blocked URL can still be indexed from links, without its content. Crawl rate adapts to your server: slow responses, five hundreds and four-twenty-nines make Googlebot back off. And Google documents file size limits. Googlebot fetches the first two megabytes of an HTML file and up to sixty-four megabytes of a PDF. Anything after the cut-off is ignored, so keep critical content and links early.
2:21 Stage 3: Rendering
Stage three, rendering. Google's Web Rendering Service runs an up-to-date version of Chromium to execute JavaScript and build the rendered page. It doesn't click, scroll like a user or fill in forms, and it doesn't keep cookies or local storage between loads. It won't run scripts blocked by robots.txt. And in December twenty twenty-five Google clarified two things: pages with a non-two-hundred status may not be rendered at all, and a noindex in the initial HTML may cause rendering to be skipped. So never rely on JavaScript to remove a noindex.
3:01 Stage 4: Indexing
Stage four, indexing. Google analyses the content, clusters duplicates, chooses a canonical for each cluster, and decides whether to store the page at all. Indexing isn't guaranteed. Learn the Page indexing statuses. Discovered, currently not indexed, means known but not yet crawled. Crawled, currently not indexed, usually means quality or duplication concerns. Duplicate, Google chose a different canonical than the user, means your signals disagree with Google's choice. And soft four-oh-four means a page that returns two hundred but looks empty or like an error.
3:38 Stage 5 + the routine
Stage five, serving. Only indexed canonical URLs can rank. Then ranking systems weigh relevance, quality, usefulness and many other signals. Technical SEO's job is to make sure the right URLs reach this stage with their full content. It doesn't guarantee rankings on its own. Now, the diagnostic routine. When a URL underperforms, walk the pipeline in order: is it discoverable, is crawling allowed, does it return a clean two hundred, does the rendered HTML contain the content and links, are the index directives right, what does URL Inspection say, and only then, look at content and competition.
4:20 Hands-on: URL Inspection API
Hands-on. The lesson text includes a Python script using the Search Console URL Inspection API. You give it a list of URLs from a property you own, and it records the verdict, coverage state, robots state, last crawl time, which crawler was used, your declared canonical, and Google's chosen canonical. Keep an eye on the quotas; at the time of writing they're documented as two thousand calls per day per property. Then sort the output. Patterns jump out: a template where Google's canonical differs from yours, or a batch of pages never crawled.
5:00 Example 1: the orphaned Leeds page
Worked example one, simple. A plumber's website in Leeds has a new page for boiler repairs in Headingley that doesn't show up anywhere. Walk the pipeline. Discovery: it isn't linked from anywhere and isn't in the sitemap. That's it. The page is an orphan. Add it to the services hub and the sitemap, and within days URL Inspection shows it crawled and indexed. No keyword research needed. The failure was at stage one.
5:32 Example 2: Riyadh publisher (illustrative)
Worked example two, with illustrative details. A Riyadh news publisher says new Arabic articles don't appear in Google. Walk fifty recent URLs. They're in the news sitemap. robots.txt allows them. They return two hundred quickly. But URL Inspection shows the rendered HTML has no article body. The body is injected by a script loaded from a subdomain, and that subdomain's robots.txt blocks it. Rendering is the failed stage. The fix is allowing the script host or server-rendering the body. Within days, verdicts change to indexed, and the script host goes into the robots.txt regression tests.
6:13 Watch me do it: one URL through the pipeline
Watch me do it. I'll walk one underperforming URL through the pipeline live. The URL is a new service page that gets no impressions. Discovery first: I open the XML sitemap and search for the URL. It's there. Then I check my crawler's inlinks report. Zero internal links from indexable pages. That's already suspicious, but I keep going. Crawl permission: in Search Console's URL Inspection, I test the live URL. Crawl allowed: yes. Response: I run curl dash capital I with a Googlebot smartphone user agent. Two hundred, no redirect, fast. Render: in the live test, I click View tested page and open the HTML tab, then search for a unique sentence from the page body. It's there, and so are the links. Index directives: I search the rendered HTML for noindex and check the response headers for X-Robots-Tag. Clean. Canonical: self-referencing. Index status: URL Inspection says discovered, currently not indexed. Put that together with zero internal links, and the failed stage is discovery and prioritisation: Google knows the URL but has little reason to crawl it soon. The fix is linking it from the services hub and two relevant guides, then re-checking in two weeks.
7:38 Mistakes and measures
Common mistakes. Blaming the algorithm for pages that were never indexed. Using robots.txt to remove pages from the index, which stops Google seeing the noindex. Assuming what you see in your browser is what Googlebot sees, when cookie banners, geo-redirects and A-B tests can change the response. And treating crawled, currently not indexed as a technical bug when it's usually a quality or duplication signal. How do you measure success? Track the share of priority URLs indexed with Google's canonical matching yours, and the days from publishing to first crawl.
8:17 Recap and try this now
Recap. Discovery, crawling, rendering, indexing, serving. Find the failed stage before you touch content. Remember the two-megabyte HTML limit, that robots.txt doesn't control indexing, and that non-two-hundred pages or an initial noindex may skip rendering. Try this now. Take five URLs that underperform, run them through the seven-step routine or the inspection script, and write down the stage each one fails at. You'll be surprised how often it isn't ranking at all.
Why the pipeline matters
Every technical SEO problem you will ever debug is a failure at one stage of a pipeline: discovery → crawling → rendering → indexing → serving. When a page "doesn't rank", the first question an expert asks is not "what keywords?" but "which stage did this URL fail at?". This lesson gives you the mental model, using Google's documented behaviour as the reference, because Google publishes the most detailed guidance and other engines (Bing, and the crawlers feeding AI answer engines) broadly follow the same shape.
Stage 1: Discovery
A search engine can only crawl URLs it knows about. It learns about them from:
- Links from pages it has already crawled (internal and external) — still the dominant discovery route.
- XML sitemaps submitted in Search Console or referenced in robots.txt.
- Redirect targets and canonical targets it encounters.
- Other feeds, such as IndexNow for engines that support it (Bing, Yandex and others; Google does not use IndexNow at the time of writing).
A URL that is not linked from anywhere and not in a sitemap is an orphan. Orphans are often discovered slowly or not at all.
Stage 2: Crawling
Googlebot takes URLs from its crawl queue, checks robots.txt for permission, and requests the URL. It records the HTTP status code, response headers and the raw HTML. Key facts:
- robots.txt controls crawling, not indexing. A disallowed URL can still be indexed (without its content) if other pages link to it.
- Crawl rate adapts to your server. If responses slow down or return 5xx/429 errors, Googlebot backs off.
- Google crawls primarily with its smartphone user agent (mobile-first indexing).
- Google documents file size limits: Googlebot fetches the first 2 MB of an HTML (or other supported text) file, including headers, and up to 64 MB of a PDF; other Google crawlers default to 15 MB unless documented otherwise. Content after the cut-off is ignored, so keep critical content and links early and HTML lean.
Stage 3: Rendering
Modern pages depend on JavaScript. Google's Web Rendering Service (WRS) runs an evergreen version of Chromium to execute scripts and build the rendered DOM. Pages that return a 200 status are queued for rendering; the rendered HTML is then parsed again for links and content. Google clarified in December 2025 that pages with non-200 status codes may not be rendered at all, and that a noindex in the initial HTML may cause rendering to be skipped.
What rendering does not do:
- It does not click, scroll like a user, or fill forms. Content that requires interaction (a "Load more" button without a real link, a tab loaded only on click) may never be seen.
- It does not keep state between page loads — no cookies, localStorage or session persistence.
- It will not execute resources blocked by robots.txt. Blocking
/static/js/can make your page render blank for Google.
Stage 4: Indexing
After rendering, Google analyses the content, clusters duplicates, chooses a canonical URL for each cluster, and decides whether the page is worth storing. Indexing is not guaranteed. The Page indexing report in Search Console will show statuses such as:
| Status | What it usually means |
|---|---|
| Discovered – currently not indexed | Known URL, not yet crawled (often crawl demand/capacity or low perceived value) |
| Crawled – currently not indexed | Crawled, judged not worth indexing right now (often quality or duplication) |
| Duplicate, Google chose different canonical than user | Your canonical signals disagree with Google's choice |
| Excluded by 'noindex' tag | A noindex directive was found — check it is intended |
| Soft 404 | Returns 200 but looks like an error or empty page |
Stage 5: Serving (ranking)
Only indexed canonical URLs are eligible to be served. Ranking systems then evaluate relevance, quality, usefulness and many other signals. Technical SEO's job is to make sure the right URLs reach this stage with their full content — it does not, by itself, guarantee rankings.
A diagnostic routine you can use today
When a URL underperforms, walk the pipeline in order:
- Discovery — Is it in the sitemap? Is it linked internally from an indexable page? (Check a crawler's inlinks report.)
- Crawl permission — Test the URL against robots.txt (Search Console's robots.txt report shows the fetched file; a crawler can test rules).
- Response — Does it return 200 directly, not via a redirect chain? Use
curl -I:
curl -sI -A "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://example.com/page- Render — Use URL Inspection → Test live URL → View tested page to see the rendered HTML and screenshot. Is the main content and are the links present?
- Index directives — Look for
noindexin the rendered HTML and theX-Robots-Tagheader. Check the canonical in both. - Index status — Read the URL Inspection verdict and the Google-selected canonical.
- Serving — Only now look at content quality, intent match and competition.
Hands-on: walk the pipeline for 50 URLs with the URL Inspection API
The Search Console URL Inspection API returns the same verdicts as the interface, for URLs in properties you own. It has per-property quotas (documented as 2,000 calls per day and 600 per minute at the time of writing), so sample wisely.
# pip install google-api-python-client google-auth
import csv, os
from google.oauth2 import service_account
from googleapiclient.discovery import build
from googleapiclient.errors import HttpError
SITE = "sc-domain:example.com" # Domain property; the service account must be added as a user
creds = service_account.Credentials.from_service_account_file(
os.environ["GSC_KEY_FILE"], scopes=["https://www.googleapis.com/auth/webmasters.readonly"])
gsc = build("searchconsole", "v1", credentials=creds, cache_discovery=False)
with open("urls.txt") as f, open("inspection.csv", "w", newline="") as out:
w = csv.writer(out)
w.writerow(["url", "verdict", "coverage", "robots", "indexing", "last_crawl",
"crawled_as", "user_canonical", "google_canonical"])
for url in (u.strip() for u in f if u.strip()):
try:
r = gsc.urlInspection().index().inspect(
body={"inspectionUrl": url, "siteUrl": SITE}).execute()
except HttpError as e:
w.writerow([url, "ERROR", e.status_code]); continue
s = r["inspectionResult"]["indexStatusResult"]
w.writerow([url, s.get("verdict"), s.get("coverageState"), s.get("robotsTxtState"),
s.get("indexingState"), s.get("lastCrawlTime"), s.get("crawledAs"),
s.get("userCanonical"), s.get("googleCanonical")])Sort the output by coverage and look for patterns: a template where googleCanonical differs from userCanonical, pages never crawled, or robotsTxtState blocking something important. Each pattern maps to one pipeline stage.
Worked example 2: a Riyadh news site's "missing" articles
A Riyadh news publisher (illustrative) says new Arabic articles "don't appear in Google". Walking the pipeline for 50 recent URLs: all are in the news sitemap (discovery fine); robots.txt allows them; each returns 200 in under 300 ms; but URL Inspection shows the rendered HTML lacks the article body, because the body is injected by a script loaded from a subdomain blocked in that subdomain's robots.txt. Rendering is the failed stage. The fix: allow the script host (or server-render the body). Within days the inspection verdicts change to indexed — and the team adds the script host to its robots.txt regression tests.
Measuring success
For a sample of priority URLs, track the share that are indexed with Google-selected canonical equal to the declared canonical, the median days from publication to first crawl (from logs or lastCrawlTime), and the count of URLs per Page indexing status by sitemap segment.
Common mistakes
- Blaming "the algorithm" for pages that were never indexed.
- Using robots.txt to "remove" pages from the index (it prevents Google from seeing the noindex).
- Assuming what you see in your browser is what Googlebot sees — cookie banners, geo-redirects and A/B tests can all change the response.
- Treating "Crawled – currently not indexed" as a technical bug when it is usually a quality or duplication signal.
Key takeaways
- Debug by stage: discovery, crawl, render, index, serve — in that order.
- robots.txt controls crawling, not indexing; noindex controls indexing but must be crawlable to be seen.
- Rendering does not click, scroll or keep state, and cannot run resources blocked by robots.txt.
- Indexing is selective; Search Console's Page indexing statuses tell you which stage failed.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick five important URLs on a site you manage and run them through the seven-step diagnostic routine. Record which stage, if any, each one fails.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.