Technical SEO MasteryHow search engines crawl, render and index · Lesson 1 of 18

The crawl, render and index pipeline

Article · 14 min · 9 min lecture

Video lecture

The crawl, render and index pipeline

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

The crawl, render and index pipeline

  • Five stages every URL passes through
  • Where URLs get stuck
  • A diagnostic routine you can run today

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why the pipeline matters

Every technical SEO problem you will ever debug is a failure at one stage of a pipeline: discovery → crawling → rendering → indexing → serving. When a page "doesn't rank", the first question an expert asks is not "what keywords?" but "which stage did this URL fail at?". This lesson gives you the mental model, using Google's documented behaviour as the reference, because Google publishes the most detailed guidance and other engines (Bing, and the crawlers feeding AI answer engines) broadly follow the same shape.

Stage 1: Discovery

A search engine can only crawl URLs it knows about. It learns about them from:

  • Links from pages it has already crawled (internal and external) — still the dominant discovery route.
  • XML sitemaps submitted in Search Console or referenced in robots.txt.
  • Redirect targets and canonical targets it encounters.
  • Other feeds, such as IndexNow for engines that support it (Bing, Yandex and others; Google does not use IndexNow at the time of writing).

A URL that is not linked from anywhere and not in a sitemap is an orphan. Orphans are often discovered slowly or not at all.

Stage 2: Crawling

Googlebot takes URLs from its crawl queue, checks robots.txt for permission, and requests the URL. It records the HTTP status code, response headers and the raw HTML. Key facts:

  • robots.txt controls crawling, not indexing. A disallowed URL can still be indexed (without its content) if other pages link to it.
  • Crawl rate adapts to your server. If responses slow down or return 5xx/429 errors, Googlebot backs off.
  • Google crawls primarily with its smartphone user agent (mobile-first indexing).
  • Google documents file size limits: Googlebot fetches the first 2 MB of an HTML (or other supported text) file, including headers, and up to 64 MB of a PDF; other Google crawlers default to 15 MB unless documented otherwise. Content after the cut-off is ignored, so keep critical content and links early and HTML lean.

Stage 3: Rendering

Modern pages depend on JavaScript. Google's Web Rendering Service (WRS) runs an evergreen version of Chromium to execute scripts and build the rendered DOM. Pages that return a 200 status are queued for rendering; the rendered HTML is then parsed again for links and content. Google clarified in December 2025 that pages with non-200 status codes may not be rendered at all, and that a noindex in the initial HTML may cause rendering to be skipped.

What rendering does not do:

  • It does not click, scroll like a user, or fill forms. Content that requires interaction (a "Load more" button without a real link, a tab loaded only on click) may never be seen.
  • It does not keep state between page loads — no cookies, localStorage or session persistence.
  • It will not execute resources blocked by robots.txt. Blocking /static/js/ can make your page render blank for Google.

Stage 4: Indexing

After rendering, Google analyses the content, clusters duplicates, chooses a canonical URL for each cluster, and decides whether the page is worth storing. Indexing is not guaranteed. The Page indexing report in Search Console will show statuses such as:

StatusWhat it usually means
Discovered – currently not indexedKnown URL, not yet crawled (often crawl demand/capacity or low perceived value)
Crawled – currently not indexedCrawled, judged not worth indexing right now (often quality or duplication)
Duplicate, Google chose different canonical than userYour canonical signals disagree with Google's choice
Excluded by 'noindex' tagA noindex directive was found — check it is intended
Soft 404Returns 200 but looks like an error or empty page

Stage 5: Serving (ranking)

Only indexed canonical URLs are eligible to be served. Ranking systems then evaluate relevance, quality, usefulness and many other signals. Technical SEO's job is to make sure the right URLs reach this stage with their full content — it does not, by itself, guarantee rankings.

A diagnostic routine you can use today

When a URL underperforms, walk the pipeline in order:

  1. Discovery — Is it in the sitemap? Is it linked internally from an indexable page? (Check a crawler's inlinks report.)
  2. Crawl permission — Test the URL against robots.txt (Search Console's robots.txt report shows the fetched file; a crawler can test rules).
  3. Response — Does it return 200 directly, not via a redirect chain? Use curl -I:
curl -sI -A "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://example.com/page
  1. Render — Use URL Inspection → Test live URL → View tested page to see the rendered HTML and screenshot. Is the main content and are the links present?
  2. Index directives — Look for noindex in the rendered HTML and the X-Robots-Tag header. Check the canonical in both.
  3. Index status — Read the URL Inspection verdict and the Google-selected canonical.
  4. Serving — Only now look at content quality, intent match and competition.

Hands-on: walk the pipeline for 50 URLs with the URL Inspection API

The Search Console URL Inspection API returns the same verdicts as the interface, for URLs in properties you own. It has per-property quotas (documented as 2,000 calls per day and 600 per minute at the time of writing), so sample wisely.

# pip install google-api-python-client google-auth
import csv, os
from google.oauth2 import service_account
from googleapiclient.discovery import build
from googleapiclient.errors import HttpError

SITE = "sc-domain:example.com"          # Domain property; the service account must be added as a user
creds = service_account.Credentials.from_service_account_file(
    os.environ["GSC_KEY_FILE"], scopes=["https://www.googleapis.com/auth/webmasters.readonly"])
gsc = build("searchconsole", "v1", credentials=creds, cache_discovery=False)

with open("urls.txt") as f, open("inspection.csv", "w", newline="") as out:
    w = csv.writer(out)
    w.writerow(["url", "verdict", "coverage", "robots", "indexing", "last_crawl",
                "crawled_as", "user_canonical", "google_canonical"])
    for url in (u.strip() for u in f if u.strip()):
        try:
            r = gsc.urlInspection().index().inspect(
                body={"inspectionUrl": url, "siteUrl": SITE}).execute()
        except HttpError as e:
            w.writerow([url, "ERROR", e.status_code]); continue
        s = r["inspectionResult"]["indexStatusResult"]
        w.writerow([url, s.get("verdict"), s.get("coverageState"), s.get("robotsTxtState"),
                    s.get("indexingState"), s.get("lastCrawlTime"), s.get("crawledAs"),
                    s.get("userCanonical"), s.get("googleCanonical")])

Sort the output by coverage and look for patterns: a template where googleCanonical differs from userCanonical, pages never crawled, or robotsTxtState blocking something important. Each pattern maps to one pipeline stage.

Worked example 2: a Riyadh news site's "missing" articles

A Riyadh news publisher (illustrative) says new Arabic articles "don't appear in Google". Walking the pipeline for 50 recent URLs: all are in the news sitemap (discovery fine); robots.txt allows them; each returns 200 in under 300 ms; but URL Inspection shows the rendered HTML lacks the article body, because the body is injected by a script loaded from a subdomain blocked in that subdomain's robots.txt. Rendering is the failed stage. The fix: allow the script host (or server-render the body). Within days the inspection verdicts change to indexed — and the team adds the script host to its robots.txt regression tests.

Measuring success

For a sample of priority URLs, track the share that are indexed with Google-selected canonical equal to the declared canonical, the median days from publication to first crawl (from logs or lastCrawlTime), and the count of URLs per Page indexing status by sitemap segment.

Common mistakes

  • Blaming "the algorithm" for pages that were never indexed.
  • Using robots.txt to "remove" pages from the index (it prevents Google from seeing the noindex).
  • Assuming what you see in your browser is what Googlebot sees — cookie banners, geo-redirects and A/B tests can all change the response.
  • Treating "Crawled – currently not indexed" as a technical bug when it is usually a quality or duplication signal.

Key takeaways

  • Debug by stage: discovery, crawl, render, index, serve — in that order.
  • robots.txt controls crawling, not indexing; noindex controls indexing but must be crawlable to be seen.
  • Rendering does not click, scroll or keep state, and cannot run resources blocked by robots.txt.
  • Indexing is selective; Search Console's Page indexing statuses tell you which stage failed.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A URL is disallowed in robots.txt but still appears in Google results with no description. Why?
  2. Which statement about Google's rendering is correct?
  3. 'Crawled – currently not indexed' most often points to…

Put it into practice

Pick five important URLs on a site you manage and run them through the seven-step diagnostic routine. Record which stage, if any, each one fails.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.