---
title: "Crawling, indexing, sitemaps and structured data"
description: "How search engines process your site 1. Crawling: bots such as Googlebot and Bingbot discover URLs by following links and reading sitemaps. 2. Rendering…"
url: https://optimizeall.com/learn/seo-and-content-strategy/crawling-indexing-structured-data
updated: 2026-10-05
---

SEO & Content Strategy · Technical SEO basics · lesson 6 of 17 · 12 min

# Crawling, indexing, sitemaps and structured data

## How search engines process your site

1. **Crawling:** bots such as Googlebot and Bingbot discover URLs by following links and reading sitemaps.
2. **Rendering:** they load the page, including JavaScript, to see the content as a user would.
3. **Indexing:** they analyse the content and store eligible pages in the index.
4. **Ranking/serving:** when someone searches, they select and order indexed pages.

A problem at any step stops a page from appearing, no matter how good the content is.

## robots.txt

`robots.txt` sits at the root (`example.com/robots.txt`) and tells crawlers which paths they may crawl. Example:

```
User-agent: *
Disallow: /cart/
Disallow: /account/
Sitemap: https://example.com/sitemap.xml
```

Key facts:

- robots.txt controls **crawling**, not indexing. A blocked URL can still appear in results (usually without a description) if other pages link to it.
- To keep a page out of the index, allow crawling and use a **noindex** meta robots tag or HTTP header instead.
- Never block CSS or JavaScript files needed to render pages.
- A single mistake (such as `Disallow: /` left over from a staging site) can block your whole site.

## XML sitemaps

A sitemap lists the canonical URLs you want indexed, optionally with last-modified dates.

- Include only indexable, canonical, 200-status URLs.
- Keep it updated automatically (most CMSs and e-commerce platforms do this).
- Submit it in Google Search Console and Bing Webmaster Tools, and reference it in robots.txt.
- Large sites can use a sitemap index that links to multiple sitemaps.

## Canonical tags and duplicate content

Duplicate or near-duplicate URLs are common: filter parameters (`?colour=red`), tracking parameters, HTTP vs HTTPS, www vs non-www, printer versions. The **canonical tag** (`rel="canonical"`) tells search engines which version is preferred. Combine it with consistent internal linking and redirects (301) from old or alternate URLs to the preferred one.

## Status codes to know

| Code | Meaning | SEO implication |
|---|---|---|
| 200 | OK | Page can be indexed |
| 301 | Permanent redirect | Passes signals to the new URL; use for moved pages |
| 302 | Temporary redirect | Use only for genuinely temporary moves |
| 404 | Not found | Fine for truly removed pages; fix internal links pointing to them |
| 410 | Gone | Signals intentional permanent removal |
| 5xx | Server error | Frequent errors can reduce crawling and cause pages to drop |

## International sites

For sites serving multiple languages or countries (for example English and Arabic versions for the UAE, or UK and US versions), **hreflang** annotations help search engines show the right version to each audience. Each version should reference the others and itself.

## Structured data

Structured data (usually **JSON-LD** using the schema.org vocabulary) describes page content in a machine-readable way. It can make pages eligible for **rich results** such as product prices and ratings, breadcrumbs, events, recipes, videos, local business details and organisation information.

Example (simplified) for a local business:

```json
{
  "@context": "https://schema.org",
  "@type": "LocalBusiness",
  "name": "Studio Noor Interiors",
  "address": {"@type": "PostalAddress", "addressLocality": "Dubai", "addressCountry": "AE"},
  "telephone": "+971-4-000-0000",
  "url": "https://example.com"
}
```

Rules:

- Mark up only content visible on the page, and keep it accurate.
- Use Google's Rich Results Test and the Schema Markup Validator to check.
- Rich results are never guaranteed, and eligibility changes (Google has reduced or retired some rich result types over time, such as limiting FAQ rich results to certain authoritative sites). Check the current documentation.
- Misleading markup (for example fake review stars) can lead to manual actions.

## Worked example

A Karachi e-commerce site notices many product pages are "Discovered – currently not indexed". Investigation shows thousands of filter URLs being crawled, wasting crawl capacity, and missing canonicals. Fixes: canonical tags on filter pages, blocking crawl of infinite filter combinations in robots.txt where appropriate, cleaner internal links and an accurate sitemap of product URLs. Over the following weeks, more product pages are indexed.

## AI crawlers and the index behind AI answers

Google's AI Overviews and AI Mode draw on Google's normal search index, crawled by Googlebot. That has two practical consequences for content teams:

- If a page cannot be crawled or indexed, it cannot appear in classic results **or** be used as a source in Google's AI features.
- The `Google-Extended` robots.txt token controls whether content is used for training future Gemini models and for grounding in some Gemini products — **blocking it does not remove pages from AI Overviews or AI Mode**. Google's documented controls for Search AI features are snippet controls (`nosnippet`, `max-snippet`), which also affect classic snippets, and — since 2026 — a **Search generative AI control** in Search Console that lets a site opt out of AI features in Search and Discover without affecting normal results. Check the current Search Console help before using it, and weigh visibility against content-use concerns.

## Hands-on: a five-minute indexability check

1. `https://www.example.com/robots.txt` — any `Disallow: /` or rules blocking CSS/JS? Is the sitemap listed?
2. Search Console → **URL Inspection** on three important URLs: indexed? Google-selected canonical = yours? Last crawl date? Mobile rendering OK?
3. Search Console → **Pages** (indexing report): which reasons affect pages you *want* indexed?
4. Search Console → **Sitemaps**: submitted vs indexed counts reasonable?

A structured data snippet for an article with author information (supports E-E-A-T signals and rich result eligibility where applicable):

```json
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Interior Designer Cost in Dubai (2026 Guide)",
  "datePublished": "2026-02-10",
  "dateModified": "2026-09-01",
  "author": {"@type": "Person", "name": "Layla Haddad", "url": "https://example.com/team/layla-haddad/"},
  "publisher": {"@type": "Organization", "name": "Studio Noor Interiors", "url": "https://example.com/"}
}
```

Only change `dateModified` when the content materially changes. For deeper coverage — rendering, JavaScript SEO, crawl budget, international SEO and structured data at scale — continue with **Technical SEO Mastery (technical-seo-mastery)**.

## Common mistakes

- Using robots.txt to try to remove pages from the index.
- Staging-site `noindex` or `Disallow: /` accidentally shipped to production.
- Sitemaps full of redirected or noindexed URLs.
- Structured data that does not match visible content.

## Video lecture: Crawling, indexing, sitemaps and structured data

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Crawling, indexing and structured data
2. Why it matters
3. The pipeline
4. The controls
5. Structured data
6. AI crawlers: the confusing bit
7. A deliberate decision
8. Example 1 (simple)
9. Example 2 (illustrative)
10. Five-minute check
11. Watch me do it: the 5-minute check
12. Common mistakes
13. Recap

## Lecture transcript

### Crawling, indexing and structured data

You can write the best page on the internet, and if search engines can't crawl or index it, it doesn't exist. Not in the blue links, and not in AI answers either. In this lecture you'll learn how search engines process your site, the controls that decide what gets crawled and indexed, how structured data works, and what's changed with AI crawlers, including a control that trips up almost everyone.

### Why it matters

Why does this matter for content strategists, not just developers? Because a single technical mistake can erase months of content work overnight. A staging site's disallow rule shipped to production. Thousands of filter URLs wasting crawl capacity so new articles aren't found. A canonical pointing to the wrong page. Knowing the basics lets you spot these problems early and brief developers clearly.

### The pipeline

Here's the pipeline. Crawling: bots like Googlebot and Bingbot discover URLs by following links and reading sitemaps. Rendering: they load the page, including JavaScript, to see it as a user would. Indexing: they analyse and store eligible pages. Ranking and serving: when someone searches, they select and order indexed pages. Think of it like a library. Finding the book, reading it, cataloguing it, and recommending it. A failure at any step means no recommendation.

### The controls

Now the controls. robots.txt sits at the root and tells crawlers which paths they may crawl. Here's the key idea: robots.txt controls crawling, not indexing. A blocked URL can still appear in results if other pages link to it. To keep a page out of the index, allow crawling and use a noindex tag. Never block CSS or JavaScript needed for rendering. XML sitemaps list the canonical, indexable URLs you want found. And canonical tags tell search engines which duplicate version you prefer, supported by redirects and consistent internal links.

### Structured data

Structured data, usually JSON-LD using the schema dot org vocabulary, describes page content in a machine readable way. It can make pages eligible for rich results like product prices, breadcrumbs, events, recipes, videos and organisation details. The rules: mark up only what's visible on the page, keep it accurate, and test it with the Rich Results Test and the Schema Markup Validator. Rich results are never guaranteed, and Google has reduced or retired several types over time. Misleading markup, like fake review stars, can lead to manual actions.

### AI crawlers: the confusing bit

Now the AI part, and the control that confuses everyone. Google's AI Overviews and AI Mode draw on the normal search index, crawled by Googlebot. There's a robots token called Google Extended. Blocking it controls whether your content is used to train future Gemini models and to ground some Gemini products. But blocking it does not remove your pages from AI Overviews or AI Mode. Google's controls for those are snippet controls, like nosnippet, which also affect classic snippets, and, since twenty twenty-six, a Search generative AI control in Search Console.

### A deliberate decision

That Search Console control lets a site opt out of AI features in Search and Discover without affecting normal results. It's a genuine business decision. Opting out may protect content from being summarised, but it also removes you as a potential cited source. Weigh visibility against your content use concerns, check the current Search Console help, and decide deliberately rather than by accident. Other AI companies publish their own crawler names, like GPTBot, OAI SearchBot, ClaudeBot and PerplexityBot, each with its own documentation.

### Example 1 (simple)

Worked example one, simple. A small UK charity relaunches its website and traffic falls to almost nothing within two weeks. Someone checks robots.txt: it says disallow everything, copied from the staging site. They also find a site wide noindex tag in the theme settings. Both are fixed, the sitemap is resubmitted, and key pages are inspected in Search Console. Over the following weeks, pages return to the index. The lesson: add a robots and noindex check to every launch checklist.

### Example 2 (illustrative)

Worked example two, realistic and illustrative. A Karachi e-commerce site sees many product pages marked discovered, currently not indexed. Investigation shows thousands of filter URLs being crawled, wasting crawl capacity, and missing canonicals. Fixes: canonical tags on filter pages, blocking crawl of infinite filter combinations where appropriate, cleaner internal links, and an accurate product sitemap. Over the following weeks more product pages are indexed. And the new pages start appearing as sources in some AI answers too, because they're finally in the index.

### Five-minute check

Here's a five minute indexability check you can run today. Open your robots.txt: any disallow everything, or rules blocking CSS and JavaScript? Is the sitemap listed? Inspect three important URLs in Search Console: indexed, Google selected canonical matching yours, last crawl date, and mobile rendering. Check the page indexing report for reasons affecting pages you want indexed. And compare submitted and indexed counts in the sitemaps report.

### Watch me do it: the 5-minute check

Watch me do it. I'll run the five minute indexability check on a charity website that just relaunched. First, I type the domain followed by slash robots dot txt. There it is: user agent star, disallow slash. The whole site is blocked, left over from staging. That's the emergency, and I message the developer immediately. Next, while that gets fixed, I open Search Console and inspect the donate page with URL Inspection. It says the page isn't indexed, and the live test shows a noindex tag. So there's a second problem: a site wide noindex setting in the theme. I add that to the developer message. Then I open the Pages report and check the reasons affecting pages the charity wants indexed. Most say blocked by robots dot txt, which matches. Then the Sitemaps report: the new sitemap was never submitted, so I submit it. Once the developer confirms both fixes are live, I run URL Inspection again on three key pages and request indexing. Finally, I test the charity's Organization structured data in the Rich Results Test, and it's valid. I add robots and noindex checks to their launch checklist, so this never happens again.

### Common mistakes

Common mistakes. Using robots.txt to try to remove pages from the index. Shipping staging noindex or disallow rules to production. Sitemaps full of redirected or noindexed URLs. Structured data that doesn't match visible content. And assuming blocking Google Extended removes you from AI Overviews. It doesn't.

### Recap

Recap. Search engines crawl, render, index and rank, and AI answers in Google draw on the same index. robots.txt controls crawling, noindex controls indexing, sitemaps and canonicals keep things clean, and structured data must match what's visible. Google Extended is about Gemini training and grounding, not AI Overviews. Try this now. Run the five minute indexability check on your site. For the deep dive, take Technical SEO Mastery next.

## Key takeaways

- Search engines crawl, render, index and then rank; a failure at any step blocks visibility.
- robots.txt controls crawling, not indexing; use noindex to keep pages out of the index.
- Keep sitemaps clean, use canonicals and 301 redirects for duplicates and hreflang for international versions.
- JSON-LD structured data can enable rich results but must match visible content; eligibility is never guaranteed.

## Try it

Check a website's robots.txt and sitemap, inspect three important URLs in Search Console (or a crawler tool) and note any crawling, indexing or canonical issues.

- [Previous: Google spam policies every content team must know](https://optimizeall.com/learn/seo-and-content-strategy/spam-policies-for-content-teams)
- [Next: Core Web Vitals, speed and mobile experience](https://optimizeall.com/learn/seo-and-content-strategy/core-web-vitals-and-mobile)
- [All lessons of SEO & Content Strategy](https://optimizeall.com/learn/seo-and-content-strategy)
