SEO & Content StrategyTechnical SEO basics · Lesson 6 of 17

Crawling, indexing, sitemaps and structured data

Article · 12 min · 8 min lecture

Video lecture

Crawling, indexing, sitemaps and structured data

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Crawling, indexing and structured data

  • Crawl → render → index → rank
  • robots.txt, noindex, sitemaps, canonicals
  • Structured data
  • AI crawlers and controls

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

How search engines process your site

  1. Crawling: bots such as Googlebot and Bingbot discover URLs by following links and reading sitemaps.
  2. Rendering: they load the page, including JavaScript, to see the content as a user would.
  3. Indexing: they analyse the content and store eligible pages in the index.
  4. Ranking/serving: when someone searches, they select and order indexed pages.

A problem at any step stops a page from appearing, no matter how good the content is.

robots.txt

robots.txt sits at the root (example.com/robots.txt) and tells crawlers which paths they may crawl. Example:

User-agent: *
Disallow: /cart/
Disallow: /account/
Sitemap: https://example.com/sitemap.xml

Key facts:

  • robots.txt controls crawling, not indexing. A blocked URL can still appear in results (usually without a description) if other pages link to it.
  • To keep a page out of the index, allow crawling and use a noindex meta robots tag or HTTP header instead.
  • Never block CSS or JavaScript files needed to render pages.
  • A single mistake (such as Disallow: / left over from a staging site) can block your whole site.

XML sitemaps

A sitemap lists the canonical URLs you want indexed, optionally with last-modified dates.

  • Include only indexable, canonical, 200-status URLs.
  • Keep it updated automatically (most CMSs and e-commerce platforms do this).
  • Submit it in Google Search Console and Bing Webmaster Tools, and reference it in robots.txt.
  • Large sites can use a sitemap index that links to multiple sitemaps.

Canonical tags and duplicate content

Duplicate or near-duplicate URLs are common: filter parameters (?colour=red), tracking parameters, HTTP vs HTTPS, www vs non-www, printer versions. The canonical tag (rel="canonical") tells search engines which version is preferred. Combine it with consistent internal linking and redirects (301) from old or alternate URLs to the preferred one.

Status codes to know

CodeMeaningSEO implication
200OKPage can be indexed
301Permanent redirectPasses signals to the new URL; use for moved pages
302Temporary redirectUse only for genuinely temporary moves
404Not foundFine for truly removed pages; fix internal links pointing to them
410GoneSignals intentional permanent removal
5xxServer errorFrequent errors can reduce crawling and cause pages to drop

International sites

For sites serving multiple languages or countries (for example English and Arabic versions for the UAE, or UK and US versions), hreflang annotations help search engines show the right version to each audience. Each version should reference the others and itself.

Structured data

Structured data (usually JSON-LD using the schema.org vocabulary) describes page content in a machine-readable way. It can make pages eligible for rich results such as product prices and ratings, breadcrumbs, events, recipes, videos, local business details and organisation information.

Example (simplified) for a local business:

{
  "@context": "https://schema.org",
  "@type": "LocalBusiness",
  "name": "Studio Noor Interiors",
  "address": {"@type": "PostalAddress", "addressLocality": "Dubai", "addressCountry": "AE"},
  "telephone": "+971-4-000-0000",
  "url": "https://example.com"
}

Rules:

  • Mark up only content visible on the page, and keep it accurate.
  • Use Google's Rich Results Test and the Schema Markup Validator to check.
  • Rich results are never guaranteed, and eligibility changes (Google has reduced or retired some rich result types over time, such as limiting FAQ rich results to certain authoritative sites). Check the current documentation.
  • Misleading markup (for example fake review stars) can lead to manual actions.

Worked example

A Karachi e-commerce site notices many product pages are "Discovered – currently not indexed". Investigation shows thousands of filter URLs being crawled, wasting crawl capacity, and missing canonicals. Fixes: canonical tags on filter pages, blocking crawl of infinite filter combinations in robots.txt where appropriate, cleaner internal links and an accurate sitemap of product URLs. Over the following weeks, more product pages are indexed.

AI crawlers and the index behind AI answers

Google's AI Overviews and AI Mode draw on Google's normal search index, crawled by Googlebot. That has two practical consequences for content teams:

  • If a page cannot be crawled or indexed, it cannot appear in classic results or be used as a source in Google's AI features.
  • The Google-Extended robots.txt token controls whether content is used for training future Gemini models and for grounding in some Gemini products — blocking it does not remove pages from AI Overviews or AI Mode. Google's documented controls for Search AI features are snippet controls (nosnippet, max-snippet), which also affect classic snippets, and — since 2026 — a Search generative AI control in Search Console that lets a site opt out of AI features in Search and Discover without affecting normal results. Check the current Search Console help before using it, and weigh visibility against content-use concerns.

Hands-on: a five-minute indexability check

  1. https://www.example.com/robots.txt — any Disallow: / or rules blocking CSS/JS? Is the sitemap listed?
  2. Search Console → URL Inspection on three important URLs: indexed? Google-selected canonical = yours? Last crawl date? Mobile rendering OK?
  3. Search Console → Pages (indexing report): which reasons affect pages you want indexed?
  4. Search Console → Sitemaps: submitted vs indexed counts reasonable?

A structured data snippet for an article with author information (supports E-E-A-T signals and rich result eligibility where applicable):

{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Interior Designer Cost in Dubai (2026 Guide)",
  "datePublished": "2026-02-10",
  "dateModified": "2026-09-01",
  "author": {"@type": "Person", "name": "Layla Haddad", "url": "https://example.com/team/layla-haddad/"},
  "publisher": {"@type": "Organization", "name": "Studio Noor Interiors", "url": "https://example.com/"}
}

Only change dateModified when the content materially changes. For deeper coverage — rendering, JavaScript SEO, crawl budget, international SEO and structured data at scale — continue with Technical SEO Mastery (technical-seo-mastery).

Common mistakes

  • Using robots.txt to try to remove pages from the index.
  • Staging-site noindex or Disallow: / accidentally shipped to production.
  • Sitemaps full of redirected or noindexed URLs.
  • Structured data that does not match visible content.

Key takeaways

  • Search engines crawl, render, index and then rank; a failure at any step blocks visibility.
  • robots.txt controls crawling, not indexing; use noindex to keep pages out of the index.
  • Keep sitemaps clean, use canonicals and 301 redirects for duplicates and hreflang for international versions.
  • JSON-LD structured data can enable rich results but must match visible content; eligibility is never guaranteed.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. You want a thank-you page kept out of search results. What should you use?
  2. Which URLs belong in an XML sitemap?
  3. Which structured data practice can lead to a manual action?

Put it into practice

Check a website's robots.txt and sitemap, inspect three important URLs in Search Console (or a crawler tool) and note any crawling, indexing or canonical issues.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.