AI Search Optimization: SEO for AI Overviews & Answer EnginesAI crawler controls, llms.txt and policy decisions · Lesson 9 of 17

AI crawlers and robots.txt: tokens and what they control

Article · 14 min · 8 min lecture

Video lecture

AI crawlers and robots.txt: tokens and what they control

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

AI crawlers and robots.txt

  • Training vs search vs user-initiated
  • What each token really controls
  • Test your file against your policy

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why this is confusing

AI companies operate several different bots with different purposes: collecting training data, building a search index for their answer engine, and fetching pages on demand when a user asks. Blocking one doesn't necessarily block the others, and some controls affect training but not search features. Getting this wrong can either remove you from AI answers you wanted, or fail to stop the use you wanted to prevent.

Important: providers update names and policies frequently. Verify against each provider's current documentation before changing production rules.

The main tokens (as documented by providers at the time of writing)

ProviderTokenPurpose (as described by the provider)
OpenAIGPTBotCrawling content that may be used to train OpenAI's models
OpenAIOAI-SearchBotSurfacing sites in ChatGPT's search features
OpenAIChatGPT-UserFetches made on behalf of users during a conversation
AnthropicClaudeBotCollecting web content that may contribute to model training
AnthropicClaude-SearchBotImproving search result quality for Claude's search features
AnthropicClaude-UserFetching pages when a Claude user asks
GoogleGoogle-ExtendedA robots.txt control token (not a separate crawler) governing whether content crawled by Google may be used for training future Gemini models and for grounding in Gemini apps and Vertex AI
GoogleGooglebotSearch crawling — also what feeds AI Overviews and AI Mode
GoogleGoogle-AgentUser-triggered fetcher for AI agents running on Google infrastructure (added to Google's fetcher list in March 2026); as a user-triggered fetcher it generally does not follow robots.txt
PerplexityPerplexityBotIndexing for Perplexity's search results
PerplexityPerplexity-UserUser-initiated fetches
AppleApplebot-ExtendedControl token for whether Applebot-crawled content can be used to train Apple's generative models
Common CrawlCCBotOpen web crawl widely used in AI training datasets
MetaMeta-ExternalAgentCrawling for AI training and related uses
ByteDanceBytespiderCrawler associated with ByteDance
AmazonAmazonbotCrawler used for Amazon services including Alexa-related answers

The key nuance about Google

Blocking Google-Extended does not remove your content from Google Search or from AI Overviews/AI Mode, which use Googlebot-crawled content under Search's controls. For those surfaces you now have two kinds of lever:

  • The Search generative AI control in Search Console settings (rolled out globally by the end of August 2026) excludes a property from AI Overviews, AI Mode and AI Overviews in Discover while keeping it in regular results. It is not a robots.txt rule and has no effect on training.
  • The older snippet controls — nosnippet, data-nosnippet, max-snippet and noindex — still limit how content is shown in AI features, but also affect normal search snippets.

Microsoft has described using noarchive/nocache meta directives to control how content is used in Bing's chat experiences; check Microsoft's current guidance.

User-initiated fetchers

Providers treat fetchers acting on a user's explicit request differently, and their documentation has changed over time. At the time of writing: OpenAI's documentation says robots.txt rules may not apply to ChatGPT-User because its actions are user-initiated; Perplexity describes Perplexity-User as generally ignoring robots.txt for user-requested fetches; Google's user-triggered fetchers (including Google-Agent) generally ignore robots.txt; Anthropic states that its bots, including Claude-User, honour robots.txt (and it documents Crawl-delay support). Re-read each provider's page before relying on any of this. If you need hard enforcement, use server-side controls (authentication, firewall rules, rate limiting) rather than relying only on robots.txt.

Compliance and verification

robots.txt is a voluntary standard. Major providers state they respect it for their crawlers, but compliance across the whole ecosystem varies, and there have been public disputes (for example, Cloudflare alleged in 2025 that Perplexity used undeclared crawlers to access blocked content; Perplexity disputed the claims). Verify bots via published IP ranges where available and monitor logs.

Example robots.txt patterns

Allow AI search and assistants, block training-focused crawlers:

# Search & user-initiated access: allowed (no rules = allowed)
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# Training-focused crawlers and tokens: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
Disallow: /

User-agent: *
Disallow: /cart/
Disallow: /account/

Sitemap: https://www.example.com/sitemap_index.xml

Note: grouping several User-agent lines before one set of rules is valid under RFC 9309. Remember a bot obeys its most specific matching group only — the OAI-SearchBot group above does not inherit the * group's cart and account rules, so repeat any shared rules if they matter.

Block a section from all AI crawlers but keep it in Google Search:

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /members-research/

CDN-level controls

Several CDNs now offer AI-crawler management (block, allow, or charge-per-crawl style schemes in some cases) and dashboards showing AI bot activity. These can override what your robots.txt implies, so align both layers.

Content Signals: expressing use preferences

robots.txt says whether a bot may fetch a URL, not what it may do with the content afterwards. Cloudflare's Content Signals Policy (2025) adds a machine-readable line to robots.txt expressing preferences for three uses — search (building a search index), ai-input (feeding content into AI answers at query time) and ai-train (training or fine-tuning models):

User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /

Cloudflare applies search=yes, ai-train=no by default on its managed robots.txt. These are signals of preference, not blocks; honouring them is up to each crawler operator, and legal effect varies by jurisdiction. Use them to state your policy clearly, and pair them with the per-bot rules above for bots that document compliance.

Hands-on: test your robots.txt against every AI token

# Python standard library only. Note: urllib.robotparser applies rules in file
# order (first match), not Google's longest-match precedence, so use it for
# per-bot allow/deny checks, and confirm tricky Allow/Disallow overlaps with
# Search Console's robots.txt report or a Google-compatible parser.
from urllib import robotparser

SITE = "https://www.example.com"
TOKENS = ["Googlebot", "Bingbot", "OAI-SearchBot", "GPTBot", "ChatGPT-User",
          "Claude-SearchBot", "ClaudeBot", "Claude-User", "PerplexityBot",
          "Perplexity-User", "Google-Extended", "Applebot-Extended", "CCBot"]
PATHS = ["/", "/pricing/", "/blog/ai-search-guide/", "/members-research/report.pdf"]

rp = robotparser.RobotFileParser(SITE + "/robots.txt")
rp.read()
print("token".ljust(18), *[p[:22].ljust(24) for p in PATHS])
for t in TOKENS:
    print(t.ljust(18), *[("allow" if rp.can_fetch(t, SITE + p) else "BLOCK").ljust(24) for p in PATHS])

Compare the grid with your written policy. Any mismatch is either a policy decision nobody recorded or a bug.

Worked example 2: an agency's "block all AI" mistake

A London agency (illustrative) copied a "block every AI bot" list into a client's robots.txt after a news story about training. Three months later the client asked why ChatGPT and Claude never cited its guides. The grid above showed OAI-SearchBot, Claude-SearchBot and PerplexityBot blocked alongside the training crawlers. The corrected policy allowed search and user-initiated access, kept GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended blocked per the client's training preference, and documented the reasoning with a six-month review date.

Common mistakes

  • Blocking Google-Extended and expecting to disappear from AI Overviews.
  • Blocking OAI-SearchBot when you only meant to block training (GPTBot).
  • Copying a "block all AI" list from a blog without a business decision.
  • Forgetting the most-specific-group rule when adding AI bot groups.

Key takeaways

  • AI providers run separate bots for training, search indexing and user-initiated fetches; control each deliberately.
  • Google-Extended controls Gemini training/grounding uses, not AI Overviews; Search features use Googlebot and Search controls.
  • User-initiated fetchers may treat robots.txt differently; use server-side controls for hard enforcement.
  • Verify bots, monitor logs, align CDN settings with robots.txt, and re-check provider documentation regularly.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. You block Google-Extended. What happens to AI Overviews?
  2. You want to appear in ChatGPT search but not contribute to OpenAI model training. Which rules fit?
  3. Which Google lever excludes a site from AI Overviews and AI Mode while keeping normal search snippets?

Put it into practice

Document which AI user agents appear in your logs over two weeks, their status codes, and whether each matches your intended policy.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.