AI Search Optimization: SEO for AI Overviews & Answer EnginesAI crawler controls, llms.txt and policy decisions · Lesson 9 of 17
AI crawlers and robots.txt: tokens and what they control
Video lecture
AI crawlers and robots.txt: tokens and what they control
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 AI crawlers and robots.txt
Here's a mistake I see every month. A company reads a headline about AI companies training on websites, copies a list of bot names into robots.txt, and blocks all of them. Months later, they ask why ChatGPT and Claude never cite their guides. They blocked the search bots along with the training bots. In this lecture you'll learn the difference between training crawlers, search crawlers and user-initiated fetchers, what each robots.txt token actually controls, and how to test your file so it matches the policy you meant to set.
0:39 Three kinds of AI bot
Here's the mental model. AI companies run up to three kinds of bot. Training crawlers collect content that may be used to train future models. Search crawlers build the index an answer engine retrieves from when it answers questions. And user-initiated fetchers visit a page because a person in a chat just asked the assistant to look at it. Think of a library. One visitor photocopies books for an archive, one writes the catalogue, and one fetches a single book because a reader asked for it. Blocking one doesn't block the others. And blocking the cataloguer means you never appear in the catalogue.
1:24 The main tokens
Let's map the main tokens. OpenAI: GPTBot for training, OAI-SearchBot for ChatGPT search, and ChatGPT-User for user requests. Anthropic: ClaudeBot for training, Claude-SearchBot for search, and Claude-User for user requests. Perplexity: PerplexityBot for its index and Perplexity-User for user requests. Apple: Applebot-Extended, a control token for training Apple's models. Common Crawl's CCBot feeds many training datasets. And Google, which needs its own slide, because it's the one people get wrong most often.
1:55 Google: two switches, two jobs
Google-Extended is not a crawler. It's a control token that governs whether content Google crawls can be used for training future Gemini models and for grounding in the Gemini apps and Vertex AI. It does not remove you from Google Search, and it does not remove you from AI Overviews or AI Mode. Those run on Googlebot and Search's own controls. In twenty twenty-six, Google added a separate lever for them: the Search generative AI control in Search Console settings, which excludes you from AI Overviews, AI Mode and AI Overviews in Discover while keeping you in regular results. So remember: two switches, two jobs.
2:41 User-initiated fetchers
Now the tricky bit: user-initiated fetchers. Providers treat them differently, and their documentation has changed. At the time of writing, OpenAI says robots.txt rules may not apply to ChatGPT-User, because a person initiated the action. Perplexity describes its user fetcher similarly. Google's user-triggered fetchers, including a new one called Google-Agent added in March twenty twenty-six for AI agents, generally ignore robots.txt. Anthropic, by contrast, says all its bots, including Claude-User, honour robots.txt. So if you need hard enforcement for sensitive content, robots.txt isn't enough. Use authentication, firewall rules or rate limits.
3:21 Two rules to get right
Two robots.txt rules to get right. First, grouping. You can list several user-agent lines above one set of rules; that's valid under the standard, RFC ninety-three-oh-nine. Second, and this one catches experts too, a bot obeys only its most specific matching group. If you add a group for OAI-SearchBot that says allow everything, it no longer inherits the rules in the star group, like disallow slash account. So repeat any shared rules inside each specific group. Otherwise you may accidentally open your account pages to the very bots you just configured.
4:01 Content Signals
There's a newer layer too: Content Signals. robots.txt controls whether a bot may fetch a page, not what it may do with it afterwards. Cloudflare's Content Signals Policy adds a line to robots.txt stating your preferences for three uses: search, AI input, meaning using content in answers at query time, and AI training. For example, search yes, AI input yes, AI train no. Cloudflare applies search yes and AI train no by default on its managed robots.txt. But be clear with clients: these are signals of preference, not blocks. It's up to each crawler operator to honour them.
4:44 Example 1: Dubai design studio
Worked example one, simple. A Dubai interior design studio wants to appear in ChatGPT and Perplexity answers, but doesn't want its portfolio text used to train models. The rules are straightforward. Allow OAI-SearchBot, Claude-SearchBot and PerplexityBot. Block GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended. Keep Googlebot and Bingbot untouched. Repeat the account and cart disallows in each specific group. Then test it, which brings us to the hands-on.
5:13 Hands-on: the robots grid
The lesson text includes a short Python script using only the standard library. It reads your live robots.txt and prints a grid: every AI token down the side, key paths across the top, and allow or block in each cell. One caveat: Python's parser applies rules in file order, not Google's longest-match precedence, so use it for per-bot checks and confirm tricky allow and disallow overlaps in Search Console's robots.txt report. Compare the grid with your written policy. Any mismatch is either a decision nobody recorded, or a bug.
5:52 Watch me do it: robots.txt vs policy
Watch me do it. I'll audit a live robots.txt against a written policy. The policy says: allow search and user-initiated AI bots on public pages, block training crawlers, keep account and cart private. Step one: I open the site's robots.txt and read it top to bottom. There's a star group with disallows for account and cart, and a new group for OAI-SearchBot with allow slash. Step two: I run the grid script from the lesson with thirteen tokens and four paths. The grid shows OAI-SearchBot allowed on slash account. That's the most-specific-group rule biting: the new group doesn't inherit the star group's disallows. Step three: GPTBot shows allowed everywhere, but the policy says block. Nobody added a GPTBot group. Step four: I draft the corrected file. I copy the account and cart disallows into the OAI-SearchBot group, and add a grouped block for GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended with disallow slash. I add a Content-Signal line stating ai-train equals no. Step five: I run the grid again on the draft. Every cell now matches the policy. I attach the before and after grids to the ticket, and schedule a CDN check, because robots.txt isn't the only layer.
7:18 Example 2: the 'block all AI' mistake (illustrative)
Worked example two, with illustrative details. A London agency copied a block every AI bot list into a client's robots.txt after a news story. Three months later the client asked why ChatGPT and Claude never cited its guides. The grid showed the search bots blocked alongside the training crawlers. The corrected policy allowed search and user-initiated access, kept the training crawlers blocked because the client preferred that, and recorded the reasoning with a six-month review date. And the team also checked the CDN, because edge settings can override what robots.txt implies.
7:58 Recap and try this now
Recap. Separate training, search and user-initiated bots. Google-Extended is about Gemini training and grounding, not AI Overviews; the Search Console control handles AI surfaces. User fetchers may ignore robots.txt, so use authentication for anything truly private. Mind the most-specific-group rule, and use Content Signals to state preferences. Try this now. Run the robots grid script on your site, put it next to your written AI policy, and fix or document every mismatch.
Why this is confusing
AI companies operate several different bots with different purposes: collecting training data, building a search index for their answer engine, and fetching pages on demand when a user asks. Blocking one doesn't necessarily block the others, and some controls affect training but not search features. Getting this wrong can either remove you from AI answers you wanted, or fail to stop the use you wanted to prevent.
Important: providers update names and policies frequently. Verify against each provider's current documentation before changing production rules.
The main tokens (as documented by providers at the time of writing)
| Provider | Token | Purpose (as described by the provider) |
|---|---|---|
| OpenAI | GPTBot | Crawling content that may be used to train OpenAI's models |
| OpenAI | OAI-SearchBot | Surfacing sites in ChatGPT's search features |
| OpenAI | ChatGPT-User | Fetches made on behalf of users during a conversation |
| Anthropic | ClaudeBot | Collecting web content that may contribute to model training |
| Anthropic | Claude-SearchBot | Improving search result quality for Claude's search features |
| Anthropic | Claude-User | Fetching pages when a Claude user asks |
Google-Extended | A robots.txt control token (not a separate crawler) governing whether content crawled by Google may be used for training future Gemini models and for grounding in Gemini apps and Vertex AI | |
Googlebot | Search crawling — also what feeds AI Overviews and AI Mode | |
Google-Agent | User-triggered fetcher for AI agents running on Google infrastructure (added to Google's fetcher list in March 2026); as a user-triggered fetcher it generally does not follow robots.txt | |
| Perplexity | PerplexityBot | Indexing for Perplexity's search results |
| Perplexity | Perplexity-User | User-initiated fetches |
| Apple | Applebot-Extended | Control token for whether Applebot-crawled content can be used to train Apple's generative models |
| Common Crawl | CCBot | Open web crawl widely used in AI training datasets |
| Meta | Meta-ExternalAgent | Crawling for AI training and related uses |
| ByteDance | Bytespider | Crawler associated with ByteDance |
| Amazon | Amazonbot | Crawler used for Amazon services including Alexa-related answers |
The key nuance about Google
Blocking Google-Extended does not remove your content from Google Search or from AI Overviews/AI Mode, which use Googlebot-crawled content under Search's controls. For those surfaces you now have two kinds of lever:
- The Search generative AI control in Search Console settings (rolled out globally by the end of August 2026) excludes a property from AI Overviews, AI Mode and AI Overviews in Discover while keeping it in regular results. It is not a robots.txt rule and has no effect on training.
- The older snippet controls —
nosnippet,data-nosnippet,max-snippetandnoindex— still limit how content is shown in AI features, but also affect normal search snippets.
Microsoft has described using noarchive/nocache meta directives to control how content is used in Bing's chat experiences; check Microsoft's current guidance.
User-initiated fetchers
Providers treat fetchers acting on a user's explicit request differently, and their documentation has changed over time. At the time of writing: OpenAI's documentation says robots.txt rules may not apply to ChatGPT-User because its actions are user-initiated; Perplexity describes Perplexity-User as generally ignoring robots.txt for user-requested fetches; Google's user-triggered fetchers (including Google-Agent) generally ignore robots.txt; Anthropic states that its bots, including Claude-User, honour robots.txt (and it documents Crawl-delay support). Re-read each provider's page before relying on any of this. If you need hard enforcement, use server-side controls (authentication, firewall rules, rate limiting) rather than relying only on robots.txt.
Compliance and verification
robots.txt is a voluntary standard. Major providers state they respect it for their crawlers, but compliance across the whole ecosystem varies, and there have been public disputes (for example, Cloudflare alleged in 2025 that Perplexity used undeclared crawlers to access blocked content; Perplexity disputed the claims). Verify bots via published IP ranges where available and monitor logs.
Example robots.txt patterns
Allow AI search and assistants, block training-focused crawlers:
# Search & user-initiated access: allowed (no rules = allowed)
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
# Training-focused crawlers and tokens: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: *
Disallow: /cart/
Disallow: /account/
Sitemap: https://www.example.com/sitemap_index.xmlNote: grouping several User-agent lines before one set of rules is valid under RFC 9309. Remember a bot obeys its most specific matching group only — the OAI-SearchBot group above does not inherit the * group's cart and account rules, so repeat any shared rules if they matter.
Block a section from all AI crawlers but keep it in Google Search:
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /members-research/CDN-level controls
Several CDNs now offer AI-crawler management (block, allow, or charge-per-crawl style schemes in some cases) and dashboards showing AI bot activity. These can override what your robots.txt implies, so align both layers.
Content Signals: expressing use preferences
robots.txt says whether a bot may fetch a URL, not what it may do with the content afterwards. Cloudflare's Content Signals Policy (2025) adds a machine-readable line to robots.txt expressing preferences for three uses — search (building a search index), ai-input (feeding content into AI answers at query time) and ai-train (training or fine-tuning models):
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /Cloudflare applies search=yes, ai-train=no by default on its managed robots.txt. These are signals of preference, not blocks; honouring them is up to each crawler operator, and legal effect varies by jurisdiction. Use them to state your policy clearly, and pair them with the per-bot rules above for bots that document compliance.
Hands-on: test your robots.txt against every AI token
# Python standard library only. Note: urllib.robotparser applies rules in file
# order (first match), not Google's longest-match precedence, so use it for
# per-bot allow/deny checks, and confirm tricky Allow/Disallow overlaps with
# Search Console's robots.txt report or a Google-compatible parser.
from urllib import robotparser
SITE = "https://www.example.com"
TOKENS = ["Googlebot", "Bingbot", "OAI-SearchBot", "GPTBot", "ChatGPT-User",
"Claude-SearchBot", "ClaudeBot", "Claude-User", "PerplexityBot",
"Perplexity-User", "Google-Extended", "Applebot-Extended", "CCBot"]
PATHS = ["/", "/pricing/", "/blog/ai-search-guide/", "/members-research/report.pdf"]
rp = robotparser.RobotFileParser(SITE + "/robots.txt")
rp.read()
print("token".ljust(18), *[p[:22].ljust(24) for p in PATHS])
for t in TOKENS:
print(t.ljust(18), *[("allow" if rp.can_fetch(t, SITE + p) else "BLOCK").ljust(24) for p in PATHS])Compare the grid with your written policy. Any mismatch is either a policy decision nobody recorded or a bug.
Worked example 2: an agency's "block all AI" mistake
A London agency (illustrative) copied a "block every AI bot" list into a client's robots.txt after a news story about training. Three months later the client asked why ChatGPT and Claude never cited its guides. The grid above showed OAI-SearchBot, Claude-SearchBot and PerplexityBot blocked alongside the training crawlers. The corrected policy allowed search and user-initiated access, kept GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended blocked per the client's training preference, and documented the reasoning with a six-month review date.
Common mistakes
- Blocking Google-Extended and expecting to disappear from AI Overviews.
- Blocking OAI-SearchBot when you only meant to block training (GPTBot).
- Copying a "block all AI" list from a blog without a business decision.
- Forgetting the most-specific-group rule when adding AI bot groups.
Key takeaways
- AI providers run separate bots for training, search indexing and user-initiated fetches; control each deliberately.
- Google-Extended controls Gemini training/grounding uses, not AI Overviews; Search features use Googlebot and Search controls.
- User-initiated fetchers may treat robots.txt differently; use server-side controls for hard enforcement.
- Verify bots, monitor logs, align CDN settings with robots.txt, and re-check provider documentation regularly.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Document which AI user agents appear in your logs over two weeks, their status codes, and whether each matches your intended policy.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.