AI Search Optimization: SEO for AI Overviews & Answer EnginesAI crawler controls, llms.txt and policy decisions · Lesson 10 of 17
llms.txt, snippet controls and choosing your AI access policy
Video lecture
llms.txt, snippet controls and choosing your AI access policy
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 llms.txt and your AI access policy
Every few months a new file or setting promises to control how AI uses your website. The best known is llms.txt. Clients ask about it constantly. Should we add it? Will it get us into ChatGPT? In this lecture you'll get an honest status report on llms.txt as of twenty twenty-six, a map of every control you actually have, and a decision framework for writing an AI access policy your whole team can follow.
0:32 What llms.txt is
What is llms.txt? It's a proposal, published in twenty twenty-four by Jeremy Howard of Answer dot AI, for a Markdown file at the root of a website. It gives language models a concise, curated overview of a site: a summary and links to the most useful pages, often with Markdown versions of documentation. The idea is sensible. Help models and agents find the right content without wading through complex HTML. Think of it as a hand-written guide left at a hotel reception for visitors who don't have time to explore.
1:11 Status in 2026
Now the honest status. It's not an official standard. No major search or answer engine has publicly committed to using it for rankings, retrieval or citations. Google is explicit: its May twenty twenty-six guide for generative AI features lists llms.txt among things you don't need, and a June twenty twenty-six clarification says Google Search ignores it, so it neither helps nor harms. Where it does have a niche is developer documentation, where people point coding assistants and agents at it directly. So it's cheap and harmless, but an experiment, not a strategy.
1:51 The real control panel
So what controls do you actually have? Robots.txt per bot, which stops compliant crawlers fetching content. Content-Signal lines, which state your use preferences for search, AI input and training. Google's Search generative AI control, which removes you from AI Overviews, AI Mode and AI in Discover. Snippet controls like nosnippet and max-snippet, which also cut normal snippets. Noindex, which removes pages from search entirely. Authentication and paywalls, which are real enforcement. And CDN bot settings, which must stay in sync with everything else. Each one has a trade-off.
2:29 Decision framework
How do you choose? Start from the business model, not from fear or hype. Question one: what do you gain from AI visibility? Lead generation businesses, SaaS, local services and e-commerce usually gain a lot, because customers ask assistants for recommendations. Publishers and paid content businesses face a harder trade, because answers can substitute for visits. Question two: what are you trying to prevent? Training use, appearance in answers, or server load? Each has a different lever. Question three: which content is sensitive? Premium research deserves stricter controls than marketing pages.
3:09 Typical policy profiles
Here are common policy profiles. A local service business, agency or SaaS company typically allows search and user-initiated bots, and makes a conscious decision on training bots. A subscription publisher allows Googlebot, weighs each AI bot carefully, publishes public summaries, and may pursue licensing. An e-commerce store allows search bots so products can be recommended, and protects account and checkout paths. A research firm with paid reports allows public summaries, gates full reports, and blocks training crawlers.
3:42 Example 1: Lahore restaurant
Worked example one, simple. A family-run restaurant in Lahore asks whether it needs llms.txt. We look at what matters: its menu and opening hours aren't in plain HTML, and its Business Profile has old hours. So the answer is no, not now. Fix the menu page and the Business Profile first, allow the search bots, and leave training bots as a conscious choice. llms.txt would change nothing for a restaurant's visibility today.
4:13 Example 2: Islamabad research firm (illustrative)
Worked example two, with illustrative details. A policy research firm in Islamabad sells detailed reports but wants its public briefings cited. Full reports sit behind a login. Public two-page summaries show key findings and methodology. Search and user-initiated bots are allowed on the summaries. Training crawlers are blocked, and the Content-Signal line says AI train no. The Search Console AI control stays off, because AI citations of summaries drive report enquiries. And it adds llms.txt only for its open data documentation, where developers genuinely use it.
4:50 Hands-on: the one-page policy
The hands-on for this lesson is a one-page AI access policy. The lesson text has a template: goals, which search and answer engines are allowed, how user-initiated fetchers are handled, which training crawlers are blocked, the Content-Signal line, your Search Console AI control decision with a name and date, what's gated, how the CDN is aligned, the status of llms.txt, and how you'll monitor. Keep it in version control next to robots.txt, and review it every six months or whenever a provider changes its documentation.
5:27 Common mistakes
Common mistakes. Selling llms.txt as a guaranteed AI ranking factor. Blanket-blocking every AI bot, then wondering why the brand never appears in assistants. Using nosnippet sitewide to stop AI and losing normal search snippets. And having no written policy, so marketing, IT and the CDN contractor all change settings independently. The written policy is the cure for all four.
5:53 Watch me do it: a policy in one meeting
Watch me do it. I'll write a real AI access policy with a client in one meeting. Step one: goals. I ask the founder of a Dubai interior design studio, what do you want from AI assistants? She says: be recommended when people ask for designers in Dubai, and don't let my portfolio text train models. I write both lines. Step two: I fill the allow list: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, and the user-initiated fetchers. Step three: the block list for training: GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended. Step four: Content-Signal: search yes, AI input yes, AI train no. Step five: the Google decision. I explain the Search generative AI control. She wants to stay in AI Overviews, so the control stays off, and I write her name and today's date next to it. Step six: gated content: none. Step seven: CDN. I open her Cloudflare dashboard with her, find the AI crawler setting, and align it with the lists. Step eight: llms.txt. We skip it; she has no documentation site. Step nine: monitoring, a monthly log check. She approves, and I commit the policy next to robots.txt.
7:15 Who owns which lever?
One more practical point: where these decisions actually live in a company. robots.txt is usually owned by developers. The CDN bot settings are often owned by IT or an outside contractor. The Search Console AI control sits with whoever has owner access, often marketing. And llms.txt might be added by the documentation team. That's four owners for one policy, which is exactly how contradictions creep in. So name one accountable owner for the policy, list who controls each lever, and make changes only through that owner. It sounds bureaucratic, but it takes ten minutes to set up and prevents months of invisible damage.
8:00 Recap and try this now
Recap. llms.txt is a cheap experiment with a real niche in developer docs, but Google Search ignores it and no engine has committed to using it. Your real controls are robots.txt, Content-Signal, the Search Console AI control, snippet controls, authentication and your CDN. Choose from the business model. Try this now. Fill in the one-page policy template for your site, get one decision-maker to approve it, and store it next to robots.txt.
What llms.txt is
llms.txt is a proposal (published in 2024 by Jeremy Howard of Answer.AI) for a Markdown file at the root of a website — /llms.txt — that gives large language models a concise, curated overview of a site: a summary and links to the most useful pages, often with Markdown versions of documentation. The idea is to help models and agents find the right content at inference time without parsing complex HTML.
# Optimize All
> Digital marketing agency and learning platform for SEO, local SEO and paid media,
> serving the UAE, Saudi Arabia, Pakistan and the UK.
## Services
- [Technical SEO](https://www.example.com/services/technical-seo/): audits, migrations, JavaScript SEO
- [Local SEO](https://www.example.com/services/local-seo/): Google Business Profile, reviews, citations
## Academy
- [Technical SEO Mastery](https://www.example.com/academy/technical-seo-mastery/): advanced course
## Optional
- [Company facts](https://www.example.com/about/)Its status in 2026: be honest
- It is not an official standard (not an IETF RFC or W3C recommendation).
- As of the time of writing, no major search or answer engine has publicly committed to using llms.txt to decide rankings, retrieval or citations.
- Google is explicit: its May 2026 guide for generative AI features lists llms.txt among things you don't need, and a June 2026 clarification added to that guide says Google Search ignores it — keeping one neither helps nor harms visibility. Google representatives had earlier compared it to the long-ignored keywords meta tag.
- It has been adopted by some developer documentation sites, where coding assistants and agents may be pointed to it directly by users or tools.
- There is little public evidence that adding it changes AI visibility.
Practical stance: it's cheap and harmless to add, particularly for documentation-heavy or developer-facing sites, but treat it as an experiment, not a strategy. Don't sell it to clients as an "AI ranking factor", and don't prioritise it over accessibility, content quality and third-party presence. Revisit if providers announce support.
Other controls to know
| Control | What it does | Trade-off |
|---|---|---|
nosnippet / max-snippet / data-nosnippet | Limit text shown from your pages in Google results, including AI features | Also reduces normal snippets and can lower click-through |
noindex | Removes pages from the index | Removes them from Search entirely |
| robots.txt per AI bot | Stops compliant crawlers fetching content | May remove you from that engine's answers |
| Authentication/paywall | Hard access control | Content can't be cited except via public summaries |
| CDN AI-bot controls | Block or manage AI bots at the edge | Must be kept in sync with robots.txt intent |
| Search generative AI control (Search Console) | Excludes a property from Google's AI Overviews, AI Mode and AI in Discover | Removes AI exposure on Google; regular results unaffected |
| Content-Signal line in robots.txt | States preferences for search, ai-input and ai-train use | A signal only; depends on crawler operators honouring it |
| TDM reservation (EU) | Machine-readable opt-out from text-and-data mining under EU copyright rules (e.g. via robots.txt-style or metadata protocols) | Legal landscape evolving; seek legal advice |
Choosing a policy: a decision framework
Start from business model, not fear or hype.
1. What does your business gain from AI visibility?
- Lead-generation businesses, SaaS, local services, e-commerce: being mentioned and cited in answers is generally valuable — customers ask assistants for recommendations.
- Publishers and paid content businesses: answers may substitute for visits; the value exchange is contested. Licensing deals exist for some publishers.
2. What are you trying to prevent?
- Use of content for training future models? Consider blocking training-focused bots (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent).
- Appearance in answer engines? Blocking search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) will reduce that — including your citations and referral traffic.
- Server load from aggressive crawlers? Rate limit or block specific abusive bots at the CDN.
3. Which content is sensitive? Premium research, member-only content and proprietary data may warrant stricter controls than marketing pages.
Common policy profiles
| Profile | Typical choice |
|---|---|
| Local service business / agency / SaaS | Allow search and user-initiated bots; decide on training bots (many allow) |
| Publisher with subscriptions | Allow Googlebot; carefully consider each AI bot; public summaries; licensing discussions |
| E-commerce | Allow search bots so products can be recommended; protect account/checkout paths |
| Research firm with paid reports | Public summaries allowed; full reports gated; training bots blocked |
Document the policy, the reasons, and a review date (for example every six months), because the landscape shifts quickly.
Hands-on: a one-page AI access policy
Write it down, get it approved, and keep it next to robots.txt in version control:
AI ACCESS POLICY — example.com Owner: Head of Marketing Approved: 2026-09-15
Review: every 6 months or when a provider changes its crawler documentation
1. Goals: be discoverable and accurately cited in AI search; do not supply
training data for paid research reports.
2. Search & answer engines: ALLOW Googlebot, Bingbot, OAI-SearchBot,
Claude-SearchBot, PerplexityBot (all public paths except /account/, /cart/).
3. User-initiated fetchers: ALLOW ChatGPT-User, Claude-User, Perplexity-User.
Sensitive paths protected by authentication (robots.txt is not enforcement).
4. Training: BLOCK GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended,
Meta-ExternalAgent sitewide. Content-Signal: search=yes, ai-input=yes, ai-train=no
5. Google AI surfaces: Search generative AI control = OFF (we stay in AI Overviews
/ AI Mode). Reason: lead generation. Decided by: CMO, 2026-09-15.
6. /research/ full reports: behind login; public summaries allowed.
7. CDN: "AI bot" settings aligned with sections 2-4; checked monthly.
8. llms.txt: published for /docs/ as a low-cost experiment; no KPI attached.
9. Monitoring: monthly log review of AI bot status codes (Lesson 4.3).Worked example 2: a research firm in Islamabad
A policy research firm (illustrative) sells detailed reports but wants its public briefings cited. Its policy: full reports behind a login; public two-page summaries with key findings and methodology; search and user-initiated bots allowed on summaries; training crawlers blocked; Content-Signal set to ai-train=no; the Search Console AI control left off because AI citations of summaries drive report enquiries. It adds llms.txt only for its open data documentation, where developers point coding assistants at it directly.
Common mistakes
- Selling llms.txt as a guaranteed AI ranking factor.
- Blanket-blocking all AI bots then wondering why the brand never appears in assistants.
- Using
nosnippetsitewide to "stop AI" and losing normal search snippets. - Having no written policy, so different teams change CDN and robots rules inconsistently.
Key takeaways
- llms.txt is a 2024 proposal, not a standard; no major engine has committed to using it for ranking or citation.
- Treat llms.txt as a low-cost experiment, especially for documentation sites — never as a core strategy.
- Snippet controls, noindex, robots rules, paywalls and CDN controls each have trade-offs.
- Choose an AI access policy from your business model, document it and review it regularly.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a one-page AI access policy for a business: goals, bots allowed or blocked, sensitive sections, CDN settings and review date.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.