AI Search Optimization: SEO for AI Overviews & Answer Engines · Generative engine optimisation: content and page practices · lesson 5 of 17 · 12 min
Technical accessibility for AI retrieval
If they can't fetch it, they can't cite it
Most AI answer engines rely on crawlers and fetchers that are, on average, less capable than Googlebot. Many do not execute JavaScript, may time out on slow pages, and may be blocked by bot-protection rules you didn't know you had. Technical accessibility is the least glamorous and most reliable GEO lever.
Checklist: make content retrievable
- Server-render primary content. Key text, headings, prices, tables and links should be in the initial HTML response. Test with JavaScript disabled or with
curl:
curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" \
https://www.example.com/pricing/ | grep -i "per month"
(Using a crawler's user-agent string with curl only simulates the header; it doesn't reproduce the provider's IPs, so CDN rules may treat the real bot differently.)
- Check your CDN and bot protection. Some CDN and security providers offer one-click settings to block AI crawlers, and some have changed defaults for new sites. Review the settings deliberately. Check firewall logs for blocked requests from AI user agents you intended to allow.
- Keep pages fast and light. Slow time to first byte or huge HTML increases the chance of timeouts. Google's documentation says Googlebot fetches only the first 2 MB of an HTML file (PDFs up to 64 MB); other crawlers publish less, so keep key facts early in lean HTML.
- Use semantic HTML. Real headings (
h2,h3), lists, tables with header rows, and paragraphs. Content locked in images, canvas or PDFs-only is harder to use. - Avoid content behind interactions. Tabs and accordions are fine if the content is in the HTML; content loaded only on click is not.
- Make key facts text, not images. Pricing tables as images, infographics without text equivalents, and text inside videos without transcripts are invisible to many systems. Provide alt text and transcripts.
- Stable, canonical URLs. Consistent canonicals help all systems consolidate signals.
- Be indexed where it matters. Google and Bing indexing support multiple AI experiences; check both.
- Avoid accidental blocking. robots.txt rules,
noindex,nosnippetand login walls all limit use.
Structured data: helpful, not magic
Structured data (JSON-LD) helps search engines understand entities, products, organisations and articles. Google's 2026 guide for its AI features says no special schema is needed. Lesson 2.3 covers what structured data and product feeds do help with. It remains worth doing for rich results and clarity, but there's no public evidence that adding schema alone causes AI citations. Keep it accurate and consistent with visible content.
Handling paywalls and gated content
If your valuable content is behind a paywall or login, most AI systems can't retrieve it (by design). Options:
- Provide a meaningful public summary with key facts and a clear path to the full content.
- For news and subscription content in Google, use the paywalled content structured data so Google understands the paywall isn't cloaking.
- Decide deliberately what's public; you can't be cited for what engines can't see.
Server logs: are AI crawlers visiting?
Filter logs for AI user agents and check status codes:
203.0.113.24 - - [02/Mar/2026:10:14:07 +0000] "GET /guides/hreflang/ HTTP/1.1" 200 48211 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)"
198.51.100.7 - - [02/Mar/2026:10:15:31 +0000] "GET /pricing/ HTTP/1.1" 403 1024 "-" "Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)"
The second line shows a 403 — perhaps a firewall rule blocking a crawler you intended to allow. As with Googlebot, user agents can be spoofed; several providers publish IP ranges for verification.
Multimedia and multilingual content
- Transcripts for videos and podcasts make spoken expertise retrievable.
- Descriptive captions and alt text for charts and diagrams.
- For bilingual audiences, publish genuine Arabic or Urdu versions with hreflang; AI systems answering in those languages need sources in them.
A quick audit for five key pages
| Check | Page 1 | Page 2 | Page 3 | Page 4 | Page 5 | |---|---|---|---|---|---| | Key facts in raw HTML | | | | | | | Indexed in Google / Bing | | | | | | | AI crawlers allowed (robots + CDN) | | | | | | | Answer-first headings | | | | | | | Facts as text (not images) | | | | | | | Updated in last 12 months | | | | | |
Hands-on: test what AI crawlers receive (Python)
This script requests your key pages with several crawler user-agent strings and reports status, HTML size and whether must-have facts appear in the raw HTML. It simulates headers only — real bots come from their providers' IP ranges, so CDN rules may still treat them differently. Check each provider's documentation for the current full user-agent strings.
# pip install requests
import requests
PAGES = {
"https://www.example.com/pricing/": ["per month", "Arabic"],
"https://www.example.com/services/technical-seo/": ["Dubai", "audit"],
}
AGENTS = {
"browser": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/126.0 Safari/537.36",
"OAI-SearchBot": "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)",
"PerplexityBot": "Mozilla/5.0 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)",
"ClaudeBot": "Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)",
}
for url, facts in PAGES.items():
for name, ua in AGENTS.items():
try:
r = requests.get(url, headers={"User-Agent": ua}, timeout=15, allow_redirects=True)
except requests.RequestException as exc:
print(f"{name:14} {url} ERROR {exc}")
continue
html = r.text
missing = [f for f in facts if f.lower() not in html.lower()]
print(f"{name:14} {r.status_code} {len(r.content)/1024:6.0f} KB "
f"hops={len(r.history)} missing={missing or 'none'} {url}")
Read the output like this: a 403 or 503 for a bot you meant to allow is a CDN/WAF rule; "missing" facts with a 200 usually mean client-side rendering; a large difference in size between the browser and bot rows suggests bot-specific handling that you should understand (and that must not become cloaking).
Worked example 2: a Karachi electronics retailer
A Karachi electronics retailer (illustrative) sees plenty of Googlebot activity but almost no AI referrals. The script shows 403 for PerplexityBot and OAI-SearchBot — the CDN's "block AI bots" toggle had been switched on by an IT contractor — and product specs missing from raw HTML because they load from an API after render. Fixes: align the CDN setting with the written AI policy, server-render specification tables, and move key product facts above heavy scripts. The team then watches logs for 200 responses to those bots on product URLs before expecting any change in answers.
Measuring success
- Share of priority pages where all must-have facts are in raw HTML (target: all).
- Status codes served to allowed AI bots in logs (target: 200/304, not 403/429/5xx).
- Pages indexed in both Google and Bing (Search Console and Bing Webmaster Tools).
Common mistakes
- Client-side rendered pricing and comparison pages.
- CDN "block AI bots" toggles switched on without a decision.
- Key content only in PDFs or images.
- Expecting schema to compensate for inaccessible content.
Video lecture: Technical accessibility for AI retrieval
Lecture coming soon · 13 chapters · about 9 minutes. Read the full transcript below.
- Technical accessibility for AI retrieval
- The mental model
- Checks 1 and 2
- Checks 3 and 4
- Hands-on: the fetch test
- Reading the output
- Example 1: Abu Dhabi yoga studio
- Example 2: Karachi retailer (illustrative)
- More checks and common mistakes
- Confirm with logs
- Watch me do it: the fetch test
- Paywalls and gated content
- Recap and try this now
Lecture transcript
Technical accessibility for AI retrieval
Here's an uncomfortable truth. You can write the most citable page on the internet, and still never be cited, because the machine that was supposed to read it got a blank page, a timeout, or a four-oh-three error. In this lecture you'll learn how to make sure AI crawlers can actually fetch your content, how to test it yourself in minutes, and what to fix first when they can't.
The mental model
Here's the mental model. Googlebot is a highly capable reader. It renders JavaScript, retries politely, and has decades of engineering behind it. Many AI crawlers are more like a speed-reader with no patience. Many don't run JavaScript. They may time out on slow pages. And your own security settings may block them without anyone noticing. So the least glamorous GEO lever is also the most reliable: make sure the important facts are in the first HTML response, fast, and not blocked.
Checks 1 and 2
First check: server-render your primary content. Prices, specs, headings, tables and links should be in the initial HTML. The fastest test is curl, or simply disabling JavaScript in your browser. If your pricing table disappears, many answer engines never saw it. Second check: size and speed. Google's documentation says Googlebot fetches only the first two megabytes of an HTML file. Other crawlers publish less. So keep key facts early, and keep pages lean.
Checks 3 and 4
Third check, and it catches more sites than you'd believe: your CDN and bot protection. Several CDN and security providers offer one-click settings to block AI crawlers, and some have changed defaults for new sites. So a contractor flips a switch, and your AI visibility quietly disappears. Look at your firewall events for blocked requests from AI user agents you intended to allow. Fourth: semantic HTML. Real headings, lists and tables with header rows are far easier to use than content locked inside images, canvas elements or PDFs.
Hands-on: the fetch test
Now, hands-on. The lesson text includes a short Python script. You list your key pages and the facts that must appear on each one, like per month or Arabic. The script requests each page with a normal browser user agent and with the user-agent strings of a few AI crawlers, then prints the status code, the size of the HTML, the number of redirects, and which facts are missing. One caveat: this only simulates the header. Real bots come from their providers' IP ranges, so your CDN may still treat them differently. That's why you confirm with logs.
Reading the output
How do you read the results? A four-oh-three or five-oh-three for a bot you meant to allow means a CDN or firewall rule. A two hundred with facts missing usually means client-side rendering. And a big size difference between the browser row and the bot rows means the site treats bots differently. Understand why, because serving bots different content from users can cross into cloaking.
Example 1: Abu Dhabi yoga studio
Worked example one, simple. A yoga studio in Abu Dhabi puts its timetable and prices in an embedded image because it looked nicer. The fetch test shows two hundred status everywhere, but the facts price, Arabic class and Saturday are missing. The fix is a plain HTML timetable with the image kept as decoration. Ten minutes of work, and the studio's key facts are now readable by every crawler.
Example 2: Karachi retailer (illustrative)
Worked example two, with illustrative details. A Karachi electronics retailer sees lots of Googlebot activity but almost no AI referrals. The fetch test returns four-oh-three for PerplexityBot and OAI-SearchBot. An IT contractor had switched on the CDN's block AI bots toggle. On top of that, product specs load from an API after rendering, so the raw HTML has no specs at all. Fixes: align the CDN setting with the written AI policy, server-render the specification tables, and move key facts above heavy scripts. Then check logs for two hundred responses to those bots before expecting answers to change.
More checks and common mistakes
A few more items for your checklist. Content behind tabs and accordions is fine if it's in the HTML, but content loaded only on click isn't. Videos and podcasts need transcripts. Charts need text equivalents. Paywalled content needs a meaningful public summary if you want to be cited at all. And make sure you're indexed in both Google and Bing, because several AI experiences rely on those indexes. Common mistakes: client-side rendered pricing pages, AI-blocking toggles switched on without a decision, and expecting schema to rescue content that crawlers can't reach.
Confirm with logs
Let's talk about logs, because they're where you confirm what the fetch test suggests. Filter your server or CDN logs for AI user agents and look at the status codes they receive. A healthy picture is mostly two hundreds and three-oh-fours on the pages you care about. A worrying picture is a wall of four-oh-threes, four-twenty-nines or five hundreds for a bot you intended to allow. Remember user agents can be faked, so where providers publish IP ranges, verify against them. We'll build a proper log script in module four. For now, just know that logs are the ground truth for whether a bot actually reached your content.
Watch me do it: the fetch test
Watch me do it. I'll run the fetch test on a real pricing page and act on the result. First, I open the script and set two pages: the pricing page, with the facts per month and Arabic, and the technical SEO service page, with Dubai and audit. I run it. The browser row for pricing says two hundred, eighty kilobytes, nothing missing. The OAI-SearchBot row says two hundred, twelve kilobytes, missing per month and Arabic. That size gap is the clue. Next, I view the raw source in the browser and search for per month. It isn't there; the pricing table is an empty container filled by JavaScript. So step three: I note the finding, pricing is client-side rendered, and I check the PerplexityBot row. That one says four-oh-three. Different problem. Step four: I open the CDN's firewall events and filter by user agent containing PerplexityBot. I see a managed rule blocking it. Step five: I write two tickets. One: server-render the pricing table. Two: align the CDN rule with our AI policy. And I add a note to confirm in logs next week that PerplexityBot gets two hundreds. Two different causes, found in about ten minutes.
Paywalls and gated content
A special case: paywalls and gated content. If your best material sits behind a login, most AI systems can't retrieve it, by design. That might be exactly what you want. But then decide deliberately what's public. A meaningful public summary with the key facts and a clear path to the full content lets you be cited without giving everything away. For news and subscription content in Google, use the paywalled content structured data so Google understands the paywall isn't cloaking. You can't be cited for what engines can't see.
Recap and try this now
Recap. If they can't fetch it, they can't cite it. Put key facts in the first HTML response, keep pages lean and fast, check CDN and firewall settings, and use semantic HTML. Try this now. Copy the script from the lesson text, add your five most important pages with two or three must-have facts each, and run it. Anything that isn't a clean two hundred with no missing facts goes straight onto your fix list.
Key takeaways
- Many AI crawlers don't run JavaScript; server-render key content and test with JavaScript off.
- Review CDN and bot-protection settings so you don't block crawlers by accident.
- Use semantic HTML, text-based facts, transcripts and alt text; ensure Google and Bing indexing.
- Structured data helps understanding but isn't required for AI features and isn't a magic citation lever.
Try it
Run the five-page accessibility audit on your site and fix at least one issue that affects retrievability.