AI Search Optimization: SEO for AI Overviews & Answer EnginesAI crawler controls, llms.txt and policy decisions · Lesson 11 of 17

AI bot log analysis: who's crawling, what they get, and are they real?

Article · 16 min · 8 min lecture

Video lecture

AI bot log analysis: who's crawling, what they get, and are they real?

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

AI bot log analysis

  • What bots actually did
  • Verify they're real
  • Catch silent blocks in days

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why logs are the ground truth for AI access

Your policy says which bots are allowed; robots.txt and the CDN try to enforce it; logs show what actually happened. For AI search specifically, logs answer questions no dashboard does: Are AI search bots reaching our key pages at all? What status codes do they get? Did a CDN change silently start blocking them? Are "GPTBot" requests really from OpenAI, or a scraper borrowing the name?

What to collect

  • Where: your CDN's logs (Cloudflare Logpush, Fastly, Akamai, CloudFront) if traffic is cached at the edge; otherwise origin server logs. Origin-only logs miss cached hits.
  • Fields: timestamp, client IP, method, URL (with query string), status, bytes, user agent, and ideally the CDN's bot or firewall action.
  • How long: at least 2-4 weeks; AI crawlers are bursty.
  • Privacy: logs contain IP addresses (personal data under UK/EU GDPR and similar laws in the Gulf and Pakistan). Restrict access, keep only what you need, and follow your retention policy.

Classify the bots

ClassUser-agent tokens to matchQuestion to answer
Search enginesGooglebot, bingbotBaseline crawling health
AI searchOAI-SearchBot, Claude-SearchBot, PerplexityBotCan answer engines index our key pages?
User-initiatedChatGPT-User, Claude-User, Perplexity-User, Google-AgentAre people asking assistants to open our pages?
AI trainingGPTBot, ClaudeBot, CCBot, Applebot, Meta-ExternalAgent, Bytespider, AmazonbotAre training crawlers behaving per policy?

ChatGPT-User and Claude-User hits are an interesting demand signal: each one generally means a person asked an assistant to read that page.

Verify before you trust

User agents are trivially spoofed. Verify in one of two ways:

  1. Reverse DNS then forward DNS (Google and Microsoft document this): the IP should resolve to a hostname under the provider's domain (for Googlebot, googlebot.com, google.com or googleusercontent.com; for Bingbot, search.msn.com), and that hostname should resolve back to the same IP.
  2. Published IP ranges: Google publishes JSON files of crawler IP ranges, and OpenAI publishes IP range files for its bots (linked from its crawler documentation). Other providers vary. Load the current files from each provider's documentation — URLs change, so don't hard-code them from a blog post.

Hands-on: a Python AI-bot log report

# pip install pandas
import ipaddress, json, re, socket, sys, urllib.request
import pandas as pd

LOG = sys.argv[1] if len(sys.argv) > 1 else "access.log"
LINE = re.compile(r'(?P<ip>\S+) \S+ \S+ \[(?P<ts>[^\]]+)\] "(?P<method>\S+) (?P<url>\S+) [^"]*" '
                  r'(?P<status>\d{3}) (?P<bytes>\S+) "[^"]*" "(?P<ua>[^"]*)"')
CLASSES = {
    "search":   r"Googlebot|bingbot",
    "ai_search": r"OAI-SearchBot|Claude-SearchBot|PerplexityBot",
    "ai_user":  r"ChatGPT-User|Claude-User|Perplexity-User|Google-Agent",
    "ai_train": r"GPTBot|ClaudeBot|CCBot|Applebot|Meta-ExternalAgent|Bytespider|Amazonbot",
}
# Paste the CURRENT IP-range JSON URLs from each provider's documentation.
IP_RANGE_FILES = {"GPTBot": "", "OAI-SearchBot": ""}

def load_ranges(url):
    if not url:
        return []
    with urllib.request.urlopen(url, timeout=20) as r:
        data = json.load(r)
    return [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
            for p in data.get("prefixes", [])]

RANGES = {bot: load_ranges(u) for bot, u in IP_RANGE_FILES.items()}

def rdns_ok(ip, suffixes):
    try:
        host = socket.gethostbyaddr(ip)[0]
        return host.endswith(suffixes) and ip in socket.gethostbyname_ex(host)[2]
    except (socket.herror, socket.gaierror):
        return False

rows = [m.groupdict() for m in map(LINE.match, open(LOG, encoding="utf-8", errors="ignore")) if m]
df = pd.DataFrame(rows)
df["cls"] = "other"
for cls, pat in CLASSES.items():
    df.loc[df.ua.str.contains(pat, regex=True), "cls"] = cls
df["bot"] = df.ua.str.extract("(" + "|".join(CLASSES.values()) + ")", expand=False)
bots = df[df.cls != "other"].copy()

def verify(row):
    if row.bot == "Googlebot":
        return rdns_ok(row.ip, (".googlebot.com", ".google.com", ".googleusercontent.com"))
    if row.bot == "bingbot":
        return rdns_ok(row.ip, (".search.msn.com",))
    nets = RANGES.get(row.bot)
    if nets:
        return any(ipaddress.ip_address(row.ip) in n for n in nets)
    return None   # unknown: no published method configured

unique_ips = bots.drop_duplicates(["bot", "ip"]).copy()
unique_ips["verified"] = unique_ips.apply(verify, axis=1)
bots = bots.merge(unique_ips[["bot", "ip", "verified"]], on=["bot", "ip"], how="left")

bots["path"] = bots.url.str.split("?").str[0]
print(bots.pivot_table(index="bot", columns="status", values="url", aggfunc="count", fill_value=0))
print(bots.groupby(["bot", "verified"], dropna=False).size())
print(bots[bots.cls.isin(["ai_search", "ai_user"])].path.value_counts().head(25))

What to look for in the output:

  • Status by bot: allowed AI search bots should see mostly 200/304. A wall of 403 means a firewall rule; 429 means rate limiting; 5xx means your server struggled.
  • Verified = False: spoofed traffic — exclude it from analysis and consider blocking it.
  • Top paths for AI search and user bots: are your priority pages among them? Pages users ask assistants to open are strong candidates for accuracy checks.

Worked example: a Dubai marketplace discovers a silent block

A Dubai property marketplace (illustrative) runs the report monthly. In August, OAI-SearchBot and PerplexityBot requests drop to almost zero and the few remaining get 403. The CDN vendor had enabled a new managed rule for "AI crawlers" on the account. The team re-aligns the rule with the policy, confirms 200 responses the following week, and adds an alert: if verified AI search bot hits fall by more than half week on week, notify the SEO lead. They also find heavy "GPTBot" traffic from IPs outside OpenAI's published ranges — scrapers — and block those at the edge.

Measuring success

  • Share of verified AI search bot requests returning 200/304 (target: close to all).
  • Coverage: share of priority URLs requested by at least one AI search bot in the last 30 days.
  • Spoofed-bot share, and time-to-detect for access regressions (target: days, not months).

Common mistakes

  • Analysing unverified traffic and reporting scrapers as "AI interest".
  • Using origin logs when most bot hits are served from the CDN cache.
  • Treating crawl volume as proof of citations — it shows access, not selection.
  • Keeping raw logs with IP addresses longer than your privacy policy allows.

Key takeaways

  • Logs (preferably from the CDN) are the only ground truth for what AI bots actually requested and received.
  • Classify bots as search, AI search, user-initiated and training; user-initiated hits signal real people asking assistants about a page.
  • Verify bots with reverse-then-forward DNS or provider-published IP ranges; exclude spoofed traffic.
  • Allowed AI search bots should receive 200/304; 403, 429 and 5xx patterns point to firewall, rate-limit or server problems.
  • Monitor weekly with a drop alert; crawl volume shows access, not citations.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A request claims to be GPTBot but its IP isn't in OpenAI's published ranges. What should you conclude?
  2. Why prefer CDN logs over origin logs for bot analysis?
  3. OAI-SearchBot requests suddenly return 403. What's the most likely cause?

Put it into practice

Run the log script on two weeks of CDN or server logs, verify bots, and list the status codes allowed AI search bots received on your five most important pages.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.