---
title: "AI bot log analysis: who's crawling, what they get, and…"
description: "Why logs are the ground truth for AI access Your policy says which bots are allowed; robots.txt and the CDN try to enforce it; logs show what actually…"
url: https://optimizeall.com/learn/ai-search-optimization-geo/ai-bot-log-analysis-and-verification
updated: 2026-10-05
---

AI Search Optimization: SEO for AI Overviews & Answer Engines · AI crawler controls, llms.txt and policy decisions · lesson 11 of 17 · 16 min

# AI bot log analysis: who's crawling, what they get, and are they real?

## Why logs are the ground truth for AI access

Your policy says which bots are allowed; robots.txt and the CDN try to enforce it; **logs show what actually happened**. For AI search specifically, logs answer questions no dashboard does: Are AI search bots reaching our key pages at all? What status codes do they get? Did a CDN change silently start blocking them? Are "GPTBot" requests really from OpenAI, or a scraper borrowing the name?

## What to collect

- **Where:** your CDN's logs (Cloudflare Logpush, Fastly, Akamai, CloudFront) if traffic is cached at the edge; otherwise origin server logs. Origin-only logs miss cached hits.
- **Fields:** timestamp, client IP, method, URL (with query string), status, bytes, user agent, and ideally the CDN's bot or firewall action.
- **How long:** at least 2-4 weeks; AI crawlers are bursty.
- **Privacy:** logs contain IP addresses (personal data under UK/EU GDPR and similar laws in the Gulf and Pakistan). Restrict access, keep only what you need, and follow your retention policy.

## Classify the bots

| Class | User-agent tokens to match | Question to answer |
|---|---|---|
| Search engines | Googlebot, bingbot | Baseline crawling health |
| AI search | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Can answer engines index our key pages? |
| User-initiated | ChatGPT-User, Claude-User, Perplexity-User, Google-Agent | Are people asking assistants to open our pages? |
| AI training | GPTBot, ClaudeBot, CCBot, Applebot, Meta-ExternalAgent, Bytespider, Amazonbot | Are training crawlers behaving per policy? |

`ChatGPT-User` and `Claude-User` hits are an interesting demand signal: each one generally means a person asked an assistant to read that page.

## Verify before you trust

User agents are trivially spoofed. Verify in one of two ways:

1. **Reverse DNS then forward DNS** (Google and Microsoft document this): the IP should resolve to a hostname under the provider's domain (for Googlebot, `googlebot.com`, `google.com` or `googleusercontent.com`; for Bingbot, `search.msn.com`), and that hostname should resolve back to the same IP.
2. **Published IP ranges:** Google publishes JSON files of crawler IP ranges, and OpenAI publishes IP range files for its bots (linked from its crawler documentation). Other providers vary. Load the current files from each provider's documentation — URLs change, so don't hard-code them from a blog post.

## Hands-on: a Python AI-bot log report

```python
# pip install pandas
import ipaddress, json, re, socket, sys, urllib.request
import pandas as pd

LOG = sys.argv[1] if len(sys.argv) > 1 else "access.log"
LINE = re.compile(r'(?P<ip>\S+) \S+ \S+ \[(?P<ts>[^\]]+)\] "(?P<method>\S+) (?P<url>\S+) [^"]*" '
                  r'(?P<status>\d{3}) (?P<bytes>\S+) "[^"]*" "(?P<ua>[^"]*)"')
CLASSES = {
    "search":   r"Googlebot|bingbot",
    "ai_search": r"OAI-SearchBot|Claude-SearchBot|PerplexityBot",
    "ai_user":  r"ChatGPT-User|Claude-User|Perplexity-User|Google-Agent",
    "ai_train": r"GPTBot|ClaudeBot|CCBot|Applebot|Meta-ExternalAgent|Bytespider|Amazonbot",
}
# Paste the CURRENT IP-range JSON URLs from each provider's documentation.
IP_RANGE_FILES = {"GPTBot": "", "OAI-SearchBot": ""}

def load_ranges(url):
    if not url:
        return []
    with urllib.request.urlopen(url, timeout=20) as r:
        data = json.load(r)
    return [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
            for p in data.get("prefixes", [])]

RANGES = {bot: load_ranges(u) for bot, u in IP_RANGE_FILES.items()}

def rdns_ok(ip, suffixes):
    try:
        host = socket.gethostbyaddr(ip)[0]
        return host.endswith(suffixes) and ip in socket.gethostbyname_ex(host)[2]
    except (socket.herror, socket.gaierror):
        return False

rows = [m.groupdict() for m in map(LINE.match, open(LOG, encoding="utf-8", errors="ignore")) if m]
df = pd.DataFrame(rows)
df["cls"] = "other"
for cls, pat in CLASSES.items():
    df.loc[df.ua.str.contains(pat, regex=True), "cls"] = cls
df["bot"] = df.ua.str.extract("(" + "|".join(CLASSES.values()) + ")", expand=False)
bots = df[df.cls != "other"].copy()

def verify(row):
    if row.bot == "Googlebot":
        return rdns_ok(row.ip, (".googlebot.com", ".google.com", ".googleusercontent.com"))
    if row.bot == "bingbot":
        return rdns_ok(row.ip, (".search.msn.com",))
    nets = RANGES.get(row.bot)
    if nets:
        return any(ipaddress.ip_address(row.ip) in n for n in nets)
    return None   # unknown: no published method configured

unique_ips = bots.drop_duplicates(["bot", "ip"]).copy()
unique_ips["verified"] = unique_ips.apply(verify, axis=1)
bots = bots.merge(unique_ips[["bot", "ip", "verified"]], on=["bot", "ip"], how="left")

bots["path"] = bots.url.str.split("?").str[0]
print(bots.pivot_table(index="bot", columns="status", values="url", aggfunc="count", fill_value=0))
print(bots.groupby(["bot", "verified"], dropna=False).size())
print(bots[bots.cls.isin(["ai_search", "ai_user"])].path.value_counts().head(25))
```

What to look for in the output:

- **Status by bot:** allowed AI search bots should see mostly `200`/`304`. A wall of `403` means a firewall rule; `429` means rate limiting; `5xx` means your server struggled.
- **Verified = False:** spoofed traffic — exclude it from analysis and consider blocking it.
- **Top paths for AI search and user bots:** are your priority pages among them? Pages users ask assistants to open are strong candidates for accuracy checks.

## Worked example: a Dubai marketplace discovers a silent block

A Dubai property marketplace (illustrative) runs the report monthly. In August, `OAI-SearchBot` and `PerplexityBot` requests drop to almost zero and the few remaining get `403`. The CDN vendor had enabled a new managed rule for "AI crawlers" on the account. The team re-aligns the rule with the policy, confirms `200` responses the following week, and adds an alert: if verified AI search bot hits fall by more than half week on week, notify the SEO lead. They also find heavy "GPTBot" traffic from IPs outside OpenAI's published ranges — scrapers — and block those at the edge.

## Measuring success

- Share of verified AI search bot requests returning 200/304 (target: close to all).
- Coverage: share of priority URLs requested by at least one AI search bot in the last 30 days.
- Spoofed-bot share, and time-to-detect for access regressions (target: days, not months).

## Common mistakes

- Analysing unverified traffic and reporting scrapers as "AI interest".
- Using origin logs when most bot hits are served from the CDN cache.
- Treating crawl volume as proof of citations — it shows access, not selection.
- Keeping raw logs with IP addresses longer than your privacy policy allows.

## Video lecture: AI bot log analysis: who's crawling, what they get, and are they real?

Lecture coming soon · 12 chapters · about 8 minutes. Read the full transcript below.

1. AI bot log analysis
2. Collect the right logs
3. Four classes of bot
4. Verify before you trust
5. Script walkthrough
6. Reading the output
7. Example 1: UK charity
8. Example 2: Dubai marketplace (illustrative)
9. From report to monitoring
10. Watch me do it: the AI bot report
11. Common mistakes
12. Recap and try this now

## Lecture transcript

### AI bot log analysis

Your AI policy says which bots are welcome. Robots.txt and your CDN try to enforce it. But how do you know what actually happened last Tuesday? Did OpenAI's search bot reach your pricing page, or did it get a four-oh-three? And was that GPTBot really OpenAI, or a scraper borrowing the name? Logs answer those questions, and nothing else does. In this lecture you'll learn how to classify AI bots in your logs, how to verify they're real, and how to run a Python report that catches silent blocks within days.

### Collect the right logs

First, where do your logs come from? If you use a CDN, like Cloudflare, Fastly, Akamai or CloudFront, many requests are answered at the edge and never reach your server. So origin logs can miss a big share of bot traffic. Get the CDN's logs where you can. You need timestamp, IP address, method, URL with the query string, status code, bytes, user agent, and ideally the CDN's firewall or bot action. Collect at least two to four weeks, because AI crawlers are bursty. And remember IP addresses are personal data under GDPR and similar laws, so restrict access and follow your retention policy.

### Four classes of bot

Next, classify. I use four classes. Search engines, like Googlebot and bingbot, give you a baseline. AI search bots, like OAI-SearchBot, Claude-SearchBot and PerplexityBot, tell you whether answer engines can index your key pages. User-initiated fetchers, like ChatGPT-User, Claude-User, Perplexity-User and Google-Agent, tell you when real people asked an assistant to open your page. That's a lovely demand signal. And training crawlers, like GPTBot, ClaudeBot and CCBot, tell you whether your training policy is being respected.

### Verify before you trust

Now verification, because user agents are trivially faked. Anyone can type GPTBot into a scraper. There are two ways to check. Method one is reverse DNS, then forward DNS. Look up the hostname for the IP, check it ends in the provider's domain, like googlebot dot com for Google or search dot msn dot com for Bing, then resolve that hostname and confirm you get the same IP back. Method two is published IP ranges. Google publishes JSON files of crawler IP ranges, and OpenAI publishes ranges for its bots. Always copy the current file locations from the providers' own documentation, because they change.

### Script walkthrough

Let's walk through the script in the lesson text. It uses pandas and the standard library. A regular expression parses each line of a combined log into IP, timestamp, method, URL, status, bytes and user agent. A dictionary maps regex patterns to the four classes. A small function loads IP range files you paste in. Another does the reverse and forward DNS check. The script verifies each unique bot and IP pair once, which keeps it fast, then merges the verdict back onto every request. Finally it prints three tables: status codes per bot, verified versus unverified counts, and the top paths requested by AI search and user bots.

### Reading the output

How do you read the output? In the status table, allowed AI search bots should mostly get two hundreds and three-oh-fours. A wall of four-oh-threes means a firewall rule. Four-twenty-nines mean rate limiting. Five hundreds mean your server struggled. In the verification table, anything unverified with a big crawler's name is probably a scraper, so exclude it from analysis and consider blocking it. And in the top paths table, ask: are our priority pages here? If people keep asking assistants to open a particular page, check that page's facts carefully, because it's being read closely.

### Example 1: UK charity

Worked example one, simple. A small UK charity runs the script for the first time. Googlebot and bingbot look healthy. There are a few ChatGPT-User and Claude-User hits on its eligibility guide, which tells the team people are asking assistants whether they qualify for support. But PerplexityBot gets five-oh-three errors on that same guide, because the page is slow to generate. The fix is caching the guide. And the team now reviews that page's accuracy monthly, since people are clearly relying on it through assistants.

### Example 2: Dubai marketplace (illustrative)

Worked example two, with illustrative details. A Dubai property marketplace runs the report monthly. In August, OAI-SearchBot and PerplexityBot requests drop almost to zero, and the few that remain get four-oh-three. The CDN vendor had enabled a new managed rule for AI crawlers on the account. Nobody at the marketplace had decided that. The team re-aligns the rule with its policy and confirms two hundreds the following week. It also finds heavy GPTBot traffic from IPs outside OpenAI's published ranges. Scrapers. Those get blocked at the edge.

### From report to monitoring

Turn this into monitoring, not a one-off. Schedule the report weekly or monthly. Add a simple alert: if verified AI search bot hits fall by more than half week on week, notify the SEO lead. Track three measures. The share of verified AI search requests that return two hundred or three-oh-four. Coverage, meaning the share of priority URLs requested by at least one AI search bot in the last thirty days. And time to detect access regressions, which should be days, not months. One caution: crawl volume shows access, not selection. Being crawled doesn't mean being cited.

### Watch me do it: the AI bot report

Watch me do it. I'll run the AI bot report on two weeks of CDN logs. Step one: I export the logs, paste the current IP range URLs for GPTBot and OAI-SearchBot from OpenAI's crawler page into the script's config, and run it. Step two: the status table. Googlebot and bingbot are mostly two hundreds. OAI-SearchBot is mostly two hundreds too. But PerplexityBot shows a column of four-oh-threes that starts on the ninth. Step three: the verification table. There are thousands of GPTBot requests marked unverified: their IPs aren't in OpenAI's published ranges. Step four: I check the top paths for AI search and user bots. The pricing and delivery pages are there, which is good, and a delivery guide gets lots of ChatGPT-User hits, meaning people are asking ChatGPT to open it. Step five: I act. I open the CDN firewall events for the ninth and find a new managed rule; I ticket it against the policy. I add a rule to block the unverified GPTBot IPs at the edge. And I put the delivery guide on the monthly accuracy-check list. Finally, I set the drop alert so next time we find out within a week, not a quarter.

### Common mistakes

Common mistakes. Analysing unverified traffic and calling scrapers AI interest. Using origin logs when most bot hits are served from the CDN cache. Treating crawl volume as proof of citations. And keeping raw logs full of IP addresses longer than your privacy policy allows. Each of these either misleads your reporting or creates risk.

### Recap and try this now

Recap. Logs are the ground truth for AI access. Classify bots into search, AI search, user-initiated and training. Verify with reverse DNS or published IP ranges. Read status codes by bot, and watch your priority paths. Try this now. Export two weeks of logs, paste the current IP range URLs from the providers' docs into the script, and run it. Then answer one question: is every AI search bot you meant to allow getting two hundreds on your five most important pages?

## Key takeaways

- Logs (preferably from the CDN) are the only ground truth for what AI bots actually requested and received.
- Classify bots as search, AI search, user-initiated and training; user-initiated hits signal real people asking assistants about a page.
- Verify bots with reverse-then-forward DNS or provider-published IP ranges; exclude spoofed traffic.
- Allowed AI search bots should receive 200/304; 403, 429 and 5xx patterns point to firewall, rate-limit or server problems.
- Monitor weekly with a drop alert; crawl volume shows access, not citations.

## Try it

Run the log script on two weeks of CDN or server logs, verify bots, and list the status codes allowed AI search bots received on your five most important pages.

- [Previous: llms.txt, snippet controls and choosing your AI access policy](https://optimizeall.com/learn/ai-search-optimization-geo/llms-txt-and-choosing-a-policy)
- [Next: Measuring AI referrals and visibility](https://optimizeall.com/learn/ai-search-optimization-geo/measuring-ai-referrals-and-visibility)
- [All lessons of AI Search Optimization: SEO for AI Overviews & Answer Engines](https://optimizeall.com/learn/ai-search-optimization-geo)
