Technical SEO Audit Workshop · Running the crawl and gathering evidence · lesson 5 of 14 · 16 min
Log-file analysis lab: Cloudflare logs to crawl insights in Python
The lab brief
Kiran Home's developers have exported four weeks of Cloudflare HTTP request logs via Logpush as newline-delimited JSON (NDJSON). Your job: verify search engine bots, classify every request by template and market, quantify crawl waste, find products Google hasn't requested, and produce two charts for the stakeholder report. All data is illustrative.
What the data looks like
Each line is one request. Field names depend on the fields selected in the Logpush job; a typical selection includes:
{"EdgeStartTimestamp":"2026-03-14T06:25:24Z","ClientIP":"66.249.66.1","ClientRequestHost":"www.kiranhome.example",
"ClientRequestURI":"/en-ae/collections/cushions/?colour=indigo&sort=price","EdgeResponseStatus":200,
"ClientRequestUserAgent":"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"CacheCacheStatus":"hit"}
Check your Logpush job's field list and set its timestamp format to RFC 3339 (the default can be Unix time) — if a field is missing (for example the user agent), ask the developers to add it; you can't analyse bots without it. Logs contain IP addresses, which are personal data: store them securely and delete raw files at the end of the engagement.
Step 1: load, verify and classify
# pip install pandas
import gzip, json, re, socket, glob
import pandas as pd
rows = []
for path in glob.glob("logs/*.log.gz"):
with gzip.open(path, "rt", encoding="utf-8", errors="ignore") as f:
for line in f:
try:
rows.append(json.loads(line))
except json.JSONDecodeError:
continue
df = pd.DataFrame(rows).rename(columns={"ClientIP": "ip", "ClientRequestURI": "uri",
"EdgeResponseStatus": "status", "ClientRequestUserAgent": "ua", "EdgeStartTimestamp": "ts"})
df["ts"] = pd.to_datetime(df.ts, utc=True, errors="coerce")
bots = df[df.ua.str.contains("Googlebot|bingbot", case=False, na=False)].copy()
_cache = {}
def verified(ip, ua):
key = (ip, "google" if "Googlebot" in ua else "bing")
if key not in _cache:
sfx = (".googlebot.com", ".google.com", ".googleusercontent.com") if key[1] == "google" else (".search.msn.com",)
try:
host = socket.gethostbyaddr(ip)[0]
_cache[key] = host.endswith(sfx) and ip in socket.gethostbyname_ex(host)[2]
except (socket.herror, socket.gaierror):
_cache[key] = False
return _cache[key]
bots["verified"] = [verified(i, u) for i, u in zip(bots.ip, bots.ua)]
g = bots[bots.verified & bots.ua.str.contains("Googlebot")].copy()
SEG = [("facet", r"[?&](colour|size|price|sort)="), ("search", r"/search"), ("product", r"/products/"),
("category", r"/collections/"), ("journal", r"/journal/"), ("asset", r"\.(js|css|png|jpe?g|webp|avif|svg)(\?|$)")]
g["segment"] = g.uri.map(lambda u: next((n for n, p in SEG if re.search(p, u)), "other"))
g["market"] = g.uri.str.extract(r"^/(en-pk|en-ae|ar-ae|en-gb)/")[0].fillna("none")
g["legacy"] = g.uri.str.contains(r"\.html$|^/shop/", regex=True) # pre-relaunch URL patterns
Step 2: the four tables every log audit needs
print((g.segment.value_counts(normalize=True) * 100).round(1)) # share of crawl by segment
print(pd.crosstab(g.segment, g.status)) # status codes by segment
print(g[g.legacy].status.value_counts()) # how Google experiences old URLs
weekly = g.set_index("ts").groupby([pd.Grouper(freq="W"), "segment"]).size().unstack(fill_value=0)
weekly.to_csv("googlebot_weekly_by_segment.csv") # chart this, annotate releases
Step 3: best-sellers Google hasn't visited
products = pd.read_csv("products_revenue_prelaunch.csv") # url_path, revenue
hit_paths = set(g[g.segment == "product"].uri.str.split("?").str[0])
products["crawled_4w"] = products.url_path.isin(hit_paths)
top = products.sort_values("revenue", ascending=False).head(200)
print(f"{(~top.crawled_4w).sum()} of the top 200 revenue products had no verified Googlebot request in 4 weeks")
Reading Kiran Home's results (illustrative)
- Facet and search URLs take most verified Googlebot requests; products a small share.
- Legacy
.htmlURLs are still requested daily; a large share return 404 — confirming the redirect gap (RC3) matters. - A substantial number of the top 200 revenue products had no Googlebot request in four weeks — the headline for the founder.
- Cache status shows most bot hits served from cache — origin logs alone would have under-counted them.
Step 4: from tables to findings
Tables don't persuade; findings do. Convert each result into a finding with a number, a comparison and a consequence, and link it to a root cause:
| Log result | Finding wording for the workbook | Root cause | |---|---|---| | Facet + search share of verified Googlebot requests | "Most of Google's crawling goes to filter and search URLs that should not be indexed" | RC2 | | Legacy URL statuses | "Google still requests old URLs daily; a large share return 404, including URLs with backlinks" | RC3 | | Top-200 products not requested | "Many best-sellers were not visited by Google in four weeks" | RC4 | | Weekly trend | "Crawling of products fell after the relaunch while facet crawling rose" | RC2/RC4 |
Chart the weekly table as a stacked area by segment, annotate the relaunch and each fix release, and reuse the same chart for verification in Module 5 — before/after on one visual is the most convincing evidence you can show.
Worked example 2: a smaller site without Logpush
A Leeds bakery chain (illustrative) on shared hosting has only Apache combined logs. The same approach works with the regex parser from the Technical SEO Mastery course; with only a few thousand bot hits a month, a spreadsheet pivot after parsing is enough. The key finding there: Googlebot requests an old /menu.pdf hundreds of times a month, a file that was replaced by an HTML menu — so the team redirects it to the HTML page.
Common mistakes
- Analysing unverified "Googlebot" lines.
- Ignoring cached hits by using origin logs only.
- Reporting raw counts instead of shares and trends by segment.
- Keeping raw log files with IPs after the engagement.
Video lecture: Log-file analysis lab: Cloudflare logs to crawl insights in Python
Lecture coming soon · 13 chapters · about 9 minutes. Read the full transcript below.
- Log-file analysis lab
- The data
- Step 1: load, verify, classify
- Why verify?
- Step 2: four tables
- Step 3: best-sellers not crawled
- Kiran Home results (illustrative)
- Logs confirm root causes
- Example 2: a small bakery chain (illustrative)
- Common mistakes
- Watch me do it: the lab end to end
- From tables to findings
- Recap and try this now
Lecture transcript
Log-file analysis lab
This is a lab lesson. Kiran Home's developers have exported four weeks of Cloudflare request logs, and it's your job to turn millions of lines into three or four findings a founder will understand. In this lecture you'll load the logs in Python, verify real Googlebot, classify every request by template and market, build the four tables every log audit needs, and find the best-selling products Google hasn't visited in a month. Everything here is illustrative data, but the workflow is real.
The data
First, what does the data look like? Cloudflare Logpush exports newline-delimited JSON: one request per line, with fields like the edge timestamp, client IP, host, request URI, response status, user agent and cache status. The exact fields depend on what was selected in the Logpush job, so check the field list. If the user agent is missing, ask for it to be added, because you can't analyse bots without it. And remember, IP addresses are personal data. Store the logs securely and delete the raw files when the engagement ends.
Step 1: load, verify, classify
Step one: load, verify and classify. The script reads every gzipped file line by line, parses the JSON, and renames the fields to short names. It filters to lines claiming to be Googlebot or bingbot, then verifies each IP with a reverse DNS lookup to a Google or Microsoft domain and a forward lookup back to the same IP, caching results so each IP is only checked once. Only verified Googlebot lines go forward. Then every request gets a segment, like facet, search, product, category, journal or asset, a market from the folder, and a legacy flag for pre-relaunch URL patterns.
Why verify?
Why does verification matter so much? Because on most real sites, a noticeable share of traffic claiming to be Googlebot is scrapers and SEO tools pretending. If you include them, you'll report crawl patterns that Google isn't actually responsible for. In one audit, unverified fake Googlebot traffic hammering the internal search pages made it look like Google was obsessed with search URLs. Once verified, the real Googlebot's pattern looked very different. Verify first, always.
Step 2: four tables
Step two: the four tables every log audit needs. One: the share of crawl by segment. Two: status codes by segment. Three: how Google experiences old URLs, meaning the status codes returned to legacy URL requests. And four: weekly requests by segment, saved to a CSV so you can chart it and annotate release dates. That weekly chart will become one of the most persuasive visuals in your stakeholder report, because it shows cause and effect over time.
Step 3: best-sellers not crawled
Step three is the join that gets a founder's attention. Load the list of products with their revenue from before the relaunch. Build the set of product paths that verified Googlebot requested in the last four weeks. Mark each product as crawled or not. Then look at the top two hundred revenue products and count how many had no verified Googlebot request at all in four weeks. That sentence, in plain language, is the headline of the log section.
Kiran Home results (illustrative)
Reading Kiran Home's illustrative results. Facet and search URLs take most verified Googlebot requests, and products only a small share. Legacy dot html URLs are still requested daily, and a large share return four-oh-four, which confirms the redirect gap matters. A substantial number of the top two hundred revenue products had no Googlebot request in four weeks. And the cache status shows most bot hits were served from Cloudflare's cache, so origin logs alone would have badly under-counted them.
Logs confirm root causes
How do these results connect to your root causes? The facet and search dominance is evidence for root cause two, crawlable facets and search. The legacy four-oh-fours support root cause three, the redirect gaps. And uncrawled deep products support root cause four, pagination that isn't crawlable. Logs rarely create new root causes on their own. What they do is quantify and confirm the ones you found in the crawl, and that confirmation is what gets tickets prioritised.
Example 2: a small bakery chain (illustrative)
Worked example two, a smaller site, illustrative. A Leeds bakery chain on shared hosting has only Apache combined logs, no Logpush. The same approach works with a regular expression parser, and with only a few thousand bot hits a month, a spreadsheet pivot after parsing is enough. The key finding: Googlebot requests an old menu PDF hundreds of times a month, a file that was replaced by an HTML menu. So they redirect the PDF to the HTML menu page. Small site, same method, useful result.
Common mistakes
Common mistakes. Analysing unverified Googlebot lines. Ignoring cached hits by using origin logs only. Reporting raw counts instead of shares and trends by segment, which makes big sites look scary and small sites look fine regardless of reality. And keeping raw log files full of IP addresses after the engagement ends. Measure your log work by whether each finding links to a root cause and a ticket.
Watch me do it: the lab end to end
Watch me do it. I'll run the lab end to end on the Kiran Home export. Step one: I check the files. Twenty-eight gzipped NDJSON files, one per day. I open the first line and confirm the fields: timestamp in RFC 3339, client IP, URI, status, user agent and cache status. Step two: I run step one of the script. It loads about three million lines, finds the requests claiming to be Googlebot or bingbot, and verifies each unique IP. The verification takes a couple of minutes thanks to the cache. About one in seven Googlebot-claiming requests fail verification. I note that for the report. Step three: I run step two and read the four tables. Facets and search take most verified Googlebot requests. Legacy dot html URLs return four-oh-four about half the time. Step four: I run the weekly table, open it in a spreadsheet, and make a stacked area chart with the relaunch annotated. The product band shrinks the week of the relaunch. Step five: I run step three. Of the top two hundred revenue products, a large share had no verified Googlebot request in four weeks. Step six: I write the four finding sentences, link each to a root cause, and delete the raw log copies from my laptop.
From tables to findings
Tables don't persuade people. Findings do. So convert each table into a sentence with a number, a comparison and a consequence, and link it to a root cause. For example: most of Google's crawling goes to filter and search URLs that shouldn't be indexed. Google still requests old URLs daily, and a large share return four-oh-four, including URLs with backlinks. And many best-sellers weren't visited by Google in four weeks. Then chart the weekly table as a stacked area by segment, annotate the relaunch, and reuse that same chart after the fixes. A before and after on one visual is the most convincing evidence you can show.
Recap and try this now
Recap. Load the logs, verify bots, classify by segment and market, build the four tables, and join with revenue to find the best-sellers Google ignored. Use the results to confirm and quantify root causes. Try this now. If you have access to any site's logs, run step one and step two on a single week. Just the share of verified Googlebot requests by segment will tell you whether crawl waste is a real problem there.
Key takeaways
- Get CDN logs (e.g. Cloudflare Logpush) with user agent, status, URI and cache status; treat IPs as personal data.
- Verify Googlebot with reverse and forward DNS before any analysis; fake bots distort conclusions.
- Four core tables: crawl share by segment, status by segment, legacy URL statuses, weekly trend by segment.
- Join uncrawled products with revenue to produce a founder-level headline.
- Use log findings to confirm and quantify root causes, not as standalone issues.
Try it
Run steps 1 and 2 of the lab on one week of logs from any site you can access and write down the share of verified Googlebot requests by segment.