---
title: "Crawl budget and log-file analysis — Technical SEO Mastery"
description: "What crawl budget actually is Google describes crawl budget as the set of URLs Googlebot can and wants to crawl. It combines two things: - Crawl capacity…"
url: https://optimizeall.com/learn/technical-seo-mastery/crawl-budget-and-log-file-analysis
updated: 2026-10-05
---

Technical SEO Mastery · Crawl budget, robots.txt and indexing directives · lesson 3 of 18 · 15 min

# Crawl budget and log-file analysis

## What crawl budget actually is

Google describes crawl budget as the set of URLs Googlebot *can and wants to* crawl. It combines two things:

- **Crawl capacity limit** — how much Googlebot can crawl without overloading your server. It rises when your server responds quickly and falls with slow responses, 5xx errors and 429s.
- **Crawl demand** — how much Google *wants* to crawl, driven by perceived popularity, staleness and the size of the URL inventory it knows about.

**Who should care?** Google's own guidance says crawl budget is mainly a concern for very large sites (hundreds of thousands to millions of URLs) or sites whose content changes very often, plus sites with many URLs stuck in "Discovered – currently not indexed". A 200-page brochure site does not have a crawl budget problem — it has a quality or linking problem.

## The biggest crawl budget wasters

1. **Faceted and parameter URLs** — `?colour=red&size=9&sort=price` combinations can generate millions of near-duplicates.
2. **Infinite spaces** — calendars with endless "next month" links, internal search result pages, session IDs in URLs.
3. **Redirect chains and soft 404s** — each hop or empty page costs a fetch.
4. **Duplicate URLs** — http/https, www/non-www, trailing slash variants, uppercase variants all resolving with 200.
5. **Slow server responses** — reduce capacity directly.

## Log files: the only ground truth

Crawlers and Search Console tell you what *could* be crawled. **Server logs tell you what was crawled.** A typical combined log line:

```text
66.249.66.1 - - [14/Mar/2026:06:25:24 +0000] "GET /shoes/red-trainers?sort=price HTTP/1.1" 200 51234 "-" "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
```

Fields to extract: IP, timestamp, method, path (with query string), status, bytes, user agent. If your CDN terminates traffic, get logs from the CDN (Cloudflare, Akamai, Fastly, CloudFront) — origin logs may miss cached hits.

## Verify the bot before you trust it

Anyone can fake a Googlebot user agent. Verify by:

- **Reverse DNS** of the IP resolving to `googlebot.com`, `google.com` or `googleusercontent.com`, followed by a forward lookup that returns the same IP; or
- Matching against the **IP range JSON files** Google publishes for its crawlers.

```bash
host 66.249.66.1        # -> crawl-66-249-66-1.googlebot.com
host crawl-66-249-66-1.googlebot.com   # -> 66.249.66.1
```

## Log analysis procedure

1. Collect at least 2–4 weeks of logs (longer for large sites).
2. Filter to verified search engine bots.
3. Classify each URL by **template** (product, category, blog, facet, search, asset) using path patterns.
4. Join with your crawl data (indexability, canonical, in-sitemap) and Search Console data (clicks).
5. Answer the key questions:

| Question | What a problem looks like |
|---|---|
| What share of bot hits go to non-indexable URLs? | A large share on parameters, redirects or noindex pages |
| Are important pages crawled regularly? | Money pages with no bot hits in weeks |
| Are there status code problems? | Spikes of 5xx or 404 to URLs you thought were fixed |
| Are orphan URLs being crawled? | URLs in logs but not in your crawl — old links, forgotten sections |
| Is crawl activity changing? | Sudden drops after a release, or spikes from a parameter bug |

A quick command-line pass for a small site:

```bash
grep "Googlebot" access.log | awk '{print $9}' | sort | uniq -c | sort -rn   # status codes
grep "Googlebot" access.log | awk '{print $7}' | cut -d'?' -f1 | sort | uniq -c | sort -rn | head -50   # top paths
```

For large volumes use a dedicated log analyser or load the data into BigQuery or a spreadsheet pivot.

## Improving crawl efficiency

- Remove internal links to parameter URLs you do not want crawled; block truly useless patterns in robots.txt.
- Return 404/410 for gone content rather than redirecting everything to the home page.
- Fix redirect chains by linking directly to final URLs.
- Keep sitemaps limited to canonical, indexable, 200-status URLs.
- Improve server response times and caching; watch the **Crawl stats** report (Settings → Crawl stats) for average response time and host status.
- Use `304 Not Modified` responses with `If-Modified-Since`/ETag support where your stack allows — it saves resources on both sides.

## Hands-on: a pandas log analysis by template

The awk one-liners above are fine for a quick look. For real audits, classify every verified Googlebot hit by template and join it to your crawl.

```python
# pip install pandas
import re, socket
import pandas as pd

LINE = re.compile(r'(?P<ip>\S+) \S+ \S+ \[(?P<ts>[^\]]+)\] "\S+ (?P<url>\S+) [^"]*" (?P<status>\d{3}) \S+ "[^"]*" "(?P<ua>[^"]*)"')
TEMPLATES = [("facet", r"[?&](colour|size|price|sort)="), ("search", r"^/search"),
             ("product", r"^/products/"), ("category", r"^/collections/"),
             ("blog", r"^/blog/"), ("asset", r"\.(js|css|png|jpe?g|webp|avif|svg)(\?|$)")]

def template(url):
    for name, pat in TEMPLATES:
        if re.search(pat, url):
            return name
    return "other"

def is_google(ip, cache={}):
    if ip not in cache:
        try:
            host = socket.gethostbyaddr(ip)[0]
            cache[ip] = host.endswith((".googlebot.com", ".google.com", ".googleusercontent.com")) \
                        and ip in socket.gethostbyname_ex(host)[2]
        except (socket.herror, socket.gaierror):
            cache[ip] = False
    return cache[ip]

rows = [m.groupdict() for m in map(LINE.match, open("access.log", errors="ignore")) if m]
df = pd.DataFrame(rows)
df = df[df.ua.str.contains("Googlebot")]
df = df[df.ip.map(is_google)]                         # verified only
df["template"] = df.url.map(template)
df["date"] = pd.to_datetime(df.ts, format="%d/%b/%Y:%H:%M:%S %z").dt.date

print(df.template.value_counts(normalize=True).round(3))             # share of crawl by template
print(pd.crosstab(df.template, df.status))                            # status codes by template

crawl = pd.read_csv("crawl_export.csv")                                # from your crawler: Address, Indexability
crawl["path"] = crawl["Address"].str.replace(r"^https?://[^/]+", "", regex=True)
hits = df.groupby("url").size().rename("googlebot_hits").reset_index()
joined = crawl.merge(hits, left_on="path", right_on="url", how="left").fillna({"googlebot_hits": 0})
never = joined[(joined["Indexability"] == "Indexable") & (joined.googlebot_hits == 0)]
print(len(never), "indexable URLs with zero verified Googlebot hits in the period")
```

## Worked example 2: a Dubai classifieds site

A Dubai classifieds site (illustrative) with around two million live ads finds new listings take a week or more to be crawled. The pandas report shows most verified Googlebot requests going to `?sort=` and `?page=` combinations on search result pages and to expired ads that return 200 with "This ad has expired". Actions: disallow sort parameters, return 410 for expired ads older than 30 days (keeping a helpful 200 page with alternatives for recently expired ones), a "fresh listings" sitemap with accurate `lastmod`, and caching that halves average response time. The team measures success in the logs — share of crawl on live ads and median hours from posting to first Googlebot hit — rather than in rankings.

## Common mistakes

- Using `noindex` to save crawl budget — noindexed pages still get crawled.
- Using `crawl-delay` for Google — Googlebot ignores it.
- Analysing unverified "Googlebot" traffic that is actually scrapers.
- Treating crawl frequency as a ranking factor in itself; it is a symptom, not a lever.

## Video lecture: Crawl budget and log-file analysis

Lecture coming soon · 12 chapters · about 8 minutes. Read the full transcript below.

1. Crawl budget and log-file analysis
2. What crawl budget is
3. Who should care?
4. The big wasters
5. Logs: the ground truth
6. The procedure
7. Hands-on: pandas log analysis
8. Example 1: UK retailer
9. Example 2: Dubai classifieds (illustrative)
10. Watch me do it: do we have a crawl problem?
11. Improve efficiency; avoid mistakes
12. Recap and try this now

## Lecture transcript

### Crawl budget and log-file analysis

Crawl budget is one of the most misunderstood ideas in SEO. Small sites worry about it when they shouldn't, and huge sites ignore it when it's quietly costing them. In this lecture you'll learn what crawl budget really is, who should care, what wastes it, and how to use server logs, the only ground truth, to see what Googlebot actually does on your site, with a Python analysis you can reuse on any audit.

### What crawl budget is

Google describes crawl budget as the set of URLs Googlebot can and wants to crawl. Two parts. Crawl capacity is how much Googlebot can fetch without overloading your server; it rises with fast responses and falls with slow ones, five hundreds and four-twenty-nines. Crawl demand is how much Google wants to crawl, driven by popularity, staleness and the size of the URL inventory it knows about. Think of a delivery driver. Capacity is how many stops the van and roads allow. Demand is how many parcels are worth delivering today.

### Who should care?

Who should care? Google's own guidance says crawl budget is mainly a concern for very large sites, with hundreds of thousands to millions of URLs, sites whose content changes very often, and sites with many URLs stuck in discovered, currently not indexed. A two-hundred-page brochure site doesn't have a crawl budget problem. It has a quality or linking problem. So before you optimise crawl budget, check whether you're in the club.

### The big wasters

What wastes crawl budget? Faceted and parameter URLs are the big one: colour, size, sort and price combinations can generate millions of near-duplicates. Infinite spaces, like calendars with endless next month links, internal search results and session IDs. Redirect chains and soft four-oh-fours, where every hop or empty page costs a fetch. Duplicate host and protocol variants all returning two hundred. And slow server responses, which directly reduce capacity. Most large-site crawl problems are one of these five.

### Logs: the ground truth

Now logs. Crawlers and Search Console tell you what could be crawled. Server logs tell you what was crawled. A combined log line gives you the IP, timestamp, method, path with query string, status, bytes and user agent. If a CDN serves cached pages, get the CDN's logs, because origin logs miss cached hits. And verify before you trust. Anyone can fake a Googlebot user agent. Do a reverse DNS lookup to a Google domain, then a forward lookup that returns the same IP, or match against Google's published crawler IP ranges.

### The procedure

Here's the procedure. Collect at least two to four weeks of logs. Filter to verified search engine bots. Classify each URL by template, like product, category, blog, facet, search or asset. Join with your crawl data and Search Console. Then answer five questions. What share of bot hits go to non-indexable URLs? Are important pages crawled regularly? Are there status code problems? Are orphan URLs being crawled? And is crawl activity changing after releases?

### Hands-on: pandas log analysis

The lesson text has a pandas script that does exactly this. A regex parses each line. A small list of patterns assigns a template. A cached function verifies Googlebot with reverse and forward DNS. Then it prints the share of crawl by template, a table of status codes by template, and, after joining with your crawler export, the number of indexable URLs that got zero verified Googlebot hits in the period. That last number is often the headline of an audit.

### Example 1: UK retailer

Worked example one, simple. A mid-sized UK retailer with fifteen thousand products runs the script. Around a third of verified Googlebot hits go to sort parameters, linked from every category page. Status codes are healthy. Fix: stop linking sort parameters as crawlable links, and disallow the sort parameter in robots.txt once you've confirmed none of those URLs should be indexed. A month later, the logs show the share on sort URLs falling and product crawl frequency rising.

### Example 2: Dubai classifieds (illustrative)

Worked example two, with illustrative details. A Dubai classifieds site with around two million live ads finds new listings take a week or more to be crawled. The report shows most verified Googlebot requests going to sort and page combinations on search results, and to expired ads returning two hundred with this ad has expired. Actions: disallow sort parameters, return four-ten for ads expired more than thirty days, a fresh listings sitemap with accurate lastmod, and caching that halves response time. Success is measured in logs: share of crawl on live ads, and hours from posting to first Googlebot hit.

### Watch me do it: do we have a crawl problem?

Watch me do it. I'll answer the question, do we have a crawl problem, for a mid-sized retailer. Step one: I get two weeks of CDN logs and a fresh crawl export. Step two: I run the pandas script. It prints the share of verified Googlebot hits by template: facets forty-one percent, products twenty-two, categories fourteen, assets twelve, the rest small. All illustrative. Step three: the status crosstab. Facets are all two hundreds, and there's a cluster of five hundreds on the search template late at night. Step four: the join with the crawl. One thousand eight hundred indexable product URLs had zero verified Googlebot hits in the period. Step five: I check Crawl stats in Search Console for the same dates. Average response time rose sharply on the nights with five hundreds, which lines up with a nightly stock import. So I have two findings. Facet URLs take the biggest share of crawl, while many products go unvisited. And a nightly job slows the server enough to reduce crawl capacity. I write them as two tickets: remove crawlable links to low-value facets and disallow the sort parameter, and move the stock import off-peak or cache pages during it.

### Improve efficiency; avoid mistakes

How do you improve crawl efficiency generally? Remove internal links to parameter URLs you don't want crawled, and block truly useless patterns in robots.txt. Return four-oh-four or four-ten for gone content instead of redirecting everything to the home page. Flatten redirect chains. Keep sitemaps to canonical, indexable, two-hundred URLs. Improve response times and watch the Crawl stats report. Support three-oh-four not modified where your stack allows. Common mistakes: using noindex to save crawl budget, since noindexed pages are still crawled; using crawl-delay for Google, which ignores it; and analysing unverified Googlebot traffic.

### Recap and try this now

Recap. Crawl budget is capacity plus demand, and it matters mainly for large or fast-changing sites. The biggest wasters are facets, infinite spaces, chains, soft four-oh-fours, duplicates and slow servers. Logs are the ground truth, but only after verification. Try this now. Export two weeks of logs, run the pandas script, and write down two numbers: the share of verified Googlebot hits on non-indexable templates, and the count of indexable URLs with zero hits. Those two numbers tell you whether you have a crawl problem at all.

## Key takeaways

- Crawl budget = crawl capacity (server health) + crawl demand (Google's interest); mostly a large-site concern.
- Server logs show what bots actually crawled; always verify Googlebot by reverse DNS or published IP ranges.
- Segment logs by template and join with crawl and Search Console data to find waste.
- noindex does not save crawl budget and Google ignores crawl-delay.

## Try it

Export two weeks of logs (or a CDN log sample), verify Googlebot hits, and build a pivot of hits by page template and status code.

- [Previous: JavaScript SEO: SPAs, SSR and prerendering](https://optimizeall.com/learn/technical-seo-mastery/javascript-seo-spas-ssr-prerendering)
- [Next: robots.txt, meta robots and X-Robots-Tag](https://optimizeall.com/learn/technical-seo-mastery/robots-txt-meta-robots-x-robots-tag)
- [All lessons of Technical SEO Mastery](https://optimizeall.com/learn/technical-seo-mastery)
