---
title: "Configuring the crawl like Googlebot | Optimize All Academy"
description: "Why configuration decides the quality of the audit The same crawler produces completely different results depending on its settings. A text-only crawl of…"
url: https://optimizeall.com/learn/technical-seo-audit-in-practice/configuring-the-crawl
updated: 2026-10-05
---

Technical SEO Audit Workshop · Scoping the audit and setting up · lesson 2 of 14 · 13 min

# Configuring the crawl like Googlebot

## Why configuration decides the quality of the audit

The same crawler produces completely different results depending on its settings. A text-only crawl of a JavaScript storefront may report "no content"; a crawl that ignores robots.txt may flag issues Google never sees. Your job is to configure a crawl that approximates **how Googlebot experiences the site**, and to record those settings so the crawl can be repeated after fixes.

## The configuration sheet

For Kiran Home, the crawl configuration is:

| Setting | Choice | Reason |
|---|---|---|
| Start URL | `https://www.kiranhome.example/` | The canonical host |
| Additional seeds | All XML sitemaps, list of old URLs from the launch redirect map | Find orphans and test redirects |
| User agent | Googlebot Smartphone | Mobile-first indexing |
| Rendering | JavaScript rendering on | Headless React storefront |
| Robots.txt | Respect | See what Google may crawl; run a second crawl ignoring it to see what is hidden |
| Crawl limits | No depth limit; cap at a sensible URL count, then review patterns | Detect infinite spaces |
| Speed | Modest concurrent threads, agreed with dev team | Avoid overloading origin/CDN rules |
| Parameters | Crawl all initially | Discover facet and tracking parameter problems |
| Extraction | Canonical, meta robots, X-Robots-Tag, hreflang, JSON-LD, H1, word count, price element | Template checks |
| Integrations | Search Console API, GA4 | Join traffic and index data to URLs |

Two practical safeguards:

- **Allow-list your crawler** with the client's CDN or bot-protection tool, or you will crawl a wall of challenge pages. Ask the developers to create a WAF rule for your crawler's IP or user agent for the audit window.
- **Crawl staging separately** with authentication and a noindex check, not the other way round.

## Custom extraction: check what matters to the business

Default reports catch generic issues. Custom extraction answers client-specific questions. For Kiran Home we want to know if product prices and "Add to basket" buttons exist in the rendered HTML, and whether JSON-LD price matches the visible price.

```text
Extractor 1 — visible price     CSS: .product-price__amount
Extractor 2 — JSON-LD blocks     XPath: //script[@type="application/ld+json"]
Extractor 3 — add to basket      CSS: button[data-action="add-to-basket"]
Extractor 4 — breadcrumb links   CSS: nav.breadcrumb a
```

## Segmentation before you crawl

Configure segments by URL pattern so every report can be read per template:

```text
product    ^/(en-pk|en-ae|ar-ae|en-gb)/products/
category   ^/(en-pk|en-ae|ar-ae|en-gb)/collections/
blog       ^/(en-pk|en-ae|ar-ae|en-gb)/journal/
facet      \?(.*&)?(colour|size|price|sort)=
search     ^/(en-pk|en-ae|ar-ae|en-gb)/search
```

Segmenting early is the difference between "3,200 pages have duplicate titles" and "all 3,100 duplicate titles are facet URLs, which should not be crawlable anyway".

## Two crawls are better than one

Run:

1. **Rendered crawl** — JavaScript on, respecting robots.txt, Googlebot Smartphone. This is your primary dataset.
2. **Raw crawl** — JavaScript off. Comparing the two reveals what depends on rendering: links, canonicals, content, structured data.

Optionally, a **list crawl** of old URLs from the pre-relaunch site to test redirects, and a **sitemap-only crawl** to test sitemap hygiene.

## Sanity-check the first minutes

Watch the crawl as it starts, and stop to fix config if you see:

- All URLs returning 403 or a challenge page (bot protection).
- The crawl exploding into parameter URLs within minutes (facet trap — cap it, note it as a finding, and exclude for the main crawl).
- Rendered pages with no content (rendering timeouts; increase the AJAX timeout or check blocked resources).
- Redirects to a login or geo page (geo-redirects by IP — itself a finding for an international site).

## Recording the configuration

Save the crawler's configuration file and add a short note to your findings workbook: crawler version, date, settings, seeds, exclusions. When you re-crawl for verification in Module 5, you will use the **identical** configuration, so any difference is caused by fixes and not by settings.

## Hands-on: a repeatable headless crawl

Desktop crawlers can run from the command line, which makes crawls repeatable and schedulable. For example, Screaming Frog SEO Spider (licensed) supports headless crawls with a saved configuration:

```bash
screamingfrogseospider --headless \
  --crawl https://www.kiranhome.example/ \
  --config ./kiranhome-v1.2.seospiderconfig \
  --save-crawl --output-folder ./crawls/2026-09-15 \
  --export-tabs "Internal:All,Response Codes:Client Error (4xx),Canonicals:All,Hreflang:All" \
  --timestamped-output
```

Check `--help` for your version's exact flags. Sitebulb, Lumar, JetOctopus and Oncrawl offer equivalent scheduled or cloud crawls. Whatever the tool, store the config file in version control with the audit and name exports by date and config version.

## Checking the crawl isn't being blocked

Before the full run, test a few URLs as the crawler will request them:

```bash
for p in / /en-ae/products/brass-lantern-large/ /en-gb/collections/cushions/; do
  curl -s -o /dev/null -w "%{http_code} %{size_download}B $p\n" \
    -A "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
    "https://www.kiranhome.example$p"
done
```

A `403`, `503` or a suspiciously small size means a challenge page. Remember: a spoofed Googlebot user agent from your IP is **not** Googlebot — CDNs may treat it differently from the real crawler, which is why you ask for an allow-list rule for your crawler during the audit window.

## Common mistakes

- Crawling with a desktop user agent for a mobile-first world.
- Leaving rendering off on a JavaScript site, or on with a timeout so short that content never loads.
- Crawling from a blocked IP and reporting "site returns 403" as a finding.
- Not seeding with old URLs, so redirect problems from the relaunch are never tested.

## Video lecture: Configuring the crawl like Googlebot

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Configuring the crawl
2. The configuration sheet
3. Extraction + safeguards
4. Custom extraction
5. Segment first
6. Four crawls
7. Hands-on: headless, versioned crawls
8. Watch the first minutes
9. Example 1: Kiran Home kick-off (illustrative)
10. Example 2: the too-short timeout (illustrative)
11. Watch me do it: launching the crawl
12. Record it; avoid mistakes
13. Recap and try this now

## Lecture transcript

### Configuring the crawl

Before you look at a single finding, look at your crawler settings. They decide what you'll find. The same tool gives completely different results depending on configuration. A text-only crawl of a JavaScript storefront may report no content. A crawl blocked by bot protection reports a site full of four-oh-threes. In this lecture you'll configure a crawl for Kiran Home that approximates how Googlebot experiences the site, run it in a repeatable way, and record every setting so you can re-crawl identically after fixes.

### The configuration sheet

Here's the configuration sheet for Kiran Home. Start at the canonical host. Seed the crawl with all XML sitemaps and the list of old URLs from the launch redirect map, so you find orphans and test redirects. Use a Googlebot smartphone user agent, because of mobile-first indexing. Turn JavaScript rendering on, because it's a headless React storefront. Respect robots.txt, and optionally run a second crawl ignoring it to see what's hidden. No depth limit, but cap the URL count and review patterns. Modest speed, agreed with developers. Crawl all parameters initially.

### Extraction + safeguards

Extraction and integrations. Extract canonical, meta robots, X-Robots-Tag, hreflang, JSON-LD, H1, word count and the price element. Connect the Search Console API and GA4, so traffic and index data join the crawl by URL. Then two safeguards. First, get your crawler allow-listed with the CDN or bot protection for the audit window, otherwise you'll crawl a wall of challenge pages. Second, crawl staging separately with authentication, and check that it's noindexed, not the other way round.

### Custom extraction

Custom extraction checks what matters to this business. For Kiran Home, we want to know whether product prices and add to basket buttons exist in the rendered HTML, and whether the JSON-LD price matches the visible price. So we extract the visible price with a CSS selector, all JSON-LD blocks with XPath, the add to basket button, and breadcrumb links. Default reports catch generic issues. Custom extraction answers the client's questions.

### Segment first

Segment before you crawl. Configure segments by URL pattern: product, category, blog, facet and search, using regular expressions like the ones in the lesson text. Here's why that matters. Without segments, you report three thousand two hundred duplicate titles. With segments, you report that all three thousand one hundred duplicate titles are facet URLs, which shouldn't be crawlable anyway. Same data, completely different conclusion.

### Four crawls

Two crawls are better than one. The rendered crawl, JavaScript on, respecting robots.txt, with a Googlebot smartphone agent, is your primary dataset. The raw crawl, JavaScript off, shows what depends on rendering: links, canonicals, content and structured data. Add a list crawl of old URLs from the pre-relaunch site to test redirects, and a sitemap-only crawl to test sitemap hygiene. Four views of the same site, each answering a different question.

### Hands-on: headless, versioned crawls

Hands-on. Crawlers can run headless from the command line, which makes crawls repeatable and schedulable. The lesson text shows a Screaming Frog example with a saved configuration file, an output folder by date, selected exports and timestamped output. Check the help for your version's exact flags, and other crawlers offer equivalent scheduled or cloud crawls. Store the config file in version control with the audit. There's also a quick curl loop to check key URLs aren't being blocked before you start.

### Watch the first minutes

Watch the first minutes of the crawl and stop to fix the config if you see these signs. All URLs returning four-oh-three or a challenge page means bot protection. The crawl exploding into parameter URLs within minutes means a facet trap, so cap it, note it as a finding, and exclude it from the main crawl. Rendered pages with no content means rendering timeouts, so increase the timeout or check blocked resources. And redirects to a login or geo page means IP-based geo-redirects, which is itself a finding for an international site.

### Example 1: Kiran Home kick-off (illustrative)

Worked example one, simple, and illustrative. On the Kiran Home kick-off crawl, the first two hundred URLs all return five-oh-three. The CDN's bot fight mode is challenging the crawler. The developers add a temporary allow rule for the audit crawler's IP. But note the second-order question for the audit: does the same rule set challenge Bingbot or AI search bots? That becomes evidence for hypothesis five.

### Example 2: the too-short timeout (illustrative)

Worked example two, also illustrative. A Doha fashion retailer's crawl finds only seven hundred URLs on a site with six thousand products. The rendered crawl used a two-second timeout, and the category grid loads after three. Increasing the timeout and checking resources reveals the full catalogue, and a genuine finding: the category API is slow, which also hurts real users and Google's rendering. Configuration problems often hide real problems.

### Watch me do it: launching the crawl

Watch me do it. I'll set up and launch the Kiran Home crawl. Step one: in the crawler, I set the user agent to Googlebot Smartphone. Step two: I switch rendering to JavaScript, and raise the render timeout because the category grid loads slowly. Step three: under robots settings, I leave respect robots.txt on. Step four: I add custom extraction rules: the visible price by CSS selector, JSON-LD blocks by XPath, the add to basket button, and breadcrumb links. Step five: I add segments with the regex patterns for product, category, blog, facet and search. Step six: I connect the Search Console and GA4 APIs. Step seven: I load the seeds: the home page, the sitemap index, and a list of old URLs from the launch redirect map. Step eight: before the full run, I use the curl loop to test three URLs with the same user agent. All two hundreds, so the developers' temporary allow rule is working. Step nine: I start the crawl and watch the first five minutes. After two thousand URLs, the facet segment is growing fast, so I cap facet URLs, note the trap as a finding, and continue. Step ten: I save the configuration as version one point two and record the settings in the workbook.

### Record it; avoid mistakes

Recording the configuration. Save the crawler's configuration file, and note the crawler version, date, settings, seeds and exclusions in your findings workbook. When you re-crawl for verification later in this workshop, you'll use the identical configuration, so any difference comes from fixes and not from settings. Common mistakes: crawling with a desktop user agent, leaving rendering off on a JavaScript site, crawling from a blocked IP and reporting site returns four-oh-three as a finding, and not seeding with old URLs.

### Recap and try this now

Recap. Configure the crawl to approximate Googlebot: smartphone agent, rendering on, robots respected, seeded with sitemaps and old URLs, segmented, and allow-listed. Run rendered and raw crawls, plus list crawls. Record everything and re-use it. Try this now. Write a configuration sheet for a site you know, with a reason for every setting, and run a quick curl check on three key URLs with a Googlebot smartphone user agent to see if anything blocks you.

## Video transcript

Before you look at a single finding, look at your crawler settings. They decide what you'll find. Our client, Kiran Home, runs a React storefront with server-side rendering, four market folders, and a CDN with bot protection. So our configuration has to reflect that. First, the user agent: Googlebot Smartphone, because Google indexes the mobile version. Second, rendering: switched on, with a sensible timeout, because parts of the page depend on JavaScript. Third, seeds: not just the home page, but every XML sitemap and the list of old URLs from before the relaunch. That's how we'll test the redirects and find orphans. Next, segments. We tell the crawler which URL patterns are products, categories, blog posts, facets and search pages. Every report can then be read per template, which turns noise into a story. Then custom extraction. We pull the visible price, the JSON-LD block and the add-to-basket button, so we can check that the things that make money actually exist in the rendered HTML. Finally, watch the first few minutes. If everything returns 403, the CDN is blocking you. Get allow-listed. If the URL count explodes into filter parameters, you've found your first finding. Note it, cap it, and carry on. And save the configuration. When we verify fixes later, we'll use exactly the same settings.

## Key takeaways

- Configure the crawl to approximate Googlebot Smartphone, with JavaScript rendering on JS sites.
- Seed with sitemaps and old URLs; integrate Search Console and analytics data.
- Use custom extraction and URL-pattern segments to answer business-specific questions per template.
- Record the configuration so verification crawls are directly comparable.

## Try it

Create a crawl configuration sheet for a site you manage, including user agent, rendering, seeds, segments and three custom extractions.

- [Previous: Meet the client and scope the audit](https://optimizeall.com/learn/technical-seo-audit-in-practice/meet-the-client-and-scope-the-audit)
- [Next: Reading the crawl: indexability, status and templates](https://optimizeall.com/learn/technical-seo-audit-in-practice/reading-the-crawl-findings)
- [All lessons of Technical SEO Audit Workshop](https://optimizeall.com/learn/technical-seo-audit-in-practice)
