Technical SEO Audit WorkshopScoping the audit and setting up · Lesson 2 of 14
Configuring the crawl like Googlebot
Video lecture
Configuring the crawl like Googlebot
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Configuring the crawl
Before you look at a single finding, look at your crawler settings. They decide what you'll find. The same tool gives completely different results depending on configuration. A text-only crawl of a JavaScript storefront may report no content. A crawl blocked by bot protection reports a site full of four-oh-threes. In this lecture you'll configure a crawl for Kiran Home that approximates how Googlebot experiences the site, run it in a repeatable way, and record every setting so you can re-crawl identically after fixes.
0:37 The configuration sheet
Here's the configuration sheet for Kiran Home. Start at the canonical host. Seed the crawl with all XML sitemaps and the list of old URLs from the launch redirect map, so you find orphans and test redirects. Use a Googlebot smartphone user agent, because of mobile-first indexing. Turn JavaScript rendering on, because it's a headless React storefront. Respect robots.txt, and optionally run a second crawl ignoring it to see what's hidden. No depth limit, but cap the URL count and review patterns. Modest speed, agreed with developers. Crawl all parameters initially.
1:17 Extraction + safeguards
Extraction and integrations. Extract canonical, meta robots, X-Robots-Tag, hreflang, JSON-LD, H1, word count and the price element. Connect the Search Console API and GA4, so traffic and index data join the crawl by URL. Then two safeguards. First, get your crawler allow-listed with the CDN or bot protection for the audit window, otherwise you'll crawl a wall of challenge pages. Second, crawl staging separately with authentication, and check that it's noindexed, not the other way round.
1:50 Custom extraction
Custom extraction checks what matters to this business. For Kiran Home, we want to know whether product prices and add to basket buttons exist in the rendered HTML, and whether the JSON-LD price matches the visible price. So we extract the visible price with a CSS selector, all JSON-LD blocks with XPath, the add to basket button, and breadcrumb links. Default reports catch generic issues. Custom extraction answers the client's questions.
2:21 Segment first
Segment before you crawl. Configure segments by URL pattern: product, category, blog, facet and search, using regular expressions like the ones in the lesson text. Here's why that matters. Without segments, you report three thousand two hundred duplicate titles. With segments, you report that all three thousand one hundred duplicate titles are facet URLs, which shouldn't be crawlable anyway. Same data, completely different conclusion.
2:49 Four crawls
Two crawls are better than one. The rendered crawl, JavaScript on, respecting robots.txt, with a Googlebot smartphone agent, is your primary dataset. The raw crawl, JavaScript off, shows what depends on rendering: links, canonicals, content and structured data. Add a list crawl of old URLs from the pre-relaunch site to test redirects, and a sitemap-only crawl to test sitemap hygiene. Four views of the same site, each answering a different question.
3:20 Hands-on: headless, versioned crawls
Hands-on. Crawlers can run headless from the command line, which makes crawls repeatable and schedulable. The lesson text shows a Screaming Frog example with a saved configuration file, an output folder by date, selected exports and timestamped output. Check the help for your version's exact flags, and other crawlers offer equivalent scheduled or cloud crawls. Store the config file in version control with the audit. There's also a quick curl loop to check key URLs aren't being blocked before you start.
3:55 Watch the first minutes
Watch the first minutes of the crawl and stop to fix the config if you see these signs. All URLs returning four-oh-three or a challenge page means bot protection. The crawl exploding into parameter URLs within minutes means a facet trap, so cap it, note it as a finding, and exclude it from the main crawl. Rendered pages with no content means rendering timeouts, so increase the timeout or check blocked resources. And redirects to a login or geo page means IP-based geo-redirects, which is itself a finding for an international site.
4:35 Example 1: Kiran Home kick-off (illustrative)
Worked example one, simple, and illustrative. On the Kiran Home kick-off crawl, the first two hundred URLs all return five-oh-three. The CDN's bot fight mode is challenging the crawler. The developers add a temporary allow rule for the audit crawler's IP. But note the second-order question for the audit: does the same rule set challenge Bingbot or AI search bots? That becomes evidence for hypothesis five.
5:04 Example 2: the too-short timeout (illustrative)
Worked example two, also illustrative. A Doha fashion retailer's crawl finds only seven hundred URLs on a site with six thousand products. The rendered crawl used a two-second timeout, and the category grid loads after three. Increasing the timeout and checking resources reveals the full catalogue, and a genuine finding: the category API is slow, which also hurts real users and Google's rendering. Configuration problems often hide real problems.
5:34 Watch me do it: launching the crawl
Watch me do it. I'll set up and launch the Kiran Home crawl. Step one: in the crawler, I set the user agent to Googlebot Smartphone. Step two: I switch rendering to JavaScript, and raise the render timeout because the category grid loads slowly. Step three: under robots settings, I leave respect robots.txt on. Step four: I add custom extraction rules: the visible price by CSS selector, JSON-LD blocks by XPath, the add to basket button, and breadcrumb links. Step five: I add segments with the regex patterns for product, category, blog, facet and search. Step six: I connect the Search Console and GA4 APIs. Step seven: I load the seeds: the home page, the sitemap index, and a list of old URLs from the launch redirect map. Step eight: before the full run, I use the curl loop to test three URLs with the same user agent. All two hundreds, so the developers' temporary allow rule is working. Step nine: I start the crawl and watch the first five minutes. After two thousand URLs, the facet segment is growing fast, so I cap facet URLs, note the trap as a finding, and continue. Step ten: I save the configuration as version one point two and record the settings in the workbook.
7:06 Record it; avoid mistakes
Recording the configuration. Save the crawler's configuration file, and note the crawler version, date, settings, seeds and exclusions in your findings workbook. When you re-crawl for verification later in this workshop, you'll use the identical configuration, so any difference comes from fixes and not from settings. Common mistakes: crawling with a desktop user agent, leaving rendering off on a JavaScript site, crawling from a blocked IP and reporting site returns four-oh-three as a finding, and not seeding with old URLs.
7:41 Recap and try this now
Recap. Configure the crawl to approximate Googlebot: smartphone agent, rendering on, robots respected, seeded with sitemaps and old URLs, segmented, and allow-listed. Run rendered and raw crawls, plus list crawls. Record everything and re-use it. Try this now. Write a configuration sheet for a site you know, with a reason for every setting, and run a quick curl check on three key URLs with a Googlebot smartphone user agent to see if anything blocks you.
Why configuration decides the quality of the audit
The same crawler produces completely different results depending on its settings. A text-only crawl of a JavaScript storefront may report "no content"; a crawl that ignores robots.txt may flag issues Google never sees. Your job is to configure a crawl that approximates how Googlebot experiences the site, and to record those settings so the crawl can be repeated after fixes.
The configuration sheet
For Kiran Home, the crawl configuration is:
| Setting | Choice | Reason |
|---|---|---|
| Start URL | https://www.kiranhome.example/ | The canonical host |
| Additional seeds | All XML sitemaps, list of old URLs from the launch redirect map | Find orphans and test redirects |
| User agent | Googlebot Smartphone | Mobile-first indexing |
| Rendering | JavaScript rendering on | Headless React storefront |
| Robots.txt | Respect | See what Google may crawl; run a second crawl ignoring it to see what is hidden |
| Crawl limits | No depth limit; cap at a sensible URL count, then review patterns | Detect infinite spaces |
| Speed | Modest concurrent threads, agreed with dev team | Avoid overloading origin/CDN rules |
| Parameters | Crawl all initially | Discover facet and tracking parameter problems |
| Extraction | Canonical, meta robots, X-Robots-Tag, hreflang, JSON-LD, H1, word count, price element | Template checks |
| Integrations | Search Console API, GA4 | Join traffic and index data to URLs |
Two practical safeguards:
- Allow-list your crawler with the client's CDN or bot-protection tool, or you will crawl a wall of challenge pages. Ask the developers to create a WAF rule for your crawler's IP or user agent for the audit window.
- Crawl staging separately with authentication and a noindex check, not the other way round.
Custom extraction: check what matters to the business
Default reports catch generic issues. Custom extraction answers client-specific questions. For Kiran Home we want to know if product prices and "Add to basket" buttons exist in the rendered HTML, and whether JSON-LD price matches the visible price.
Extractor 1 — visible price CSS: .product-price__amount
Extractor 2 — JSON-LD blocks XPath: //script[@type="application/ld+json"]
Extractor 3 — add to basket CSS: button[data-action="add-to-basket"]
Extractor 4 — breadcrumb links CSS: nav.breadcrumb aSegmentation before you crawl
Configure segments by URL pattern so every report can be read per template:
product ^/(en-pk|en-ae|ar-ae|en-gb)/products/
category ^/(en-pk|en-ae|ar-ae|en-gb)/collections/
blog ^/(en-pk|en-ae|ar-ae|en-gb)/journal/
facet \?(.*&)?(colour|size|price|sort)=
search ^/(en-pk|en-ae|ar-ae|en-gb)/searchSegmenting early is the difference between "3,200 pages have duplicate titles" and "all 3,100 duplicate titles are facet URLs, which should not be crawlable anyway".
Two crawls are better than one
Run:
- Rendered crawl — JavaScript on, respecting robots.txt, Googlebot Smartphone. This is your primary dataset.
- Raw crawl — JavaScript off. Comparing the two reveals what depends on rendering: links, canonicals, content, structured data.
Optionally, a list crawl of old URLs from the pre-relaunch site to test redirects, and a sitemap-only crawl to test sitemap hygiene.
Sanity-check the first minutes
Watch the crawl as it starts, and stop to fix config if you see:
- All URLs returning 403 or a challenge page (bot protection).
- The crawl exploding into parameter URLs within minutes (facet trap — cap it, note it as a finding, and exclude for the main crawl).
- Rendered pages with no content (rendering timeouts; increase the AJAX timeout or check blocked resources).
- Redirects to a login or geo page (geo-redirects by IP — itself a finding for an international site).
Recording the configuration
Save the crawler's configuration file and add a short note to your findings workbook: crawler version, date, settings, seeds, exclusions. When you re-crawl for verification in Module 5, you will use the identical configuration, so any difference is caused by fixes and not by settings.
Hands-on: a repeatable headless crawl
Desktop crawlers can run from the command line, which makes crawls repeatable and schedulable. For example, Screaming Frog SEO Spider (licensed) supports headless crawls with a saved configuration:
screamingfrogseospider --headless \
--crawl https://www.kiranhome.example/ \
--config ./kiranhome-v1.2.seospiderconfig \
--save-crawl --output-folder ./crawls/2026-09-15 \
--export-tabs "Internal:All,Response Codes:Client Error (4xx),Canonicals:All,Hreflang:All" \
--timestamped-outputCheck --help for your version's exact flags. Sitebulb, Lumar, JetOctopus and Oncrawl offer equivalent scheduled or cloud crawls. Whatever the tool, store the config file in version control with the audit and name exports by date and config version.
Checking the crawl isn't being blocked
Before the full run, test a few URLs as the crawler will request them:
for p in / /en-ae/products/brass-lantern-large/ /en-gb/collections/cushions/; do
curl -s -o /dev/null -w "%{http_code} %{size_download}B $p\n" \
-A "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
"https://www.kiranhome.example$p"
doneA 403, 503 or a suspiciously small size means a challenge page. Remember: a spoofed Googlebot user agent from your IP is not Googlebot — CDNs may treat it differently from the real crawler, which is why you ask for an allow-list rule for your crawler during the audit window.
Common mistakes
- Crawling with a desktop user agent for a mobile-first world.
- Leaving rendering off on a JavaScript site, or on with a timeout so short that content never loads.
- Crawling from a blocked IP and reporting "site returns 403" as a finding.
- Not seeding with old URLs, so redirect problems from the relaunch are never tested.
Key takeaways
- Configure the crawl to approximate Googlebot Smartphone, with JavaScript rendering on JS sites.
- Seed with sitemaps and old URLs; integrate Search Console and analytics data.
- Use custom extraction and URL-pattern segments to answer business-specific questions per template.
- Record the configuration so verification crawls are directly comparable.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create a crawl configuration sheet for a site you manage, including user agent, rendering, seeds, segments and three custom extractions.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.