Technical SEO Audit WorkshopScoping the audit and setting up · Lesson 2 of 14

Configuring the crawl like Googlebot

Video lesson · 13 min · 8 min lecture

Video lecture

Configuring the crawl like Googlebot

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Configuring the crawl

  • Settings decide findings
  • Approximate Googlebot
  • Repeatable and recorded

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why configuration decides the quality of the audit

The same crawler produces completely different results depending on its settings. A text-only crawl of a JavaScript storefront may report "no content"; a crawl that ignores robots.txt may flag issues Google never sees. Your job is to configure a crawl that approximates how Googlebot experiences the site, and to record those settings so the crawl can be repeated after fixes.

The configuration sheet

For Kiran Home, the crawl configuration is:

SettingChoiceReason
Start URLhttps://www.kiranhome.example/The canonical host
Additional seedsAll XML sitemaps, list of old URLs from the launch redirect mapFind orphans and test redirects
User agentGooglebot SmartphoneMobile-first indexing
RenderingJavaScript rendering onHeadless React storefront
Robots.txtRespectSee what Google may crawl; run a second crawl ignoring it to see what is hidden
Crawl limitsNo depth limit; cap at a sensible URL count, then review patternsDetect infinite spaces
SpeedModest concurrent threads, agreed with dev teamAvoid overloading origin/CDN rules
ParametersCrawl all initiallyDiscover facet and tracking parameter problems
ExtractionCanonical, meta robots, X-Robots-Tag, hreflang, JSON-LD, H1, word count, price elementTemplate checks
IntegrationsSearch Console API, GA4Join traffic and index data to URLs

Two practical safeguards:

  • Allow-list your crawler with the client's CDN or bot-protection tool, or you will crawl a wall of challenge pages. Ask the developers to create a WAF rule for your crawler's IP or user agent for the audit window.
  • Crawl staging separately with authentication and a noindex check, not the other way round.

Custom extraction: check what matters to the business

Default reports catch generic issues. Custom extraction answers client-specific questions. For Kiran Home we want to know if product prices and "Add to basket" buttons exist in the rendered HTML, and whether JSON-LD price matches the visible price.

Extractor 1 — visible price     CSS: .product-price__amount
Extractor 2 — JSON-LD blocks     XPath: //script[@type="application/ld+json"]
Extractor 3 — add to basket      CSS: button[data-action="add-to-basket"]
Extractor 4 — breadcrumb links   CSS: nav.breadcrumb a

Segmentation before you crawl

Configure segments by URL pattern so every report can be read per template:

product    ^/(en-pk|en-ae|ar-ae|en-gb)/products/
category   ^/(en-pk|en-ae|ar-ae|en-gb)/collections/
blog       ^/(en-pk|en-ae|ar-ae|en-gb)/journal/
facet      \?(.*&)?(colour|size|price|sort)=
search     ^/(en-pk|en-ae|ar-ae|en-gb)/search

Segmenting early is the difference between "3,200 pages have duplicate titles" and "all 3,100 duplicate titles are facet URLs, which should not be crawlable anyway".

Two crawls are better than one

Run:

  1. Rendered crawl — JavaScript on, respecting robots.txt, Googlebot Smartphone. This is your primary dataset.
  2. Raw crawl — JavaScript off. Comparing the two reveals what depends on rendering: links, canonicals, content, structured data.

Optionally, a list crawl of old URLs from the pre-relaunch site to test redirects, and a sitemap-only crawl to test sitemap hygiene.

Sanity-check the first minutes

Watch the crawl as it starts, and stop to fix config if you see:

  • All URLs returning 403 or a challenge page (bot protection).
  • The crawl exploding into parameter URLs within minutes (facet trap — cap it, note it as a finding, and exclude for the main crawl).
  • Rendered pages with no content (rendering timeouts; increase the AJAX timeout or check blocked resources).
  • Redirects to a login or geo page (geo-redirects by IP — itself a finding for an international site).

Recording the configuration

Save the crawler's configuration file and add a short note to your findings workbook: crawler version, date, settings, seeds, exclusions. When you re-crawl for verification in Module 5, you will use the identical configuration, so any difference is caused by fixes and not by settings.

Hands-on: a repeatable headless crawl

Desktop crawlers can run from the command line, which makes crawls repeatable and schedulable. For example, Screaming Frog SEO Spider (licensed) supports headless crawls with a saved configuration:

screamingfrogseospider --headless \
  --crawl https://www.kiranhome.example/ \
  --config ./kiranhome-v1.2.seospiderconfig \
  --save-crawl --output-folder ./crawls/2026-09-15 \
  --export-tabs "Internal:All,Response Codes:Client Error (4xx),Canonicals:All,Hreflang:All" \
  --timestamped-output

Check --help for your version's exact flags. Sitebulb, Lumar, JetOctopus and Oncrawl offer equivalent scheduled or cloud crawls. Whatever the tool, store the config file in version control with the audit and name exports by date and config version.

Checking the crawl isn't being blocked

Before the full run, test a few URLs as the crawler will request them:

for p in / /en-ae/products/brass-lantern-large/ /en-gb/collections/cushions/; do
  curl -s -o /dev/null -w "%{http_code} %{size_download}B $p\n" \
    -A "Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
    "https://www.kiranhome.example$p"
done

A 403, 503 or a suspiciously small size means a challenge page. Remember: a spoofed Googlebot user agent from your IP is not Googlebot — CDNs may treat it differently from the real crawler, which is why you ask for an allow-list rule for your crawler during the audit window.

Common mistakes

  • Crawling with a desktop user agent for a mobile-first world.
  • Leaving rendering off on a JavaScript site, or on with a timeout so short that content never loads.
  • Crawling from a blocked IP and reporting "site returns 403" as a finding.
  • Not seeding with old URLs, so redirect problems from the relaunch are never tested.

Key takeaways

  • Configure the crawl to approximate Googlebot Smartphone, with JavaScript rendering on JS sites.
  • Seed with sitemaps and old URLs; integrate Search Console and analytics data.
  • Use custom extraction and URL-pattern segments to answer business-specific questions per template.
  • Record the configuration so verification crawls are directly comparable.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why run both a rendered and a raw (JavaScript off) crawl?
  2. Your crawler receives 403s on every URL. What is the most likely explanation?
  3. Why seed the crawl with old pre-relaunch URLs?

Put it into practice

Create a crawl configuration sheet for a site you manage, including user agent, rendering, seeds, segments and three custom extractions.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.