Skip to content

Technical SEO Audit Workshop · Running the crawl and gathering evidence · lesson 3 of 14 · 14 min

Reading the crawl: indexability, status and templates

From thousands of rows to a handful of patterns

The Kiran Home rendered crawl has finished. The crawler found about 61,000 URLs (illustrative). Only around 4,500 are products, categories, blog posts and static pages the business actually wants indexed. The first job is to explain that gap.

Step 1: Indexability overview by segment

Pivot the crawl by segment and indexability:

| Segment | URLs | Indexable | Non-indexable reasons | |---|---|---|---| | Product | 16,200 | 3,900 | Canonicalised to another market (see below) | | Category | 720 | 690 | A few noindexed | | Facet | 41,800 | 41,800 | Nothing prevents indexing | | Search | 1,900 | 1,900 | Internal search results crawlable and indexable | | Blog | 1,200 | 290 | Canonicalised to /en-gb/ versions | | Static | 80 | 70 | Mixed |

All figures are illustrative for this exercise.

Three patterns jump out immediately:

  1. Facets and internal search are crawlable and indexable — an infinite-space problem.
  2. Product and blog canonicals point across markets, suggesting only one market version is treated as canonical.
  3. There are four market versions of each product (roughly 4,000 × 4), which is expected — but most are canonicalised away.

Step 2: Investigate each pattern with samples

Never report a pattern without opening real examples. For pattern 2, open five products in each market:

<!-- https://www.kiranhome.example/en-ae/products/brass-lantern-large -->
<link rel="canonical" href="https://www.kiranhome.example/en-gb/products/brass-lantern-large">
<link rel="alternate" hreflang="en-ae" href="https://www.kiranhome.example/en-ae/products/brass-lantern-large">
<link rel="alternate" hreflang="en-gb" href="https://www.kiranhome.example/en-gb/products/brass-lantern-large">

Finding: every non-UK product canonicalises to the UK version while declaring hreflang. The canonical tells Google the UAE and Pakistan pages are duplicates of the UK page, so hreflang cannot work. That fits the hypothesis that international signals broke at relaunch — and explains why UAE users might see UK pricing.

Step 3: Status code and redirect review

Filter the list crawl of 5,600 pre-relaunch URLs:

| Result | Count | Notes | |---|---|---| | Single 301 to relevant 200 page | 3,100 | Good | | 301 chain (2+ hops) | 900 | Old .html → new slug → market folder | | 301 to home page | 700 | Old blog category and tag URLs | | 404 | 850 | Mostly old product URLs with a different slug pattern | | 200 (old URL still live) | 50 | Legacy pages the platform kept |

The 850 old product 404s need checking against backlinks and Search Console clicks: those with links or traffic are high priority.

Step 4: Template-level checks

For each key template, open a representative URL and review the rendered HTML against a checklist:

  • Title and H1: present, unique, descriptive?
  • Canonical: self-referencing on the correct market?
  • Robots directives: none unexpected in meta or headers?
  • Main content in rendered HTML; product description, price and reviews present?
  • Internal links: breadcrumbs, related products and category links as real anchors?
  • JSON-LD: valid, matching visible data?

For Kiran Home, category pages render only the first 24 products; "Load more" uses a button with no paginated URLs. Products beyond the first 24 in each category have few internal links — a discovery issue.

Step 5: Content signals the crawler can reveal

  • Duplicate titles mostly on facets (a symptom, not a separate issue).
  • Thin pages: 400 products with fewer than 30 words of description.
  • Near-duplicate market pages where only currency changes — acceptable for regional pages if hreflang is correct, but worth noting.

Recording findings properly

Each finding in the workbook gets:

ID:        F-03
Title:     Non-UK product pages canonicalise to UK versions
Evidence:  16,200 product URLs crawled; 12,300 canonicalise to /en-gb/ equivalents.
           Examples: 5 URLs listed. URL Inspection: Google-selected canonical = /en-gb/.
Affected:  All product and blog templates in /en-pk/, /en-ae/, /ar-ae/
Hypothesis link: H3 (international signals)
Severity:  To be scored in triage

Hands-on: the segment × indexability pivot in pandas

import re
import pandas as pd
df = pd.read_csv("internal_all.csv")                 # crawler export
SEG = [("facet", r"\?(.*&)?(colour|size|price|sort)="), ("search", r"/search"),
       ("product", r"/products/"), ("category", r"/collections/"), ("blog", r"/journal/")]
def seg(u):
    for name, pat in SEG:
        if re.search(pat, u):
            return name
    return "static"
df["segment"] = df["Address"].map(seg)
df["market"] = df["Address"].str.extract(r"/(en-pk|en-ae|ar-ae|en-gb)/")[0].fillna("none")

pivot = pd.pivot_table(df, index="segment", columns="Indexability", values="Address",
                       aggfunc="count", fill_value=0, margins=True)
print(pivot)

# Where do non-indexable product URLs point their canonicals?
prod = df[(df.segment == "product") & (df["Indexability Status"] == "Canonicalised")].copy()
prod["canon_market"] = prod["Canonical Link Element 1"].str.extract(r"/(en-pk|en-ae|ar-ae|en-gb)/")[0]
print(pd.crosstab(prod.market, prod.canon_market))

The crosstab makes the Kiran Home pattern unmistakable: product URLs in /en-pk/, /en-ae/ and /ar-ae/ canonicalise to /en-gb/.

Worked example 2: the "thin content" red herring

On a different client, a Riyadh electronics store (illustrative), the crawler flags 1,200 "low content" pages. Segmenting shows 1,100 are product variant URLs (?colour=) that already canonicalise correctly to the parent product — working as designed — and 100 are genuinely thin accessory pages. The finding reported is the 100, with examples; the 1,100 go into the "notable non-issues" list so the client doesn't panic when a tool shows the same warning.

Common mistakes

  • Reporting the crawler's issue list verbatim.
  • Treating symptoms (duplicate titles on facets) as separate root issues.
  • Stating a finding without URL-level examples and the evidence source.

Video lecture: Reading the crawl: indexability, status and templates

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

  1. Reading the crawl
  2. Step 1: segment × indexability
  3. Step 2: open real examples
  4. Step 3: redirect review
  5. Step 4: template checks
  6. Step 5: content signals
  7. Hands-on: pivot and crosstab
  8. Record findings
  9. Example 2: the thin-content red herring (illustrative)
  10. Common mistakes
  11. Symptom or cause?
  12. Watch me do it: three findings in an hour
  13. Keep a non-issues list
  14. Recap and try this now

Lecture transcript

Reading the crawl

The Kiran Home crawl has finished. Around sixty-one thousand URLs, all illustrative, on a site that only wants about four and a half thousand pages indexed. Your job isn't to list every warning. It's to explain that gap with a handful of patterns, backed by real examples. In this lecture you'll pivot the crawl by segment, investigate each pattern with samples, review redirects and templates, and record findings so they survive a sceptical stakeholder.

Step 1: segment × indexability

Step one: the indexability overview by segment. Pivot the crawl by segment and indexability. For Kiran Home, the illustrative picture is striking. Around sixteen thousand product URLs, but only about four thousand indexable, because the rest canonicalise to another market. About forty-two thousand facet URLs, all indexable, nothing preventing it. Nineteen hundred internal search pages, crawlable and indexable. And blog posts mostly canonicalised to the UK versions. Three patterns jump out immediately.

Step 2: open real examples

Step two: never report a pattern without opening real examples. Open five products in each market. On the UAE brass lantern page, the canonical points to the UK version, while hreflang lists both UAE and UK. So every non-UK product tells Google it's a duplicate of the UK page, and hreflang can't work, because hreflang only connects canonical pages. That fits hypothesis three, broken international signals, and it explains why UAE users might see UK prices in results.

Step 3: redirect review

Step three: status codes and redirects. The list crawl of fifty-six hundred pre-relaunch URLs, illustrative, shows about thirty-one hundred single-hop redirects to relevant pages, which is good. Nine hundred chains of two or more hops. Seven hundred redirects to the home page, mostly old blog category and tag URLs. Eight hundred and fifty four-oh-fours, mostly old product URLs with a different slug pattern. And fifty old URLs still live. The four-oh-fours need checking against backlinks and Search Console clicks; those with links or traffic are high priority.

Step 4: template checks

Step four: template-level checks. For each key template, open a representative URL and review the rendered HTML: title and H1, canonical, robots directives, main content, internal links as real anchors, and valid JSON-LD matching visible data. For Kiran Home, category pages render only the first twenty-four products, and load more is a button with no paginated URLs. So products beyond the first twenty-four have few internal links. That's a discovery issue, and evidence for hypothesis two.

Step 5: content signals

Step five: content signals. Duplicate titles are mostly on facets, which is a symptom, not a separate issue. About four hundred products have fewer than thirty words of description, which is genuinely thin. And market pages that differ only by currency are near-duplicates, acceptable for regional pages if hreflang is correct, but worth noting.

Hands-on: pivot and crosstab

Hands-on. The lesson text has a pandas script that assigns each crawled URL a segment and a market, then builds the segment-by-indexability pivot. It then takes canonicalised product URLs and cross-tabulates their market against the market their canonical points to. For Kiran Home, that crosstab makes the pattern unmistakable in a single table: Pakistan and UAE products all pointing to the UK.

Record findings

Recording findings properly. Each finding gets an ID, a title, the evidence with counts and example URLs, the source, the affected templates and markets, the hypothesis it links to, and a severity to be scored later. For Kiran Home, finding F three: non-UK product pages canonicalise to UK versions. Evidence: sixteen thousand two hundred product URLs crawled, twelve thousand three hundred canonicalised to UK equivalents, five examples, and URL Inspection confirming Google chose the UK URL.

Example 2: the thin-content red herring (illustrative)

Worked example one is Kiran Home, which you've just seen. Worked example two, a different and illustrative client. A Riyadh electronics store's crawler flags twelve hundred low content pages. Segmenting shows eleven hundred are colour variant URLs that already canonicalise correctly to the parent product. Working as designed. Only a hundred are genuinely thin accessory pages. The finding reported is the hundred, with examples, and the eleven hundred go into a notable non-issues list, so the client doesn't panic the next time a tool shows that warning.

Common mistakes

Common mistakes. Reporting the crawler's issue list verbatim. Treating symptoms, like duplicate titles on facets, as separate root issues. And stating a finding without URL-level examples and the evidence source. Every one of these makes your audit easier to dismiss.

Symptom or cause?

Let's pause on a skill that separates senior auditors from juniors: distinguishing a symptom from a cause while you're still reading the crawl. A useful habit is to ask, if this pattern disappeared tomorrow, what would have caused it to disappear? For duplicate titles on facets, the answer is, facets stopped being crawlable. So duplicate titles aren't the finding; facet crawlability is. For thin products, the answer is, someone wrote better descriptions. So thin content is its own finding, owned by marketing. Asking that question as you read keeps your workbook short and your tickets meaningful.

Watch me do it: three findings in an hour

Watch me do it. I'll turn the Kiran Home crawl into three findings in under an hour. Step one: I export internal HTML and run the pivot script. The table shows facets and search as fully indexable, and most product URLs canonicalised. Step two: I run the crosstab of product market against canonical market. Every non-UK row points to the UK column. That's pattern one. Step three: I open five products per market. I copy the head section of one UAE product into my notes: canonical to the UK URL, hreflang listing both. Then I check URL Inspection for the same URL: Google-selected canonical is the UK page. Evidence confirmed. Step four: I filter the list crawl of old URLs by status. I sort the four-oh-fours by Search Console clicks from last year, and the top twenty include best-selling lanterns. Pattern two. Step five: I open a category page's rendered HTML and count product links: twenty-four. The category has ninety products. No paginated URLs anywhere. Pattern three. Step six: I write three finding cards, F three, F six and F fourteen, each with counts, five example URLs, the evidence source and the hypothesis it supports. Everything else goes to a parking list for triage.

Keep a non-issues list

One more practical tip: keep a notable non-issues list from the very start. Every crawler produces warnings that are harmless in context: variant URLs that canonicalise as designed, multiple H1s on a clear page, missing meta descriptions on paginated pages, low word counts on contact pages. Listing them explicitly in your report does two things. It shows the client you looked and made a judgement. And it stops them panicking when a tool flags the same thing next month.

Recap and try this now

Recap. Pivot by segment and indexability, open real examples, review redirects and templates, separate symptoms from causes, and record each finding with evidence. Try this now. Take any crawl export, add a segment column with a few regex patterns, and build the pivot. Write down the three biggest patterns and open five example URLs for each before you call any of them a finding.

Key takeaways

  • Pivot the crawl by segment and indexability to find the big patterns first.
  • Validate every pattern with real URL samples in the rendered HTML and URL Inspection.
  • Test legacy URLs with a list crawl and classify outcomes (one-hop, chain, home-page, 404, still live).
  • Record each finding with ID, evidence, examples, affected templates and linked hypothesis.

Try it

Take a crawl export from any site and build a segment × indexability pivot. Write up the three biggest patterns with five sample URLs each.