Technical SEO Audit Workshop · Triage and prioritisation · lesson 7 of 14 · 12 min
From findings to root causes
Why grouping matters
Your Kiran Home workbook now holds about 40 findings. Many are symptoms of the same underlying cause. Developers and managers act on causes, not symptoms. Grouping transforms "40 issues" into "7 problems with clear fixes".
The 5 Whys, applied to SEO
Take the finding "3,100 duplicate title tags".
- Why duplicate titles? Because facet URLs reuse the category title.
- Why are facet URLs crawled? Because filters are links with parameters, and nothing blocks them.
- Why are they linked as crawlable links? The storefront's filter component outputs anchor links for every combination.
- Why was that not caught? No SEO requirements were specified for the new filter component.
- Why? The relaunch had no SEO acceptance criteria.
Root cause: facet URLs are crawlable and indexable by default. Process cause: no SEO acceptance criteria in releases. The duplicate titles, the crawl waste in logs, and part of the "Crawled – currently not indexed" growth are all symptoms.
Kiran Home: root-cause map
| Root cause | Symptoms (finding IDs) | |---|---| | RC1: Cross-market canonicals on product and blog templates | Non-UK pages not indexed (F-03), hreflang ineffective (F-11), UAE users see UK prices in results (F-12), Page indexing "Alternate page with proper canonical" (E-08) | | RC2: Facets and internal search crawlable and indexable | Duplicate titles (F-05), crawl waste (E-07), thin search pages indexed (F-09) | | RC3: Incomplete relaunch redirects | Legacy 404s with backlinks (F-06), home-page redirects (F-07), chains (F-08) | | RC4: Category pagination not crawlable | Deep products with few inlinks (F-14), products not crawled in logs (E-09) | | RC5: Product template performance | Poor INP/CLS from reviews widget (F-16), LCP image lazy-loaded on categories (F-17) | | RC6: Product structured data incomplete | Product rich result items dropped (F-19) | | RC7: Process — no SEO release checks | Enables recurrence of all of the above | | RC8: CDN bot rules challenge some crawlers (H5) | Bingbot and AI search bots receive challenge pages on product URLs (E-12); Bing indexing of products fell (E-13) |
Distinguishing issue types
Label each root cause, because it changes who owns it and how it is fixed:
- Configuration — a setting or rule (robots.txt, canonical logic, redirect manager).
- Template/code — component behaviour (filters, pagination, JSON-LD output).
- Content — thin or duplicate copy; owned by marketing.
- Infrastructure — server response, CDN, caching.
- Process — missing checks, ownership, documentation.
Estimating affected scope
For each root cause, quantify scope with your data:
- Number of URLs and templates affected.
- Share of revenue or organic traffic flowing through those templates (from GA4 and Search Console).
- Markets affected.
RC1 affects every non-UK product and blog URL, in the three markets that the founder said account for most revenue — a very large scope. RC6 affects all products but its impact is limited to rich result presentation.
Watch for false positives
Not every crawler warning is a problem:
- "Missing meta description" on paginated category pages — low value; Google often rewrites snippets anyway.
- "Multiple H1s" — HTML5 allows it; not a meaningful issue if the page is clear.
- "Low word count" on a contact page — expected.
- "Non-indexable canonicalised URLs" that are intentionally canonicalised (UTM variants) — working as designed.
Explicitly list notable non-issues in the report so stakeholders do not receive the same alerts from a tool and panic.
Tie back to hypotheses
| Hypothesis | Verdict | Evidence | |---|---|---| | H1 Redirect gaps | Confirmed (partial) | RC3 | | H2 Rendering/indexability regressions | Confirmed | RC1, RC4 | | H3 International signals broke | Confirmed | RC1 | | H4 Mobile performance regression | Confirmed on product/category templates | RC5 | | Alternative: algorithm update | Not supported — no confirmed update aligned with the drop, and the drop is concentrated in affected templates/markets | Performance segmentation |
Hands-on: a root-cause map as data
Keep the mapping in a small structured file so it can generate the report tables and ticket links automatically:
# root_causes.yaml
RC1:
title: Cross-market canonicals on product and journal templates
type: template/code
symptoms: [F-03, F-11, F-12, E-08, F-21]
hypotheses: [H2, H3]
scope: {urls: 12300, templates: [product, journal], markets: [en-pk, en-ae, ar-ae]}
RC8:
title: CDN bot rules challenge Bingbot and AI search bots
type: infrastructure
symptoms: [E-12, E-13]
hypotheses: [H5]
scope: {urls: all, templates: [all], markets: [all]}
# pip install pyyaml
import yaml
rc = yaml.safe_load(open("root_causes.yaml"))
for k, v in rc.items():
print(f"{k} | {v['title']} | {v['type']} | symptoms: {', '.join(v['symptoms'])}")
Worked example 2: one cause behind "twelve issues"
A Riyadh furniture retailer (illustrative) receives a tool report with twelve "critical" issues: duplicate titles, duplicate descriptions, canonicalised URLs in sitemaps, thin pages, crawl depth warnings and more. Grouping shows ten of the twelve trace to one cause: a colour-swatch component that generates a crawlable URL per colour for every product. One component fix plus a robots.txt rule addresses ten tool warnings at once — and the other two (missing alt text and slow LCP on the home page) are separate, smaller tickets.
Common mistakes
- Handing developers 40 symptom tickets that all touch the same component.
- Omitting process causes — the next release reintroduces the problems.
- Presenting tool warnings as issues without judging their relevance.
Video lecture: From findings to root causes
Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.
- From findings to root causes
- Why group?
- The five whys
- Kiran Home root causes
- Label by type
- Estimate scope
- False positives
- Hypothesis verdicts
- Hands-on: root causes as data
- Example 2: one cause, ten warnings (illustrative)
- Watch me do it: 40 cards → 8 causes
- From H5 to RC8
- Mistakes, recap, try this now
Lecture transcript
From findings to root causes
Your Kiran Home workbook now holds about forty findings. If you hand those to the developers as forty tickets, most will never be done, because many are symptoms of the same problem, and several touch the same component. In this lecture you'll group findings into root causes using the five whys, label each cause by type, estimate its scope, weed out false positives, and tie everything back to the hypotheses you wrote before crawling.
Why group?
Why group? Because developers and managers act on causes, not symptoms. Imagine a leaking roof. You could put forty buckets under forty drips, or fix three broken tiles. Grouping turns forty issues into seven problems with clear fixes, and it makes your audit dramatically easier to implement and to verify.
The five whys
The tool for this is the five whys, applied to SEO. Take the finding, three thousand one hundred duplicate title tags. Why? Because facet URLs reuse the category title. Why are facet URLs crawled? Because filters are links with parameters and nothing blocks them. Why are they links? Because the new filter component outputs anchor links for every combination. Why wasn't that caught? Because no SEO requirements were specified for the component. Why? Because the relaunch had no SEO acceptance criteria. Two causes emerge: facets crawlable by default, and no SEO release checks.
Kiran Home root causes
Here's Kiran Home's root-cause map, illustratively. RC one: cross-market canonicals on product and journal templates. RC two: facets and internal search crawlable and indexable. RC three: incomplete relaunch redirects. RC four: category pagination not crawlable. RC five: product template performance. RC six: product structured data incomplete. RC seven: no SEO release checks. And a new one from hypothesis five, RC eight: CDN bot rules challenging Bingbot and AI search bots on product URLs. Each cause lists the findings and evidence IDs that support it.
Label by type
Label each root cause by type, because the type changes who owns it and how it's fixed. Configuration: a setting or rule, like robots.txt, canonical logic or the redirect manager. Template or code: component behaviour, like filters, pagination or JSON-LD output. Content: thin or duplicate copy, owned by marketing. Infrastructure: server response, CDN and caching, like RC eight. And process: missing checks, ownership and documentation, like RC seven.
Estimate scope
Estimate the scope of each cause with your data: the number of URLs and templates affected, the share of revenue or organic traffic flowing through them, from GA4 and Search Console, and which markets are affected. RC one affects every non-UK product and journal URL, in the three markets the founder said drive most revenue. A very large scope. RC six affects all products, but its impact is limited to rich result presentation and data accuracy. Scope feeds straight into prioritisation in the next lesson.
False positives
Watch for false positives. Not every crawler warning is a problem. Missing meta descriptions on paginated category pages: low value, and Google often rewrites snippets anyway. Multiple H1s: allowed in HTML, and not meaningful if the page is clear. Low word count on a contact page: expected. Canonicalised UTM variants: working as designed. List notable non-issues explicitly in the report, so stakeholders don't panic when a tool shows the same alerts next month.
Hypothesis verdicts
Then tie back to your hypotheses. Redirect gaps: confirmed, partially, by RC three. Rendering and indexability regressions: confirmed, by RC one and RC four. International signals broke: confirmed, by RC one. Mobile performance regression: confirmed on product and category templates, RC five. Bot access changed: confirmed for Bingbot and AI search bots, RC eight. And the alternative, an algorithm update: not supported, because no confirmed update aligned with the drop, and the drop is concentrated in affected templates and markets.
Hands-on: root causes as data
Hands-on. Keep the root-cause map as data, in a small YAML file: for each cause, a title, type, symptom IDs, hypotheses and scope. A few lines of Python then print the table for your report, and the same file can generate ticket links later. It sounds like overkill, until the client asks for the fifth time which findings sit under which cause, and you answer in seconds.
Example 2: one cause, ten warnings (illustrative)
Worked example two, a different client, illustrative. A Riyadh furniture retailer receives a tool report with twelve critical issues: duplicate titles, duplicate descriptions, canonicalised URLs in sitemaps, thin pages, depth warnings and more. Grouping shows ten of the twelve trace to one cause: a colour-swatch component that generates a crawlable URL per colour for every product. One component fix plus a robots rule addresses ten warnings at once. The remaining two, missing alt text and slow home-page LCP, become separate, smaller tickets.
Watch me do it: 40 cards → 8 causes
Watch me do it. I'll group Kiran Home's forty findings live. Step one: I print every finding card and lay them out, or on screen, drag them onto a whiteboard. Step two: I pick the finding with the biggest count, three thousand one hundred duplicate titles, and run the five whys out loud. Facet URLs reuse titles. Facets are crawlable links. The filter component outputs anchors. No SEO requirements. No release checks. I create two cause cards: facets crawlable, and no release checks. Step three: I drag every card that the facet cause explains under it: duplicate titles, crawl waste in logs, indexed search pages. Step four: I repeat with the cross-market canonical finding. Under it go the non-indexed UAE products, the ineffective hreflang, the UK prices shown in the UAE, and now the journal canonical that changes after hydration. Step five: after twenty minutes, most cards sit under eight causes. Six cards remain. I check each: two are real but small, like thin product descriptions, so they become content findings; four are tool warnings that are working as designed, so they go to the notable non-issues list. Step six: I write the root causes into the YAML file, with types, symptom IDs, hypotheses and scope.
From H5 to RC8
Let's look more closely at RC eight, because it's a good example of how a hypothesis turns into a root cause. The CDN firewall events showed challenge pages served to requests from verified Bingbot IPs and to OpenAI's and Perplexity's search bots on product URLs, starting the week of the relaunch. Bing Webmaster Tools showed product pages dropping out of Bing's index over the following weeks. Googlebot wasn't affected, because the CDN had a separate rule for known search engines that only covered Google. So the cause is infrastructure, the owner is the developer who manages Cloudflare, and the fix is a rule change aligned with the written crawler policy. Three pieces of evidence, one clear cause.
Mistakes, recap, try this now
Common mistakes. Handing developers forty symptom tickets that all touch the same component. Omitting process causes, so the next release reintroduces the problems. And presenting tool warnings as issues without judging their relevance. Recap: group with the five whys, label by type, estimate scope, list non-issues, and give every hypothesis a verdict. Try this now: take your last audit or tool report, and group its issues into no more than seven root causes, each with a type and an owner.
Key takeaways
- Group symptoms into root causes using a 5 Whys chain; developers act on causes.
- Label root causes by type (configuration, template, content, infrastructure, process) to route ownership.
- Quantify scope with URL counts and traffic or revenue share per template and market.
- Call out false positives and close the loop on each hypothesis with evidence.
Try it
Take ten findings from an audit and group them into root causes using the 5 Whys. Label each by type and owner.