Technical SEO MasteryHow search engines crawl, render and index · Lesson 2 of 18

JavaScript SEO: SPAs, SSR and prerendering

Video lesson · 15 min · 8 min lecture

Video lecture

JavaScript SEO: SPAs, SSR and prerendering

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

JavaScript SEO

  • Raw HTML vs rendered HTML
  • Rendering strategies compared
  • Non-negotiables + a diff script

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The core problem

A classic server-rendered page sends complete HTML: content, links and meta tags are all in the first response. A client-side rendered single-page application (SPA) often sends an almost empty shell:

<body>
  <div id="root"></div>
  <script src="/static/js/main.4f2a.js"></script>
</body>

Everything a search engine needs — text, internal links, titles, canonicals — only exists after JavaScript runs. Google can render JavaScript, but you add risk and delay at every step, and many other crawlers (including several AI crawlers and social preview bots) do not execute JavaScript at all.

Rendering strategies compared

StrategyHow it worksSEO risk
Client-side rendering (CSR)Browser builds page from JSHighest: content and links depend on successful rendering
Server-side rendering (SSR)Server returns full HTML per request, then hydratesLow, if hydration does not change content
Static site generation (SSG)HTML built at deploy timeLowest; great for content that changes infrequently
Incremental/on-demand regenerationStatic pages rebuilt on a schedule or triggerLow; watch for stale content
Dynamic renderingServe prerendered HTML to bots, CSR to usersGoogle calls this a workaround, not a long-term solution

Frameworks such as Next.js, Nuxt, SvelteKit, Remix and Astro make SSR or SSG the default. For most sites the recommendation is simple: send meaningful HTML in the initial response for anything you want indexed.

The non-negotiables for JavaScript sites

  1. Real links. Google discovers links from a elements with an href. Router links that render as proper anchors are fine; click handlers on div or span are not.
<!-- Crawlable -->
<a href="/pricing">Pricing</a>
<!-- Not reliably crawlable -->
<span onclick="router.push('/pricing')">Pricing</span>
<a href="javascript:void(0)" onclick="goTo('pricing')">Pricing</a>
  1. Real URLs, not fragments. Use the History API (/products/red-shoes), not hash routing (/#/products/red-shoes). Google generally ignores everything after #.
  2. Correct status codes. An SPA that shows "Product not found" but returns 200 creates soft 404s. Either route unknown paths to a server-generated 404, or add a noindex robots meta tag via JavaScript on error views (Google documents both approaches).
  3. Meta tags in the initial HTML where possible. Titles, canonicals and robots meta can be set with JavaScript and Google usually honours them after rendering, but conflicting values between raw and rendered HTML cause unpredictable results. Google's December 2025 documentation updates are explicit: when Google encounters noindex it may skip rendering, so never put noindex in the raw HTML and remove it with JavaScript; and pages returning non-200 status codes may not be rendered, so don't rely on JavaScript on error pages to change what Google indexes.
  4. Do not block resources. Allow crawling of JS, CSS and API endpoints the page needs to render.
  5. Lazy-load safely. Use native loading="lazy" or IntersectionObserver; do not require scroll events to load primary content.

Hydration mismatches

With SSR, the server HTML is "hydrated" by client JS. If the client renders different content (a different price, or different canonical because of a locale cookie), Google may index whichever version it saw. Audit by comparing raw HTML with rendered HTML.

How to audit a JavaScript site

  1. Compare raw vs rendered. Crawl twice with a desktop crawler (Screaming Frog, Sitebulb or similar): once text-only, once with JavaScript rendering. Diff word counts, links, titles and canonicals. Big differences mean you depend on rendering.
  2. Inspect with Google. URL Inspection → Test live URL → View tested page → HTML tab. Search the rendered HTML for a unique sentence from your main content.
  3. Check the console and resources. The "More info" tab lists page resources that could not be loaded and JavaScript console messages.
  4. Test without JavaScript. Disable JS in Chrome DevTools. What remains is roughly what non-rendering crawlers see.
  5. Check internal link discovery. In your rendered crawl, compare the number of unique internal URLs found against the text-only crawl.

Worked example 1: a Dubai React catalogue

An e-commerce brand in Dubai rebuilt its catalogue as a React SPA. Category pages returned only the shell; products loaded from /api/products, which was disallowed in robots.txt "to save crawl budget". The rendered page in URL Inspection showed an empty grid, so Google saw thin category pages with no product links. The fix: allow the API path, switch category and product templates to SSR, output product links as real anchors, and return 404 for discontinued products. Illustratively, a fix like this typically shows up first as more pages moving into "Indexed" in the Page indexing report over the following weeks — measure it there rather than assuming.

Hands-on: diff raw vs rendered HTML with Playwright (Python)

# pip install requests playwright beautifulsoup4 && playwright install chromium
import requests
from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright

URL = "https://www.example.com/collections/trainers/"
UA = ("Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) "
      "Chrome/126.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)")

def summarise(html):
    s = BeautifulSoup(html, "html.parser")
    canon = s.find("link", rel="canonical")
    robots = s.find("meta", attrs={"name": "robots"})
    return {"title": s.title.string.strip() if s.title and s.title.string else None,
            "canonical": canon.get("href") if canon else None,
            "robots": robots.get("content") if robots else None,
            "links": len({a["href"] for a in s.find_all("a", href=True)}),
            "words": len(s.get_text(" ", strip=True).split())}

raw = requests.get(URL, headers={"User-Agent": UA}, timeout=20)
with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(user_agent=UA)
    page.goto(URL, wait_until="networkidle", timeout=45000)
    rendered = page.content()
    browser.close()

a, b = summarise(raw.text), summarise(rendered)
for k in a:
    flag = "" if a[k] == b[k] else "   <-- differs"
    print(f"{k:10} raw={a[k]!s:45} rendered={b[k]!s:45}{flag}")

Differences in canonical or robots between raw and rendered HTML are high-severity; large differences in links or words mean the page depends on rendering for discovery or content. Playwright is not Google's renderer — confirm important findings with URL Inspection.

Worked example 2: a Lahore SaaS on a framework "that does SSR"

A Lahore SaaS company (illustrative) says its Next.js marketing site is server-rendered, so JavaScript SEO "isn't a concern". The diff shows the raw HTML of pricing pages contains only a loading skeleton: the pricing component fetches plans client-side after hydration. The fix is fetching plan data at build or request time so prices are in the server HTML. A second finding: the locale switcher changes the canonical client-side based on a cookie, producing raw/rendered canonical mismatches — fixed by rendering the canonical on the server from the URL path only.

Common mistakes

  • Testing only in a browser where you are logged in or have cookies set.
  • Relying on dynamic rendering indefinitely, and letting the bot version drift from the user version (which can look like cloaking).
  • Infinite scroll with no paginated URLs.
  • Forgetting that social, many AI and some search crawlers do not run JavaScript.

Key takeaways

  • Send meaningful HTML (content, links, meta tags) in the initial response for anything you want indexed.
  • Use real anchor links and History API URLs; avoid hash routing and click-only navigation.
  • SPAs must avoid soft 404s by returning real 404s or adding noindex on error views.
  • Audit JavaScript sites by diffing raw and rendered crawls and checking URL Inspection's rendered HTML.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which link is reliably crawlable by Googlebot?
  2. Google's documentation describes dynamic rendering as…
  3. Why is putting noindex in the raw HTML and removing it with JavaScript risky?

Put it into practice

Crawl one JavaScript-heavy template twice (text-only and rendered) and write down the differences in word count, internal links and meta tags.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.