Programmatic SEO and AI Content at Scale — Without Getting PenalizedBuilding page systems: templates, links and structured data · Lesson 7 of 14

Structured data at scale: JSON-LD generation and validation

Article · 8 min · 8 min lecture

Video lecture

Structured data at scale: JSON-LD generation and validation

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Structured data at scale

  • What it does and doesn't do
  • Relevant types
  • Generating JSON-LD
  • Validation at scale

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What structured data does (and does not) do

Structured data (Schema.org vocabulary, usually as JSON-LD) describes page content in a machine-readable way. Google uses it to understand pages and to make pages eligible for rich results (e.g., product snippets with price and availability, review snippets, breadcrumbs, local business details, events, job postings). It does not guarantee rich results, and it is not a direct ranking boost. It must describe content that is visible on the page and follow Google's structured data policies.

Google periodically retires rich result types — HowTo rich results were removed and FAQ rich results limited to certain authoritative sites in 2023, and several more types were phased out in 2025 — so always check Google's current Search Gallery before investing in a type. Structured data can still help other systems (other search engines, AI assistants, product feeds) understand your pages.

Types commonly relevant to page systems

Page typeSchema typesNotes
Listing/category pagesItemList, BreadcrumbList, CollectionPageGoogle's carousel support for ItemList is limited to certain types; breadcrumbs are widely useful
Product/comparisonProduct, Offer, AggregateRating, ReviewOnly real, visible offers and reviews; follow review snippet guidelines
Local entitiesLocalBusiness subtypes, PostalAddress, OpeningHoursSpecification, GeoCoordinatesMatch your Google Business Profile details where relevant
JobsJobPostingStrict requirements; eligible for the Indexing API
EventsEventDates, location, status
Articles/guidesArticle, author Person, OrganizationSupports E-E-A-T signals of accountability
DatasetsDatasetFor data pages; eligible for Dataset Search

Generating JSON-LD from your data

Generate structured data from the same data source that renders the page, so they can never disagree:

import json

def breadcrumb(items):
    return {"@context": "https://schema.org", "@type": "BreadcrumbList",
            "itemListElement": [{"@type": "ListItem", "position": i + 1, "name": n, "item": u}
                                for i, (n, u) in enumerate(items)]}

def local_business(clinic):
    data = {
        "@context": "https://schema.org", "@type": "Dentist",
        "name": clinic["name"], "url": clinic["url"], "telephone": clinic["phone"],
        "address": {"@type": "PostalAddress", "streetAddress": clinic["street"],
                    "addressLocality": clinic["city"], "addressCountry": clinic["country"]},
        "geo": {"@type": "GeoCoordinates", "latitude": clinic["lat"], "longitude": clinic["lon"]},
        "openingHoursSpecification": [
            {"@type": "OpeningHoursSpecification", "dayOfWeek": d, "opens": o, "closes": c}
            for d, o, c in clinic["hours"]],
    }
    if clinic.get("rating_count", 0) >= 5:  # only when real, visible reviews exist
        data["aggregateRating"] = {"@type": "AggregateRating",
                                   "ratingValue": round(clinic["rating"], 1),
                                   "reviewCount": clinic["rating_count"]}
    return data

def script_tag(obj):
    return '<script type="application/ld+json">' + json.dumps(obj, ensure_ascii=False) + "</script>"

Note self-serving reviews: Google does not show review rich results for reviews a business publishes about itself (LocalBusiness/Organization about itself). Aggregated third-party reviews on a directory about listed businesses are a different case — check current guidelines.

Validation at scale

  • Unit tests: validate generated JSON-LD against your own schema expectations (required fields, types).
  • Rich Results Test and Schema Markup Validator for samples of each template.
  • Search Console enhancement reports (e.g., breadcrumbs, products, merchant listings) for errors and warnings across all pages.
  • Consistency checks: price, availability and ratings in JSON-LD must match visible content and, for products, your Merchant Center feed.
def check_business(d):
    errors = []
    for f in ["name", "address", "telephone"]:
        if not d.get(f):
            errors.append(f"missing {f}")
    ar = d.get("aggregateRating")
    if ar and not (1 <= ar["ratingValue"] <= 5):
        errors.append("ratingValue out of range")
    return errors

Worked example: a Pakistani job board

"[job title] jobs in [city]" hubs use BreadcrumbList and ItemList; individual job pages use JobPosting with salary ranges (where provided), validThrough, employment type, and location or remote eligibility. Expired jobs are updated promptly (and removed from the index) — stale job postings are a known quality problem. The Indexing API is used for job posting pages, which is one of the use cases Google permits for it.

AI-driven search experiences and assistants increasingly read pages directly. Clear structured data does not guarantee a mention or citation, but it helps any system parse entities, prices, locations, dates and authorship accurately. For data pages, consider publishing a Dataset description with the methodology and license, and make key numbers visible in HTML tables rather than only in images or client-rendered charts. Consistency across your page, markup, feeds and business profiles reduces the chance of AI systems repeating outdated or conflicting facts about you.

Pitfalls

  • Markup that doesn't match visible content (spam risk, manual actions for structured data).
  • Fake or self-serving review markup.
  • Investing in retired rich result types.
  • Hand-written markup that drifts from data.

How to measure success

Zero critical errors in Search Console enhancement reports, rich result impressions where eligible (Search Console search appearance filters), and consistency tests passing in CI.

Key takeaways

  • Structured data (JSON-LD) aids understanding and rich-result eligibility; it's not a guaranteed boost.
  • Mark up only visible, truthful content and follow Google's structured data policies.
  • Check Google's Search Gallery — some rich result types were retired or limited (HowTo, FAQ in 2023; more in 2025).
  • Generate JSON-LD from the same data that renders the page; test in CI and with Google's tools.
  • Avoid self-serving review markup and keep jobs/products fresh.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does structured data guarantee?
  2. Why generate JSON-LD from the same data that renders the page?
  3. Which use of review markup is problematic?

Put it into practice

Choose the schema types for your page system, write a generator function for one type, and list the consistency checks you'll run in CI.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.