Advanced Prompt EngineeringFew-shot examples and structured outputs · Lesson 4 of 17

Structured outputs and JSON schemas

Article · 12 min · 8 min lecture

Video lecture

Structured outputs and JSON schemas

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Structured outputs and JSON schemas

  • Three levels of enforcement
  • Schema design that reduces invention
  • Current APIs
  • Validation is still your job

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why structure matters

The moment an AI output feeds another system (a CRM, a spreadsheet, a dashboard or another prompt), free text becomes a liability. You need outputs that are parseable, validated and predictable. Structured output techniques turn a language model into a dependable component.

There are three levels of enforcement, from weakest to strongest:

  1. Prompted format. "Return JSON with fields name, sentiment, reason." Usually works, occasionally breaks (a missing brace, extra commentary, a renamed key).
  2. Tool or function calling. You define a function with a JSON Schema for its parameters and the model returns arguments that fit it. Stronger, because the model is trained to produce arguments in this form.
  3. Constrained decoding (often called "structured outputs" or "JSON mode" with a schema). Many major APIs can now constrain generation so the output is guaranteed to match your schema. Check your provider's documentation for which schema features are supported, since support for things like recursive schemas or certain keywords varies.

Even at level 3, the structure is guaranteed, not the truth. A perfectly valid JSON object can still contain a hallucinated value.

Designing a good schema

{
  "type": "object",
  "properties": {
    "company_name": { "type": "string" },
    "intent": { "type": "string", "enum": ["HOT", "WARM", "COLD", "SPAM"] },
    "budget_mentioned": { "type": "boolean" },
    "budget_amount": { "type": ["number", "null"],
      "description": "Numeric amount only if explicitly stated; otherwise null" },
    "evidence": { "type": "string",
      "description": "Exact quote from the message supporting the intent label" }
  },
  "required": ["company_name", "intent", "budget_mentioned", "budget_amount", "evidence"],
  "additionalProperties": false
}

Design principles:

  • Use enums for categories. They eliminate spelling variants like "Hot", "hot lead" and "HOT ".
  • Make absence explicit. Allow null and describe when to use it. Otherwise the model will invent a value to fill a required field.
  • Descriptions are prompts. Field descriptions are read by the model; use them to state rules ("exact quote", "ISO 8601 date").
  • Add an evidence field. Requiring a supporting quote makes outputs auditable and tends to reduce unsupported claims.
  • Order fields thoughtfully. Models generate left to right. Placing evidence or a short reasoning field before the final label lets the model ground its decision first. Placing it after makes it a post-hoc justification.
  • Keep it flat where you can. Deeply nested schemas are harder for models and for humans reviewing outputs.

Validation is still your job

Always validate on your side, even with constrained decoding:

import json
from jsonschema import validate, ValidationError

def parse_lead(raw: str, schema: dict):
    try:
        data = json.loads(raw)
        validate(data, schema)
    except (json.JSONDecodeError, ValidationError) as err:
        return None, str(err)   # retry once with the error message, then route to a human
    if data["budget_amount"] is not None and not data["budget_mentioned"]:
        return None, "inconsistent budget fields"
    return data, None

Notice the semantic check after schema validation: the schema cannot express "if amount is present, the boolean must be true". Business rules like this belong in code.

Retry with feedback

When validation fails and you are not using constrained decoding, a single retry that includes the error message ("Your output failed validation: budget_amount must be a number or null") fixes most cases. Cap retries; loops of retries burn money and hide systematic prompt problems. Track the failure rate as a metric.

Worked example: extracting invoice data

A finance operations team extracts supplier, date, currency, total and line items from PDF invoices across the UK, UAE and Pakistan. Lessons learned from a pilot like this typically include:

  • Dates appear in many formats; specify ISO 8601 output and describe ambiguous cases ("if day and month are ambiguous, set date_ambiguous to true").
  • Currencies must be ISO codes (GBP, AED, PKR), not symbols, because symbols can be ambiguous across countries.
  • Totals should be numbers without thousands separators, and a code check confirms line items sum to the total within a tolerance. Mismatches go to human review rather than being "fixed" by the model.

Failure modes

  • Schema pressure hallucination. Required fields with no null option force invention.
  • Enum drift. You add a new category in code but forget the schema, or vice versa. Generate both from one source.
  • Truncation. Long arrays can hit output limits and cut off mid-object. Set sensible maximum lengths and paginate large extractions.

Hands-on: schema-enforced outputs in current APIs

Most major APIs now offer constrained decoding against a JSON Schema. The SDKs also accept Pydantic models and validate the result for you.

Claude (Anthropic Python SDK):

from typing import Literal, Optional
from pydantic import BaseModel, Field
import anthropic, os

class Lead(BaseModel):
    evidence: str = Field(description="Exact quote supporting the intent label")
    intent: Literal["HOT", "WARM", "COLD", "SPAM"]
    company_name: Optional[str] = Field(description="Only if stated; otherwise null")
    budget_amount: Optional[float] = Field(description="Only if explicitly stated; otherwise null")

client = anthropic.Anthropic()
resp = client.messages.parse(
    model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"),
    max_tokens=1024,
    messages=[{"role": "user", "content":
        "Classify this lead:\n<message>We need Snapchat ads for our Jeddah launch, budget 40k SAR.</message>"}],
    output_format=Lead,
)
lead = resp.parsed_output  # a validated Lead instance
print(lead.intent, lead.budget_amount)

The raw equivalent passes output_config={"format": {"type": "json_schema", "schema": {...}}} to messages.create. Note that evidence comes first so the model grounds its decision before labelling.

OpenAI (Responses API):

from openai import OpenAI
client = OpenAI()  # reads OPENAI_API_KEY
resp = client.responses.parse(
    model=os.environ["OPENAI_MODEL"],  # set to a current model from the docs
    input="Classify this lead: We need Snapchat ads for our Jeddah launch, budget 40k SAR.",
    text_format=Lead,
)
lead = resp.output_parsed

In the Responses API, raw schemas go under text.format; the older Chat Completions API used response_format. Gemini offers the equivalent through a response schema in its generation config.

Know the limits of your provider's schema support

Constrained decoding supports a subset of JSON Schema that differs by provider. Common restrictions include requiring additionalProperties: false, requiring every property to be listed in required (use a nullable type for optional values), and limited support for some keywords or recursion. Read the current documentation before designing complex schemas, and keep a unit test that sends your schema to the API so incompatibilities surface in CI, not production.

Refusals and stop reasons

A guaranteed schema cannot force an answer the model declines to give, and long outputs can be cut off. Always check the stop reason (for Claude, stop_reason such as max_tokens or refusal; OpenAI returns refusal content items) before trusting a parse. Route those cases to a retry with a larger limit, or to a human.

Structured outputs versus tool calling

Use structured outputs when you want the model's final answer in a fixed shape. Use tool calling when the model should decide whether and when to call a function, with arguments that fit a schema (many APIs offer a strict mode for tool arguments too). Both can be combined in one request.

Going further

Treat schemas as versioned contracts between the model and your code. When you change a schema, bump a version field, keep old parsers until traffic moves over, and rerun your evaluation set. Many incidents in AI features are not model failures but silent contract changes between components.

Key takeaways

  • Prompted formats can break; tool calling is stronger; constrained decoding can guarantee structure but not truth.
  • Use enums, explicit nulls, descriptive fields and an evidence field to reduce invention and ease auditing.
  • Put evidence or reasoning fields before the final decision so the model grounds its answer first.
  • Always validate in code, add semantic business-rule checks, and cap retries.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your schema requires a 'budget_amount' number, but many messages mention no budget. What is the likely result?
  2. Constrained decoding guarantees which of the following?
  3. Why place an 'evidence' field before the 'intent' field in the schema?

Put it into practice

Design a JSON schema for one extraction task in your work, including enums, nullable fields, and an evidence field. Write down two business-rule checks that the schema cannot express.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.