Advanced Prompt EngineeringFew-shot examples and structured outputs · Lesson 4 of 17
Structured outputs and JSON schemas
Video lecture
Structured outputs and JSON schemas
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Structured outputs and JSON schemas
The moment an AI output feeds another system, a CRM, a spreadsheet, a dashboard or another prompt, free text becomes a liability. One missing brace at two in the morning and your pipeline stops. In this lecture you will learn the three levels of structure enforcement, how to design schemas that reduce invented values, how to use schema-enforced outputs in the Claude and OpenAI APIs, and why validation is still your job even when the structure is guaranteed.
0:34 Three levels
There are three levels of enforcement. Level one, prompted format: return JSON with these fields. Usually works, occasionally breaks with a missing brace, extra commentary or a renamed key. Level two, tool or function calling: you define a function with a JSON Schema for its parameters, and the model returns arguments in that shape. Stronger, because models are trained for it. Level three, constrained decoding, usually called structured outputs: the API constrains generation so the output is guaranteed to match your schema. Most major APIs now offer this. But remember: the structure is guaranteed, not the truth. A perfectly valid object can still contain a hallucinated value.
1:21 Schema design
Schema design is where you prevent invention. Use enums for categories, so you never get hot, Hot and hot lead as three different values. Make absence explicit: allow null and describe when to use it, otherwise the model will invent a value to fill a required field. Remember that field descriptions are prompts; the model reads them, so state rules there, like exact quote or ISO eight six zero one date. Add an evidence field that requires a supporting quote. And keep it as flat as you reasonably can.
2:00 Order fields deliberately
Field order matters more than people expect. Models generate from left to right. If the evidence or a short reasoning field comes before the final label, the model grounds its decision first. If it comes after, it becomes a post-hoc justification for a decision already made. So put evidence first, label second. In the lesson's code, the Pydantic model lists evidence, then intent, then company name and budget amount, each optional value typed to allow null.
2:33 Current APIs
Let us look at the current APIs. With the Anthropic Python SDK, you can define a Pydantic model and call messages dot parse with an output format set to that model; you get back a validated object in parsed output. The raw equivalent passes an output config with a JSON schema format to messages dot create. With OpenAI's Responses API, responses dot parse with a text format parameter returns output parsed; raw schemas go under text dot format, while the older Chat Completions API used response format. Gemini offers the equivalent through a response schema in its generation config.
3:16 Know your provider's limits
Each provider supports a subset of JSON Schema, and the subsets differ. Common restrictions include requiring additional properties false, requiring every property in the required list, which means optional values become nullable types, and limited support for some keywords or recursion. So read the current docs before designing complex schemas, and keep a unit test that sends your schema to the API, so incompatibilities show up in CI rather than production. Also check stop reasons. A schema cannot force an answer the model declines to give, and long outputs can be cut off at the token limit.
3:58 Validation is still your job
Even with guaranteed structure, validate in your own code. Parse, validate against the schema, and then run semantic business-rule checks that the schema cannot express. For example: if a budget amount is present, the budget mentioned flag must be true. Or: invoice line items must sum to the total within a tolerance. When checks fail and you are not using constrained decoding, one retry that includes the error message fixes most cases. Cap retries, and track the failure rate as a metric, because a rising rate usually means a prompt or data problem.
4:38 Worked example: invoice extraction
Here is a real-world pattern. A finance operations team extracts supplier, date, currency, total and line items from invoices across the UK, UAE and Pakistan. They learn to require ISO dates and a flag for ambiguous day and month order; to require ISO currency codes like GBP, AED and PKR instead of symbols, which can be ambiguous; and to output totals as plain numbers, with a code check that line items sum to the total. Mismatches go to human review, rather than being fixed by the model.
5:16 Structured outputs vs tools
When should you use structured outputs versus tool calling? Use structured outputs when you want the model's final answer in a fixed shape. Use tool calling when the model should decide whether and when to call a function; many APIs offer a strict mode for tool arguments too. You can combine both in one request. And treat schemas as versioned contracts: generate code types and schemas from one source, bump a version when fields change, and rerun your evals.
5:50 Example 1: extracting a meeting
A simple worked example first. You want to extract a meeting from an email: title, date, time and attendees. With a prompted format, one reply in fifty might arrive with a friendly sentence before the JSON, and your parser crashes. With a schema, where date is a string described as year-month-day, time is a string in twenty-four-hour format, attendees is an array of strings, and a location field is nullable for when none is given, the structure is guaranteed. But watch the location field. Without the null option, the model would have to invent a location for an email that never mentioned one. That is schema pressure hallucination, prevented by one small design choice.
6:40 Example 2: delivery exceptions (illustrative)
Now a business scenario, with illustrative numbers. A logistics company in Jeddah processes about ten thousand delivery exception notes a month from drivers, in Arabic and English. They need each note classified into an exception type, a responsible party and a customer-contact flag, then pushed to their operations dashboard. Their first version used prompted JSON: roughly two percent of outputs failed to parse, which meant around two hundred manual fixes a month. They moved to schema-enforced output with enums for exception type and responsible party, an evidence field quoting the driver's note, and nullable fields for missing data. Parse failures went to effectively zero. But validation in code still caught about one percent of records where the customer-contact flag contradicted the exception type, and those now go to a review queue instead of silently reaching the dashboard. Illustrative figures.
7:40 Recap
To recap. Prompted formats can break, tool calling is stronger, and constrained decoding guarantees structure but not truth. Design schemas with enums, explicit nulls, instructive descriptions and an evidence field placed before the decision. Use the current parse helpers, know your provider's schema subset, check stop reasons, and always validate business rules in code. Try this now: design a schema for one extraction task with enums, nullable fields and an evidence field, and write two business-rule checks the schema cannot express. Next module: decomposition, reasoning and verification.
Why structure matters
The moment an AI output feeds another system (a CRM, a spreadsheet, a dashboard or another prompt), free text becomes a liability. You need outputs that are parseable, validated and predictable. Structured output techniques turn a language model into a dependable component.
There are three levels of enforcement, from weakest to strongest:
- Prompted format. "Return JSON with fields name, sentiment, reason." Usually works, occasionally breaks (a missing brace, extra commentary, a renamed key).
- Tool or function calling. You define a function with a JSON Schema for its parameters and the model returns arguments that fit it. Stronger, because the model is trained to produce arguments in this form.
- Constrained decoding (often called "structured outputs" or "JSON mode" with a schema). Many major APIs can now constrain generation so the output is guaranteed to match your schema. Check your provider's documentation for which schema features are supported, since support for things like recursive schemas or certain keywords varies.
Even at level 3, the structure is guaranteed, not the truth. A perfectly valid JSON object can still contain a hallucinated value.
Designing a good schema
{
"type": "object",
"properties": {
"company_name": { "type": "string" },
"intent": { "type": "string", "enum": ["HOT", "WARM", "COLD", "SPAM"] },
"budget_mentioned": { "type": "boolean" },
"budget_amount": { "type": ["number", "null"],
"description": "Numeric amount only if explicitly stated; otherwise null" },
"evidence": { "type": "string",
"description": "Exact quote from the message supporting the intent label" }
},
"required": ["company_name", "intent", "budget_mentioned", "budget_amount", "evidence"],
"additionalProperties": false
}Design principles:
- Use enums for categories. They eliminate spelling variants like "Hot", "hot lead" and "HOT ".
- Make absence explicit. Allow
nulland describe when to use it. Otherwise the model will invent a value to fill a required field. - Descriptions are prompts. Field descriptions are read by the model; use them to state rules ("exact quote", "ISO 8601 date").
- Add an evidence field. Requiring a supporting quote makes outputs auditable and tends to reduce unsupported claims.
- Order fields thoughtfully. Models generate left to right. Placing
evidenceor a shortreasoningfield before the final label lets the model ground its decision first. Placing it after makes it a post-hoc justification. - Keep it flat where you can. Deeply nested schemas are harder for models and for humans reviewing outputs.
Validation is still your job
Always validate on your side, even with constrained decoding:
import json
from jsonschema import validate, ValidationError
def parse_lead(raw: str, schema: dict):
try:
data = json.loads(raw)
validate(data, schema)
except (json.JSONDecodeError, ValidationError) as err:
return None, str(err) # retry once with the error message, then route to a human
if data["budget_amount"] is not None and not data["budget_mentioned"]:
return None, "inconsistent budget fields"
return data, NoneNotice the semantic check after schema validation: the schema cannot express "if amount is present, the boolean must be true". Business rules like this belong in code.
Retry with feedback
When validation fails and you are not using constrained decoding, a single retry that includes the error message ("Your output failed validation: budget_amount must be a number or null") fixes most cases. Cap retries; loops of retries burn money and hide systematic prompt problems. Track the failure rate as a metric.
Worked example: extracting invoice data
A finance operations team extracts supplier, date, currency, total and line items from PDF invoices across the UK, UAE and Pakistan. Lessons learned from a pilot like this typically include:
- Dates appear in many formats; specify ISO 8601 output and describe ambiguous cases ("if day and month are ambiguous, set
date_ambiguousto true"). - Currencies must be ISO codes (GBP, AED, PKR), not symbols, because symbols can be ambiguous across countries.
- Totals should be numbers without thousands separators, and a code check confirms line items sum to the total within a tolerance. Mismatches go to human review rather than being "fixed" by the model.
Failure modes
- Schema pressure hallucination. Required fields with no null option force invention.
- Enum drift. You add a new category in code but forget the schema, or vice versa. Generate both from one source.
- Truncation. Long arrays can hit output limits and cut off mid-object. Set sensible maximum lengths and paginate large extractions.
Hands-on: schema-enforced outputs in current APIs
Most major APIs now offer constrained decoding against a JSON Schema. The SDKs also accept Pydantic models and validate the result for you.
Claude (Anthropic Python SDK):
from typing import Literal, Optional
from pydantic import BaseModel, Field
import anthropic, os
class Lead(BaseModel):
evidence: str = Field(description="Exact quote supporting the intent label")
intent: Literal["HOT", "WARM", "COLD", "SPAM"]
company_name: Optional[str] = Field(description="Only if stated; otherwise null")
budget_amount: Optional[float] = Field(description="Only if explicitly stated; otherwise null")
client = anthropic.Anthropic()
resp = client.messages.parse(
model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"),
max_tokens=1024,
messages=[{"role": "user", "content":
"Classify this lead:\n<message>We need Snapchat ads for our Jeddah launch, budget 40k SAR.</message>"}],
output_format=Lead,
)
lead = resp.parsed_output # a validated Lead instance
print(lead.intent, lead.budget_amount)The raw equivalent passes output_config={"format": {"type": "json_schema", "schema": {...}}} to messages.create. Note that evidence comes first so the model grounds its decision before labelling.
OpenAI (Responses API):
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
resp = client.responses.parse(
model=os.environ["OPENAI_MODEL"], # set to a current model from the docs
input="Classify this lead: We need Snapchat ads for our Jeddah launch, budget 40k SAR.",
text_format=Lead,
)
lead = resp.output_parsedIn the Responses API, raw schemas go under text.format; the older Chat Completions API used response_format. Gemini offers the equivalent through a response schema in its generation config.
Know the limits of your provider's schema support
Constrained decoding supports a subset of JSON Schema that differs by provider. Common restrictions include requiring additionalProperties: false, requiring every property to be listed in required (use a nullable type for optional values), and limited support for some keywords or recursion. Read the current documentation before designing complex schemas, and keep a unit test that sends your schema to the API so incompatibilities surface in CI, not production.
Refusals and stop reasons
A guaranteed schema cannot force an answer the model declines to give, and long outputs can be cut off. Always check the stop reason (for Claude, stop_reason such as max_tokens or refusal; OpenAI returns refusal content items) before trusting a parse. Route those cases to a retry with a larger limit, or to a human.
Structured outputs versus tool calling
Use structured outputs when you want the model's final answer in a fixed shape. Use tool calling when the model should decide whether and when to call a function, with arguments that fit a schema (many APIs offer a strict mode for tool arguments too). Both can be combined in one request.
Going further
Treat schemas as versioned contracts between the model and your code. When you change a schema, bump a version field, keep old parsers until traffic moves over, and rerun your evaluation set. Many incidents in AI features are not model failures but silent contract changes between components.
Key takeaways
- Prompted formats can break; tool calling is stronger; constrained decoding can guarantee structure but not truth.
- Use enums, explicit nulls, descriptive fields and an evidence field to reduce invention and ease auditing.
- Put evidence or reasoning fields before the final decision so the model grounds its answer first.
- Always validate in code, add semantic business-rule checks, and cap retries.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Design a JSON schema for one extraction task in your work, including enums, nullable fields, and an evidence field. Write down two business-rule checks that the schema cannot express.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.