Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 8 of 20
Structured outputs and strict tool use
Video lecture
Structured outputs and strict tool use
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Structured outputs and strict tools
Every automation builder has seen it. You ask the model to reply in JSON. Nine hundred and ninety times it does. Then once it adds a friendly sentence before the brace, or misses a field, and the next step in your workflow falls over at two in the morning. Structured outputs fix the format problem at the source. In this lesson you'll learn what they guarantee, what they don't, and how to design schemas that give you data you can actually trust.
0:36 Analogy: describing a form vs handing one over
Here's an analogy. Asking a model to please reply in JSON is like asking a new colleague to fill in a form by describing it over the phone. Most of the time it works. Structured outputs are like handing them the actual form, with boxes, drop-downs and required fields. They can still write the wrong answer in a box, but they can't skip a box, invent a new one, or scribble in the margin.
1:08 Two features
Most major model APIs now let you supply a JSON Schema, or a typed model in your SDK like a Pydantic class, and constrain generation so the response validates against it. That's structured outputs. The same idea applies to tool calls: with strict tool use, the arguments the model passes to your tool are guaranteed to match the tool's input schema. Anthropic's API, OpenAI's Responses API and Google's Gemini API all offer versions of this. Field names differ, but the concept is identical.
1:44 Guaranteed vs not
What's guaranteed? When the request completes normally, the output parses, required fields exist, types and enums are respected, and no surprise properties appear if you forbid them. What's not guaranteed? Truth: a perfectly valid invoice total can be the wrong number. Completion: if the response hits the token limit or the model declines, you may not have a valid object, so always check the stop reason. And honesty about unknowns: if your schema forces a value, the model has to invent one.
2:20 Schema design
So design schemas that allow honesty. Make fields nullable, or add an explicit not-found status. Use enums for categories, with an escape value like other or needs review. Keep schemas flat and small, and split big jobs across two calls. Write field descriptions, because they work like instructions: ISO date, null if not stated. For high-stakes extraction, add an evidence field with short quotes from the source supporting each value. Reviewers can check it in seconds, and hallucinations become visible.
2:55 Simple example: review sentiment
A simple example. You ask for the sentiment of a review, with a schema allowing only positive, neutral or negative. Without structure, the model might reply mostly positive, with some concerns about delivery, which your code can't route. With structured output, it must pick one of the three values. Add a reason field, and you get both: a clean label for your workflow, and a short explanation a human can read when checking.
3:27 Worked example: enquiry extraction
Here's a worked example. A creator agency receives brand enquiries by email in English and Arabic. Their schema has a category enum with other, nullable brand, budget, currency and deadline, a list of deliverables, and evidence quotes. Before structured outputs, a small share of responses failed to parse and went to a retry queue. Afterwards, parsing failures disappear. The work that remains is accuracy, which they measure on a hundred labelled emails, with the evidence quotes making human review fast.
4:02 Business example (illustrative)
Deeper into the agency case, illustrative. They receive about four hundred enquiries a month. Before structured outputs, roughly one in thirty failed parsing, and those sat in a retry queue for hours. After, parsing failures vanished. Field accuracy on a hundred labelled emails was about ninety-five percent for category and eighty-eight for budget. Adding a clearer budget description and allowing null lifted budget accuracy to about ninety-four percent, mostly by stopping invented budgets.
4:34 Hands-on in the lesson
In the lesson's hands-on section you'll use the Claude API's parse helper with a Pydantic model to extract an enquiry, check the stop reason, and apply a business rule in code. You'll also see a strict tool definition for creating CRM deals, with additional properties forbidden, so an agent can never call it with malformed arguments. Swap in your own email, and try one with no budget to confirm you get null rather than an invention.
5:07 Measure accuracy per field
How do you measure success once the format is guaranteed? Build a labelled set of fifty to a hundred real inputs, and score each field separately: exact match for categories and dates, tolerance for amounts, and a separate check for fields that should be null. Track the share of records a reviewer accepts without edits, and read the evidence quotes on every miss. You'll usually find that one or two fields cause most errors, and that clearer field descriptions or an extra enum value fix them.
5:44 Common mistakes
Common mistakes, in more detail. Nesting deep objects inside objects until the model struggles to fill them well. Forgetting additional properties false, so unexpected fields slip through. Using structured outputs where you actually need citations, which some APIs don't combine. Not versioning the schema, so a downstream system breaks when a field is renamed. And skipping accuracy evaluation because the JSON always parses now.
6:12 How you'll know it's working
How will you know it's working? Parsing and validation errors drop to zero. Field-level accuracy on your labelled set meets the threshold you agreed. The share of null values matches reality, meaning the model says unknown when information really is missing, rather than inventing it. And reviewers accept most records without edits, using the evidence quotes to check quickly.
6:38 Watch me do it: typed extraction
Watch me do it. I open the extraction code. First, the Enquiry model: category is a Literal with five values including other; brand, budget, currency and deadline are Optional; evidence is a list of strings with a description. Next, extract calls messages dot parse with the system instruction never guess, use null, and passes the Enquiry model as the output format. Then I check the stop reason. If it isn't end turn, I return None, because a truncated or refused response can't be trusted. Next, a business rule in code: a negative budget is rejected. I run it on the GlowLab email. The result shows category brand deal, budget fifteen thousand, currency AED, two deliverables, and evidence quotes for each value. Then I run an email with no budget, and budget comes back as None rather than a made-up figure. Finally, I show the strict create deal tool: additional properties false, so no stray fields.
7:45 Recap
Avoid four pitfalls. Treating schema-valid as correct. Forcing values that may not exist. Ignoring the stop reason and using a partial result. And giant schemas that try to extract, classify and draft all at once. To recap: structured outputs and strict tools remove format failures, but truth still needs evaluation and business rules in code. Your next step is to pick one extraction or classification step you run, write its schema with nullable fields, an escape value and evidence, and test it on twenty real inputs.
8:22 Try this now (25 minutes)
Try this now. Pick one classification or extraction step you run today, even if it's done by hand. Write a schema for it with a category enum that includes other, nullable fields for anything that might be missing, and an evidence field. Run it on ten real inputs with the hands-on code or your platform's structured output option. Count how many values are correct, not just valid. That's your accuracy baseline.
Why "please return JSON" is not enough
For years, teams asked models to "reply in JSON" and then wrote defensive code for the times they did not: missing fields, trailing commentary, invalid enums, truncated objects. Downstream steps broke, retries piled up and costs rose. Today most major model APIs offer structured outputs: you supply a JSON Schema (or a typed model in your SDK), and the API constrains generation so the response validates against the schema. The same idea applies to tool calls through strict tool use, where the model's tool arguments are guaranteed to match the tool's input schema.
This lesson explains what these features guarantee, what they do not, and how to design schemas that give you reliable, useful data.
Two related features
| Feature | What it constrains | Typical use |
|---|---|---|
| Structured outputs (JSON Schema response format) | The model's final response body | Extraction, classification, scoring, generating records for other systems |
| Strict tool use | The arguments of tool calls | Agents and workflows where tool inputs must be valid before your code runs them |
On Anthropic's API, structured outputs are set with output_config.format (or the SDK's messages.parse() helper with a Pydantic or Zod model), and strict tools with "strict": true on a tool definition. OpenAI's Responses API and Google's Gemini API offer equivalent schema-constrained output. Field names differ; the concept is identical. Check each provider's docs for supported JSON Schema features, because not every keyword is supported everywhere.
What is guaranteed, and what is not
Guaranteed (when the feature is on and the request completes normally): the output parses, required fields exist, types are right, enums are respected and no extra properties appear (when you set additionalProperties: false).
Not guaranteed:
- Truth. A schema-valid invoice total can still be the wrong number. Validate business rules in code and evaluate accuracy.
- Completion. If the response hits the token limit or the model declines, you may not get a valid object. Check the stop reason before using the result.
- Good judgement about "unknown". If your schema forces a value, the model must fill it. Give it a legal way to say "not found".
Designing good schemas
- Allow honest uncertainty. Use nullable fields or an explicit
"status": "not_found"option. A schema that forces abudgetnumber invites invented budgets. - Prefer enums for categories and include an escape value (
OTHER,NEEDS_REVIEW). - Keep it flat and small. Deeply nested, many-field schemas are harder to fill well. Split into two calls if needed.
- Describe fields. Field descriptions in the schema act like instructions ("ISO 8601 date; null if not stated").
- Add evidence fields for high-stakes extraction, such as a short quote from the source supporting each value. This makes review fast and hallucinations visible.
- Version your schemas like APIs; downstream systems depend on them.
Worked example: lead extraction for a creator agency
An agency receives brand enquiries by email in English and Arabic. The extraction schema:
{
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["BRAND_DEAL", "PR_GIFTING", "FAN", "SPAM", "OTHER"]},
"brand": {"type": ["string", "null"]},
"budget": {"type": ["number", "null"], "description": "Amount only if explicitly stated"},
"currency": {"type": ["string", "null"], "enum": ["AED", "SAR", "PKR", "GBP", "USD", null]},
"deliverables": {"type": "array", "items": {"type": "string"}},
"deadline": {"type": ["string", "null"], "description": "ISO 8601 date if stated"},
"evidence": {"type": "array", "items": {"type": "string"}, "description": "Short quotes supporting each extracted value"}
},
"required": ["category", "brand", "budget", "currency", "deliverables", "deadline", "evidence"],
"additionalProperties": false
}Before structured outputs, roughly one in a few dozen responses (illustrative) failed parsing and went to a retry queue. After, parsing failures disappear; the remaining work is accuracy, measured on 100 labelled emails, with the evidence quotes speeding up human review.
Hands-on: typed extraction with the Claude API
import os
from typing import Literal, Optional
import anthropic
from pydantic import BaseModel, Field
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
class Enquiry(BaseModel):
category: Literal["BRAND_DEAL", "PR_GIFTING", "FAN", "SPAM", "OTHER"]
brand: Optional[str]
budget: Optional[float] = Field(description="Only if explicitly stated")
currency: Optional[Literal["AED", "SAR", "PKR", "GBP", "USD"]]
deliverables: list[str]
deadline: Optional[str] = Field(description="ISO 8601 date if stated")
evidence: list[str] = Field(description="Short quotes supporting each extracted value")
def extract(email_text: str) -> Enquiry | None:
resp = client.messages.parse(
model=MODEL, max_tokens=2000,
system="Extract enquiry details. Never guess: use null when a value is not stated.",
messages=[{"role": "user", "content": email_text}],
output_format=Enquiry,
)
if resp.stop_reason != "end_turn": # e.g. max_tokens or refusal: do not trust
return None
enquiry = resp.parsed_output
if enquiry.budget is not None and enquiry.budget < 0: # business rule in code
return None
return enquiry
print(extract("Hi! We're GlowLab (skincare). Budget AED 15,000 for 2 Reels + 3 Stories, live by 20 Oct."))And a strict tool definition, so an agent can never call create_deal with malformed arguments:
{
"name": "create_deal",
"description": "Create a CRM deal for a qualified brand enquiry. Drafts only; a human confirms.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {
"brand": {"type": "string"},
"stage": {"type": "string", "enum": ["new", "qualified"]},
"value": {"type": ["number", "null"]}
},
"required": ["brand", "stage", "value"],
"additionalProperties": false
}
}Pitfalls
- Treating schema-valid as correct; you still need accuracy evaluation.
- Forcing values that may not exist, which manufactures hallucinations.
- Ignoring the stop reason and using a partial result.
- Huge schemas that try to do extraction, classification and drafting in one go.
Key takeaways
- Structured outputs and strict tool use constrain responses and tool arguments to validate against your schema.
- Schema-valid is not the same as correct: validate business rules in code and evaluate accuracy.
- Design schemas that allow honest uncertainty, use enums with escape values, stay small and include evidence fields.
- Always check the stop reason before trusting a structured result.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick one extraction or classification step in your work. Write its schema with nullable fields, an escape enum value and an evidence field, then test it on 20 real inputs.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.