Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 8 of 20

Structured outputs and strict tool use

Article · 13 min · 9 min lecture

Video lecture

Structured outputs and strict tool use

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Structured outputs and strict tools

  • Format guaranteed at the source
  • What it does and doesn't promise
  • Schemas that invite honesty

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why "please return JSON" is not enough

For years, teams asked models to "reply in JSON" and then wrote defensive code for the times they did not: missing fields, trailing commentary, invalid enums, truncated objects. Downstream steps broke, retries piled up and costs rose. Today most major model APIs offer structured outputs: you supply a JSON Schema (or a typed model in your SDK), and the API constrains generation so the response validates against the schema. The same idea applies to tool calls through strict tool use, where the model's tool arguments are guaranteed to match the tool's input schema.

This lesson explains what these features guarantee, what they do not, and how to design schemas that give you reliable, useful data.

FeatureWhat it constrainsTypical use
Structured outputs (JSON Schema response format)The model's final response bodyExtraction, classification, scoring, generating records for other systems
Strict tool useThe arguments of tool callsAgents and workflows where tool inputs must be valid before your code runs them

On Anthropic's API, structured outputs are set with output_config.format (or the SDK's messages.parse() helper with a Pydantic or Zod model), and strict tools with "strict": true on a tool definition. OpenAI's Responses API and Google's Gemini API offer equivalent schema-constrained output. Field names differ; the concept is identical. Check each provider's docs for supported JSON Schema features, because not every keyword is supported everywhere.

What is guaranteed, and what is not

Guaranteed (when the feature is on and the request completes normally): the output parses, required fields exist, types are right, enums are respected and no extra properties appear (when you set additionalProperties: false).

Not guaranteed:

  • Truth. A schema-valid invoice total can still be the wrong number. Validate business rules in code and evaluate accuracy.
  • Completion. If the response hits the token limit or the model declines, you may not get a valid object. Check the stop reason before using the result.
  • Good judgement about "unknown". If your schema forces a value, the model must fill it. Give it a legal way to say "not found".

Designing good schemas

  1. Allow honest uncertainty. Use nullable fields or an explicit "status": "not_found" option. A schema that forces a budget number invites invented budgets.
  2. Prefer enums for categories and include an escape value (OTHER, NEEDS_REVIEW).
  3. Keep it flat and small. Deeply nested, many-field schemas are harder to fill well. Split into two calls if needed.
  4. Describe fields. Field descriptions in the schema act like instructions ("ISO 8601 date; null if not stated").
  5. Add evidence fields for high-stakes extraction, such as a short quote from the source supporting each value. This makes review fast and hallucinations visible.
  6. Version your schemas like APIs; downstream systems depend on them.

Worked example: lead extraction for a creator agency

An agency receives brand enquiries by email in English and Arabic. The extraction schema:

{
  "type": "object",
  "properties": {
    "category": {"type": "string", "enum": ["BRAND_DEAL", "PR_GIFTING", "FAN", "SPAM", "OTHER"]},
    "brand": {"type": ["string", "null"]},
    "budget": {"type": ["number", "null"], "description": "Amount only if explicitly stated"},
    "currency": {"type": ["string", "null"], "enum": ["AED", "SAR", "PKR", "GBP", "USD", null]},
    "deliverables": {"type": "array", "items": {"type": "string"}},
    "deadline": {"type": ["string", "null"], "description": "ISO 8601 date if stated"},
    "evidence": {"type": "array", "items": {"type": "string"}, "description": "Short quotes supporting each extracted value"}
  },
  "required": ["category", "brand", "budget", "currency", "deliverables", "deadline", "evidence"],
  "additionalProperties": false
}

Before structured outputs, roughly one in a few dozen responses (illustrative) failed parsing and went to a retry queue. After, parsing failures disappear; the remaining work is accuracy, measured on 100 labelled emails, with the evidence quotes speeding up human review.

Hands-on: typed extraction with the Claude API

import os
from typing import Literal, Optional
import anthropic
from pydantic import BaseModel, Field

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

class Enquiry(BaseModel):
    category: Literal["BRAND_DEAL", "PR_GIFTING", "FAN", "SPAM", "OTHER"]
    brand: Optional[str]
    budget: Optional[float] = Field(description="Only if explicitly stated")
    currency: Optional[Literal["AED", "SAR", "PKR", "GBP", "USD"]]
    deliverables: list[str]
    deadline: Optional[str] = Field(description="ISO 8601 date if stated")
    evidence: list[str] = Field(description="Short quotes supporting each extracted value")

def extract(email_text: str) -> Enquiry | None:
    resp = client.messages.parse(
        model=MODEL, max_tokens=2000,
        system="Extract enquiry details. Never guess: use null when a value is not stated.",
        messages=[{"role": "user", "content": email_text}],
        output_format=Enquiry,
    )
    if resp.stop_reason != "end_turn":           # e.g. max_tokens or refusal: do not trust
        return None
    enquiry = resp.parsed_output
    if enquiry.budget is not None and enquiry.budget < 0:   # business rule in code
        return None
    return enquiry

print(extract("Hi! We're GlowLab (skincare). Budget AED 15,000 for 2 Reels + 3 Stories, live by 20 Oct."))

And a strict tool definition, so an agent can never call create_deal with malformed arguments:

{
  "name": "create_deal",
  "description": "Create a CRM deal for a qualified brand enquiry. Drafts only; a human confirms.",
  "strict": true,
  "input_schema": {
    "type": "object",
    "properties": {
      "brand": {"type": "string"},
      "stage": {"type": "string", "enum": ["new", "qualified"]},
      "value": {"type": ["number", "null"]}
    },
    "required": ["brand", "stage", "value"],
    "additionalProperties": false
  }
}

Pitfalls

  • Treating schema-valid as correct; you still need accuracy evaluation.
  • Forcing values that may not exist, which manufactures hallucinations.
  • Ignoring the stop reason and using a partial result.
  • Huge schemas that try to do extraction, classification and drafting in one go.

Key takeaways

  • Structured outputs and strict tool use constrain responses and tool arguments to validate against your schema.
  • Schema-valid is not the same as correct: validate business rules in code and evaluate accuracy.
  • Design schemas that allow honest uncertainty, use enums with escape values, stay small and include evidence fields.
  • Always check the stop reason before trusting a structured result.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your schema requires a numeric budget, and many emails do not mention one. What is the likely result?
  2. What does strict tool use guarantee?
  3. A structured response stopped because it hit the token limit. What should your code do?

Put it into practice

Pick one extraction or classification step in your work. Write its schema with nullable fields, an escape enum value and an evidence field, then test it on 20 real inputs.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.