---
title: "Structured outputs: reliable JSON from every provider"
description: "Why structured outputs Most business uses of LLMs end in code: save fields to a database, route a ticket, render a card, call an API. Free text is…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/structured-outputs-json-schema
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Core API calls across providers · lesson 6 of 19 · 15 min

# Structured outputs: reliable JSON from every provider

## Why structured outputs

Most business uses of LLMs end in code: save fields to a database, route a ticket, render a card, call an API. Free text is fragile to parse. **Structured outputs** constrain the model to produce JSON that matches a schema you provide, so your code can trust the shape.

Three levels of reliability:

1. **Prompting for JSON** ("reply only with JSON like ..."): easy but can break (extra prose, missing fields, invalid JSON).
2. **JSON mode**: guarantees valid JSON but not your schema.
3. **Schema-constrained output**: guarantees the response matches your JSON Schema. All three major providers offer this for current models; use it for production.

## How each provider does it

**Claude**: `output_config={"format": {"type": "json_schema", "schema": {...}}}` on `messages.create`, or the SDK helper `client.messages.parse(..., output_format=PydanticModel)` which returns `response.parsed_output`. (An older top-level `output_format` parameter on `create` is deprecated in favor of `output_config`.) Structured outputs are incompatible with some features such as citations; check the docs.

**OpenAI (Responses API)**: `text={"format": {"type": "json_schema", "name": "...", "schema": {...}, "strict": True}}`, or the helper `client.responses.parse(..., text_format=PydanticModel)` with `response.output_parsed`. (In Chat Completions the equivalent lives under `response_format`.)

**Gemini**: `GenerateContentConfig(response_mime_type="application/json", response_json_schema=...)`; you can pass a Pydantic model's `model_json_schema()`.

## Hands-on: one Pydantic model, three providers

A lead-qualification extractor for a UK B2B agency reading inbound emails:

```python
import json, os
from typing import Literal
from pydantic import BaseModel, Field
import anthropic
from openai import OpenAI
from google import genai
from google.genai import types

class Lead(BaseModel):
    company: str
    contact_name: str
    budget_gbp: int | None = Field(description="Stated budget in GBP, null if not stated")
    service: Literal["seo", "paid_social", "web_design", "other"]
    urgency: Literal["low", "medium", "high"]
    needs_human: bool = Field(description="True if the email is unclear, a complaint, or legal/sensitive")

EMAIL = """Hi, I'm Tom from Greenleaf Interiors in Leeds. We need a new website before our spring launch
in March, budget around £12k. Can we talk this week?"""
SYSTEM = "Extract lead details from the email. Do not guess missing values; use null."

# Claude: SDK parse helper with a Pydantic model
c = anthropic.Anthropic().messages.parse(
    model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=1000, system=SYSTEM,
    messages=[{"role": "user", "content": EMAIL}], output_format=Lead)
print("Claude:", c.parsed_output)

# OpenAI: Responses parse helper
o = OpenAI().responses.parse(model=os.environ.get("OPENAI_MODEL", "gpt-5.5"),
                             instructions=SYSTEM, input=EMAIL, text_format=Lead)
print("OpenAI:", o.output_parsed)

# Gemini: JSON schema from the Pydantic model, then validate
g = genai.Client().models.generate_content(
    model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=EMAIL,
    config=types.GenerateContentConfig(system_instruction=SYSTEM, response_mime_type="application/json",
                                       response_json_schema=Lead.model_json_schema()))
print("Gemini:", Lead.model_validate_json(g.text))
```

Always validate on your side too (Pydantic `model_validate`), because schema support differs in edge cases and business rules (like "budget must be positive") go beyond JSON Schema.

## Schema design tips

- **Enums over free text** for anything you route on (`service`, `urgency`).
- **Nullable fields** for information that may be absent, with instructions not to guess. Otherwise models fill gaps with plausible inventions.
- **An escape hatch** such as `needs_human` or `confidence` so the model can signal uncertainty instead of forcing a wrong value.
- **Descriptions on fields**: they are part of the prompt.
- **Keep it flat and small**: deeply nested, huge schemas cost tokens and increase errors; split into two calls if needed.
- **Portable subset**: stick to widely supported JSON Schema features (types, enums, required, arrays, nested objects). Check each provider's documented limitations for things like recursive schemas, numeric ranges or string patterns.

## Structured outputs vs tool calling

Use **structured outputs** when you want the model's final answer in a shape. Use **tools** when the model needs to *take an action* or fetch data mid-reasoning. Many extraction pipelines that previously forced a fake tool call to get JSON should switch to structured outputs; on some of the newest models, forcing a specific tool is no longer supported, which makes structured outputs the right mechanism.

## Worked example: routing 5,000 inbound emails a month

The agency routes each email by `service` and `urgency` to the right team channel; `needs_human` goes to a manager. Validation failures (rare) go to a retry queue with a different model. Weekly, a sample of 50 extractions is reviewed by a human to measure field accuracy. After adding the nullable budget field and "do not guess" instruction, invented budgets disappeared from the sample (illustrative).

## Handling the rare failure

Even with constrained decoding, requests can fail: the output hits `max_tokens` mid-object, the model refuses, a network error interrupts, or your own business validation rejects a value. Build a small failure path: check the stop reason before parsing, catch validation errors, retry once (optionally with a different model), and send anything still failing to a human queue with the raw input attached. Log the failure reason so you can see patterns, such as one field that often fails validation and needs a clearer description.

## Pitfalls

- Parsing free text with regex in production.
- Required fields for information that is often missing (forces hallucination).
- Trusting the schema guarantee as a guarantee of **truth**: the JSON can be perfectly shaped and still wrong.
- Mixing structured outputs with incompatible features without checking docs.

## Measuring success

Schema-validation failure rate (should be near zero with constrained decoding), field-level accuracy against human labels, `needs_human` rate, and downstream error rates (misrouted tickets).

## Video lecture: Structured outputs: reliable JSON from every provider

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Structured outputs
2. Why it matters
3. Three reliability levels
4. Provider mechanisms
5. Simple example: restaurant reviews
6. Under the hood
7. Business example: UK agency leads
8. Schema design tips
9. Shape ≠ truth
10. Structured outputs vs tools
11. Common mistakes
12. Text plus fields
13. Performance notes
14. Deeper: 5,000 emails/month (illustrative)
15. Watch me do it: the Lead extractor
16. Recap + try this now

## Lecture transcript

### Structured outputs

Here's a quiet truth about AI in business: most outputs aren't read by people. They're read by code. A ticket gets routed, a record gets saved, a dashboard gets updated. And code hates surprises. In this lesson you'll learn structured outputs, which force a model's answer into the exact JSON shape you define, and how to do it with Claude, OpenAI and Gemini.

### Why it matters

Why does this matter? Because parsing free text is fragile. The model adds a friendly sentence before the JSON, or forgets a field, or writes high priority instead of high, and your code breaks at two in the morning. Think of a paper form versus a blank sheet of paper. If you ask someone to write their details on a blank sheet, you get every format imaginable. Give them a form with boxes and tick options, and every answer lands in the right place.

### Three reliability levels

There are three levels of reliability. Level one: asking nicely for JSON in the prompt. It usually works, until it doesn't. Level two: JSON mode, which guarantees valid JSON, but not your fields. Level three: schema constrained output, where the provider guarantees the response matches your JSON schema. All three major providers support level three on current models, and that's what you should use in production.

### Provider mechanisms

How does each provider do it? Claude takes an output config with a JSON schema format, and the SDK's parse helper accepts a Pydantic model and returns a parsed object. OpenAI's Responses API takes a text format with a JSON schema and strict mode, and its parse helper also accepts a Pydantic model. Gemini takes a response mime type of application JSON and a response JSON schema, which you can generate from the same Pydantic model. One model definition, three providers.

### Simple example: restaurant reviews

A simple example. You want to turn restaurant reviews into three fields: rating from one to five, a sentiment of positive, neutral or negative, and the dish mentioned, or null if none. With a schema, every review becomes a tidy row. Without one, you'll see four stars in one answer, eight out of ten in another, and a paragraph in a third. The schema makes your spreadsheet trustworthy.

### Under the hood

Let's look at what happens under the hood, because it explains the guarantee. With schema constrained output, the provider restricts which tokens the model is allowed to produce at each step, so it literally cannot write a field name that isn't in your schema or a value outside your enum. Think of a vending machine keypad that only lights up valid product codes. You can't press a code that doesn't exist. That's why constrained output is so much more reliable than asking nicely, and why the schema itself deserves careful design.

### Business example: UK agency leads

Now the realistic business example in the lesson. A UK B2B agency receives about five thousand inbound emails a month. A Pydantic model called Lead defines company, contact name, budget in pounds or null, the service from a fixed list, urgency, and a needs human flag. The system prompt says: don't guess missing values. The code runs the same model through Claude, OpenAI and Gemini. Each email is routed by service and urgency, and anything flagged needs human goes to a manager.

### Schema design tips

Schema design is where most of the quality comes from. Use enums for anything you route on. Make fields nullable when information may be missing, and say don't guess, otherwise models fill gaps with plausible inventions. Add an escape hatch, like needs human or a confidence level. Describe each field, because descriptions are part of the prompt. Keep schemas small and flat. And stick to widely supported schema features, so the same schema works across providers.

### Shape ≠ truth

Here's a subtle but important point. A schema guarantees shape, not truth. The JSON can be perfectly formed and still wrong: the wrong company name, the wrong urgency. So validate on your side with Pydantic for business rules, like budget must be positive, and measure field accuracy by having a person label a sample every week. In the agency example, adding the nullable budget field and the don't guess instruction made invented budgets disappear from the weekly sample.

### Structured outputs vs tools

When should you use structured outputs versus tool calling? Structured outputs shape the model's final answer. Tools let the model take an action or fetch data in the middle of its work. Many older pipelines forced a fake tool call just to get JSON. Switch those to structured outputs, especially since some of the newest models no longer allow forcing a specific tool.

### Common mistakes

Common mistakes. Parsing free text with regular expressions in production. Making fields required when the information is often missing, which forces hallucination. Treating valid JSON as correct data. And combining structured outputs with features that don't support them, like some citation features, without checking the docs.

### Text plus fields

A question from product teams: what about outputs that are mostly text, like an email, but need a few fields too? Put the text inside the schema as a string field, next to the structured fields, like subject, body, tone and needs review. You keep one call, one validated object, and your app can route or store the metadata while displaying the body. Just set sensible length limits on long text fields, so a runaway answer can't blow your token budget.

### Performance notes

A quick note on performance and cost. The first time you use a new schema, some providers need a moment to prepare it, so the very first request can be slower; later requests reuse that work. Large schemas also add input tokens to every call. Keep schemas stable, reuse them across requests, and split a giant extraction into two smaller calls if one schema grows past a few dozen fields. Your latency charts will thank you.

### Deeper: 5,000 emails/month (illustrative)

Let's deepen the UK agency's email routing with illustrative numbers. Five thousand inbound emails a month. Before structured outputs, a regex parser failed on about one email in twenty, usually when the model added a friendly sentence before the JSON, and failures landed in a manual queue. After switching to schema constrained output with the Lead model, parsing failures went to practically zero. The weekly human sample of fifty extractions showed high accuracy on company and service, and the needs human flag caught complaints and legal threats that the old system routed to sales by mistake. The operations lead's comment: we stopped noticing the AI, which is exactly what we wanted.

### Watch me do it: the Lead extractor

Watch me do it. Let's walk through the lead extractor. The Lead model has company, contact name, an optional budget in pounds with a description telling the model to use null if it isn't stated, a service limited to four values, urgency limited to three, and a needs human flag for unclear, complaint or legal emails. The system prompt says: extract lead details; don't guess missing values. For Claude, messages parse takes output format equal to the Lead class and returns parsed output as a Lead object. For OpenAI, responses parse takes text format equal to Lead and returns output parsed. For Gemini, I pass the JSON schema generated from the Lead model with the JSON mime type, then validate the text with model validate json. On the sample email from Tom at Greenleaf Interiors in Leeds, all three return web design, budget twelve thousand, urgency high, needs human false. Then I delete the budget sentence and rerun: budget comes back null.

### Recap + try this now

Quick recap. Structured outputs give you guaranteed shapes. Each provider has a mechanism, and one Pydantic model can drive all three. Design schemas with enums, nullable fields and an escape hatch, validate business rules yourself, and measure accuracy. Try this now: define a Pydantic model for one extraction task in your business, run it through two providers on twenty real samples, and measure how many fields are right.

## Key takeaways

- Structured outputs constrain responses to your JSON Schema so code can trust the shape.
- Claude: output_config format or messages.parse; OpenAI: text.format or responses.parse; Gemini: response_json_schema.
- Design schemas with enums, nullable fields, an escape hatch and field descriptions.
- Validate on your side and measure field accuracy: valid shape is not the same as truth.
- Use structured outputs for final answers and tools for actions.

## Try it

Define a Pydantic model for one extraction task in your business, run it through two providers on 20 real samples, and measure field-level accuracy.

- [Previous: Tool calling across Claude, OpenAI and Gemini](https://optimizeall.com/learn/ai-platform-apis-integration/tool-calling-across-providers)
- [Next: Multimodal inputs: images, PDFs and audio](https://optimizeall.com/learn/ai-platform-apis-integration/vision-documents-and-audio)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
