Building Production AI Agents · Tools, reasoning and multi-agent design · lesson 5 of 18 · 15 min
Planning, reasoning models and effort control
From "think step by step" to built-in reasoning
Early agents relied on prompting tricks such as "think step by step" or the ReAct pattern (interleaving Reasoning and Acting in text). Today's frontier models from Anthropic, OpenAI and Google have built-in reasoning: they can spend extra tokens thinking before (and between) tool calls, and the APIs expose controls for how much. This changes how you design agents: less scaffolding in prompts, more attention to how much reasoning each step deserves.
The controls you will meet
- Anthropic (Claude): recent models use adaptive thinking (
thinking: {"type": "adaptive"}), where the model decides how much to think, and an effort setting (output_config: {"effort": "low" | "medium" | "high" | ...}) that trades thoroughness against tokens. Older models used a fixedbudget_tokens; on current models that parameter is deprecated or rejected, so check the model's migration notes. Thinking can interleave with tool calls, so the model reasons about each result. - OpenAI: reasoning models accept a reasoning effort setting in the Responses API, and reasoning items can be carried between turns.
- Google Gemini: thinking-capable models accept a thinking configuration in
GenerateContentConfig.
Parameter names and allowed values change between model generations. Always confirm against the current docs for the exact model you deploy. The durable concept: reasoning depth is now a tunable dial with a cost.
Choosing effort by task
| Step type | Suggested starting effort | Why | |---|---|---| | Classification, routing, extraction | Low | Little reasoning needed; latency matters | | Customer chat, simple tool use | Low to medium | Responsiveness beats depth | | Multi-step research, analysis | Medium to high | Plans and cross-checks pay off | | Coding, long-horizon agentic work | High or above | Errors compound over many steps |
Measure, don't guess: run your eval set at two effort levels and compare success rate and cost per completed task (not per request; a cheaper setting that needs more retries is not cheaper).
Planning patterns that still matter
Built-in reasoning does not remove the need for explicit plans when tasks are long:
- Plan-then-execute: ask the model to produce a numbered plan (as structured output), validate it in code (e.g., no forbidden steps, fewer than 10 steps), then execute step by step. Good for auditability.
- Re-planning: after each major step, let the model revise the remaining plan given new information.
- Todo lists in state: keep a checklist in the agent's state or memory and have the model tick items off. Coding agents use this heavily; it keeps long runs on track.
- Reflection: after a draft or failed attempt, prompt the model to critique against explicit criteria, then retry. This is the evaluator–optimizer pattern applied inside an agent.
Worked example: a Karachi e-commerce analytics agent
Task: "Find why conversion dropped on mobile last week and propose fixes."
- Step 1 (effort medium): produce a plan: check traffic mix, device breakdown, checkout funnel, recent releases, payment-provider errors.
- Code validates the plan (all steps map to available tools).
- Steps 2–5 (effort low for data pulls, medium for interpretation).
- Final synthesis (effort high): correlate the drop with a release that changed the checkout form and a spike in payment timeouts from one provider; propose two fixes with expected impact labeled as hypotheses.
The team saves cost by using high effort only where judgment matters.
Hands-on: plan-then-execute with adaptive thinking
import os
import anthropic
from pydantic import BaseModel, Field
client = anthropic.Anthropic()
MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5") # confirm model + thinking support in docs
ALLOWED = {"get_traffic", "get_funnel", "get_releases", "get_payment_errors"}
class Step(BaseModel):
tool: str
purpose: str
class Plan(BaseModel):
steps: list[Step] = Field(max_length=8)
def make_plan(goal: str) -> Plan:
resp = client.messages.parse( # structured output validated against the Pydantic model
model=MODEL,
max_tokens=4000,
thinking={"type": "adaptive"},
output_config={"effort": "medium"},
system=f"Plan using only these tools: {sorted(ALLOWED)}.",
messages=[{"role": "user", "content": goal}],
output_format=Plan,
)
plan = resp.parsed_output
bad = [s.tool for s in plan.steps if s.tool not in ALLOWED]
if bad:
raise ValueError(f"Plan uses unknown tools: {bad}")
return plan
plan = make_plan("Mobile conversion dropped last week. Find likely causes.")
for i, s in enumerate(plan.steps, 1):
print(i, s.tool, "-", s.purpose)
Execution then proceeds with the normal loop from lesson 3, passing the plan as context. Keep the plan in state so a crash can resume from the last completed step (module 3).
Budgets that the model can see
max_tokens caps a single response, and the model does not know about it. Some providers now offer task budgets (Anthropic has a beta for agentic loops) that tell the model how many tokens remain so it can pace itself and finish gracefully. Where unavailable, you can inject a short status line each turn ("Budget: 3 tool calls left"). Either way, enforce hard limits in code too.
Pitfalls
- Maxing effort everywhere "to be safe"; costs and latency balloon for no measurable gain.
- Asking for long visible reasoning in the answer when the model already reasons internally; it wastes output tokens.
- Treating a plan as immutable when reality changes mid-run.
- Stripping thinking blocks from history when continuing a tool loop, which some APIs require you to pass back unchanged.
Measuring success
For each step type, compare success rate, tokens and latency across effort levels on the same eval set. Plot cost per completed task. Pick the lowest setting that holds quality, and revisit when you change models.
Video lecture: Planning, reasoning models and effort control
Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.
- Planning and reasoning models
- Why reasoning controls matter
- Reasoning controls by provider
- Effort by step type
- Planning patterns
- Simple example: clinic front desk
- Example: mobile conversion drop
- Hands-on: validated plan
- Budgets and pitfalls
- Show reasoning in the answer?
- Deeper: effort by step pays off
- Watch me do it: make_plan()
- Try this now
- Recap
Lecture transcript
Planning and reasoning models
Remember when every prompt ended with, let's think step by step? Today's leading models from Anthropic, OpenAI and Google reason internally, before and between tool calls, and the APIs let you control how much. In this lesson you'll learn how to tune reasoning depth for each step, and the planning patterns that still matter for long tasks.
Why reasoning controls matter
Why should you care about reasoning controls? Because they're now one of the biggest drivers of both quality and cost. Think of it like choosing how long to spend on an email. A quick thank you needs ten seconds. A sensitive message to a client needs ten minutes. Spending ten minutes on every email wastes your day; spending ten seconds on the sensitive one creates problems. Reasoning effort is the same trade off for a model. Set it per step, not once for the whole system, and you'll get better answers for less money.
Reasoning controls by provider
Each provider exposes reasoning differently. Claude's recent models use adaptive thinking, where the model decides how much to think, plus an effort setting that trades thoroughness against tokens. OpenAI's reasoning models take a reasoning effort setting in the Responses API. Gemini's thinking models take a thinking configuration. The names and allowed values change between model generations, so always check the current docs for the exact model you ship. The durable idea is simple: reasoning depth is now a dial, and turning it up costs money and time.
Effort by step type
So how much reasoning does each step deserve? Classification, routing and extraction usually need little, so start low. Customer chat and simple tool use: low to medium, because responsiveness matters. Multi step research and analysis: medium to high. Coding and long running agent work: high or above, because mistakes compound over many steps. But don't take a table's word for it. Run your eval set at two settings and compare cost per completed task, not cost per request. A cheaper setting that fails more often isn't cheaper.
Planning patterns
Built in reasoning doesn't remove the need for explicit plans on long tasks. Plan then execute: the model writes a numbered plan as structured data, your code checks it, then you execute step by step. Re planning: after big steps, let the model revise what's left. Todo lists: keep a checklist in the agent's state and tick items off, which keeps long runs on track. And reflection: after a failed attempt, critique against clear criteria and try again.
Simple example: clinic front desk
A simple example. You have a two step assistant for a clinic's front desk. Step one classifies each incoming message: booking, cancellation, or question for a doctor. Step two drafts a reply. Classification is easy, so run it at low effort and it's fast and cheap. Drafting a careful, accurate reply to a medical question deserves more thought, so use medium or high effort there, plus a human check. When you test both steps at two settings, you'll usually find classification accuracy doesn't move, while the reply quality does. Spend where it shows.
Example: mobile conversion drop
Here's a worked example from an e commerce team in Karachi. Mobile conversion dropped last week. The agent first writes a plan at medium effort: check traffic mix, device breakdown, checkout funnel, recent releases and payment errors. Code confirms every step maps to a real tool. Data pulls run at low effort. Interpretation runs at medium. The final synthesis runs at high effort and links the drop to a checkout form release and a spike in timeouts from one payment provider, with fixes presented as hypotheses to test.
Hands-on: validated plan
In the lesson's hands on code, the model returns its plan as a validated Pydantic object using structured outputs, with adaptive thinking and medium effort. Then your code rejects any step that uses a tool that doesn't exist. That's a small check with a big payoff: bad plans die before they cost you a single tool call. Keep the plan in the agent's state too, so a crash can resume from the last completed step.
Budgets and pitfalls
One more control: budgets. Max tokens limits a single response, and the model doesn't know it's there. Newer options, like Anthropic's task budgets beta, tell the model how many tokens remain so it can pace itself and wrap up gracefully. If your provider lacks that, add a short status line each turn, like three tool calls left. Either way, enforce hard limits in code. And avoid the classic pitfalls: maximum effort everywhere, asking for long visible reasoning you're paying for twice, and stripping thinking blocks that the API expects you to pass back.
Show reasoning in the answer?
Here's a question that comes up a lot: should you ask the model to show its reasoning in the answer? Usually not. Current reasoning models already think internally, and asking for long visible reasoning makes you pay twice, once for the thinking and again for the explanation. If users need an explanation, ask for a short, clear rationale in plain language, like three bullets on why this recommendation. If you need to debug, many APIs can return a summarized view of the thinking, which is separate from the answer and easy to hide from users.
Deeper: effort by step pays off
Let's deepen the Karachi e-commerce example. The analytics agent ran every Monday across four product categories. At high effort for every step, a full run cost noticeably more and took several minutes, illustrative figures. When the team measured, the data pulling steps gave identical results at low effort, and only the final synthesis benefited from high effort. After splitting effort by step, cost per completed report fell by more than half and runs finished in under two minutes, with no change in the analysts' quality ratings. Just as important, the validated plan meant every report showed which checks were run, so when the head of growth asked, did you look at payment errors, the answer was right there in the plan.
Watch me do it: make_plan()
Watch me do it. Let's walk through the plan then execute code. At the top, allowed lists the four tools the agent may use. Two small Pydantic models define a step, a tool and a purpose, and a plan, a list of up to eight steps. In make plan, I call messages parse with the model, adaptive thinking, and effort set to medium. The system prompt lists the allowed tools, and the output format is the plan model, so the response comes back as a validated plan object. Then comes the most important line: I collect any step whose tool isn't in the allowed set, and if there are any, I raise an error before a single tool runs. Finally, the plan prints as numbered steps. In a real system, I'd store this plan in the run state, and the execution loop would tick steps off, so a crash resumes at the right place.
Try this now
Try this now. Take your five baseline goals from the loop lesson. Run each at two effort settings on the same model, and record success, total tokens and seconds. Put the numbers in a small table, and compute cost per completed task using prices from your provider's pricing page. Then make one decision: which steps get low effort, and which get more. Write that decision down with the evidence next to it. Next time a new model launches, rerun the same table instead of guessing.
Recap
To recap: reasoning is built in, so your job is tuning it. Match effort to each step, validate plans in code, re plan when facts change, and compare settings by cost per completed task. Your next step: run your five baseline goals at two effort levels, record success, tokens and latency, and choose a setting for each step type.
Key takeaways
- Frontier models reason internally; your job is to tune how much reasoning each step gets.
- Effort and thinking parameters differ by provider and model generation, so verify against current docs.
- Use low effort for routing and extraction, higher effort for analysis and long-horizon work.
- Plan-then-execute with code validation gives auditability; re-plan when facts change.
- Judge settings by cost per completed task, not cost per request.
Try it
Run your five baseline goals from lesson 3 at two effort levels. Record success, tokens and latency, and choose a setting per step type.