Multimodal & Reasoning Models in PracticeReasoning models and test-time compute · Lesson 13 of 17
Extended and adaptive thinking: controls across APIs
Video lecture
Extended and adaptive thinking: controls across APIs
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Extended and adaptive thinking
Every major model API now lets you decide how hard the model should think. The parameter names differ, the defaults differ, and some older settings are now rejected outright. Get it wrong and you either pay for deep reasoning on trivial tasks, or get shallow answers on hard ones. In this lecture you will learn what extended and adaptive thinking mean today, the controls in the Claude, OpenAI and Gemini APIs, when extended thinking is worth it, how to run a cross-provider effort sweep, and how to handle long outputs and multi-turn tool use.
0:41 Why thinking controls matter
Why does this matter? Because thinking settings are now one of the biggest levers on your bill and your response times, and they are easy to set once and forget. A default that made sense for one model can be wasteful or too shallow on the next. Think of effort like the gears on a bicycle: the right gear depends on the hill, and riding everything in the hardest gear just wears you out.
1:13 From budgets to effort
Extended thinking describes a model spending extra tokens reasoning before, and sometimes between, its visible answers. Early implementations asked you to set a fixed thinking budget in tokens. Current flagship models increasingly use adaptive thinking: the model decides how much to think per request, guided by an effort setting. So your job has shifted from choosing a number of thinking tokens to choosing an effort level per route, and measuring the result.
1:44 The controls
Here are the controls. On Anthropic's Claude, you set thinking to adaptive and choose an effort level, from low up to max, inside the output configuration. The fixed budget tokens setting is deprecated, or rejected, on the newest models, and the thinking output is a summary or omitted, never the raw chain of thought. On OpenAI, the Responses API takes a reasoning effort, and you can request reasoning summaries; values vary by model. On Google's Gemini three models, you set a thinking level; the older numeric thinking budget is kept for backward compatibility, and you should not use both in one request. Always confirm names and defaults for your exact model.
2:32 When to think harder
When is extended thinking worth it? When the task has many interacting constraints or steps, errors are costly, and the answer can be checked. Multi-step financial or pricing analysis, schedules and allocation problems. Code changes, debugging and review. Contract or policy analysis against several rules. And agentic work with many tools, where planning quality compounds. When is it not? Rewriting, formatting and tone changes. Simple classification and extraction at high volume. And real-time voice turns, where latency dominates.
3:06 Hands-on: one function per provider
The lesson gives you one function per provider with the same shape. The Claude version requests adaptive thinking with summarised display and an effort level, and returns the answer, the thinking summary and the output tokens. The OpenAI version sets reasoning effort with an automatic summary and returns the output text and output tokens. The Gemini version sets a thinking level in the generation config and returns the text and token count. Use them in an effort sweep: the same questions, several settings, recording accuracy, latency and tokens. Reasoning tokens are billed as output, so the token column is your cost signal.
3:50 Long outputs and timeouts
High effort on hard problems can run for minutes and produce long outputs. Use streaming for long requests; some SDKs require it for very large maximum token settings. Set generous but finite timeouts, and give users progress feedback. In agents, enforce turn, time and spend limits in code as well. Nothing is more expensive than a stuck loop at maximum effort.
4:17 Multi-turn and tools
Thinking in multi-turn conversations and tool use has one golden rule: keep history append-only. When a model thinks between tool calls, providers may return thinking blocks, sometimes encrypted or signed, that must be passed back unchanged in the next request for the conversation to continue correctly. Add the model's full content blocks to history rather than editing them. Do not show thinking content to end users as an authoritative explanation, and never parse it for business decisions.
4:50 Worked example: media-plan checks
A worked example from Riyadh. An agency checks media plans against client rules, budget caps per channel, flighting dates and minimum frequency, before sending them to clients. On their test set, low effort misses roughly one violation in ten, an illustrative figure, while high effort catches nearly all, at several times the tokens and latency. So they run high effort only for plans above a spend threshold, and low effort for small plans. And they re-run the sweep every quarter, because newer models often reach the same accuracy at lower effort.
5:30 Example 1: a homework helper
A simple worked example. You build a small app that helps students check maths homework, and you use the same model for two routes: explaining a concept in simple words, and checking a multi-step algebra solution. With the same high effort for both, explanations are slow and no better. With low effort for explanations and high effort for solution checks, explanations come back quickly, and the solution checker catches more mistakes where it matters. One model, two effort settings, matched to the task.
6:06 Example 2: financial promotion checks (illustrative)
Now a business scenario, with illustrative numbers. A compliance team at a payments company in London reviews about two thousand marketing messages a month against financial promotion rules before they go out. They run an effort sweep on one hundred and twenty labelled messages across two providers. At low effort, both providers miss around one in eight rule breaches. At high effort, one provider misses about one in forty, and the other about one in twenty-five, with the first taking roughly twice as long. Since messages are reviewed in a daily batch, latency is not critical, so they choose the first provider at high effort for anything promoting interest rates or returns, and low effort for simple announcements that have no financial claims. A compliance officer still signs off every flagged message. They re-run the sweep each quarter. Illustrative figures.
7:07 Recap
To recap. Current models use adaptive thinking steered by effort on Claude, reasoning effort on OpenAI, and thinking level on Gemini three. Think harder for multi-constraint, costly, checkable tasks, and lightly for wording and simple, high-volume work. Decide with sweeps that measure accuracy, latency and tokens. Stream long requests, keep tool loops append-only, and never parse thinking. Try this now: run an effort sweep on fifteen real tasks across two settings, and two providers if you can, then write a one-paragraph routing decision. Next module: computer-use agents.
What "extended thinking" means now
Extended thinking describes a model spending extra tokens reasoning before (and sometimes between) its visible answers. Early implementations asked you to set a fixed thinking budget in tokens. Current flagship models increasingly use adaptive thinking: the model decides how much to think per request, guided by an effort setting. The practical controls in 2026:
| Provider | Control | Notes |
|---|---|---|
| Anthropic Claude | thinking: {"type": "adaptive"} plus output_config.effort (low to max) | Fixed budget_tokens is deprecated or rejected on the newest models; thinking output is summarised or omitted, never the raw chain of thought |
| OpenAI | reasoning.effort in the Responses API | Reasoning summaries can be requested; values vary by model |
| Google Gemini | thinking_level on Gemini 3 models | Older numeric thinking_budget kept for backward compatibility; do not use both |
Always confirm parameter names and defaults for your exact model; they have changed between generations and will change again.
When extended thinking is worth it
Use more thinking when the task has many interacting constraints or steps, errors are costly, and the answer can be checked:
- Multi-step financial or pricing analysis, schedules, allocation problems.
- Code changes, debugging and review.
- Contract or policy analysis against several rules.
- Agentic work with many tools, where planning quality compounds.
Use little or none for:
- Rewriting, formatting, tone changes, translation of short text.
- Simple classification and extraction at high volume.
- Real-time voice turns, where latency dominates.
Hands-on: one function, three providers
import os
def think_claude(prompt, effort="medium"):
import anthropic
client = anthropic.Anthropic()
r = client.messages.create(
model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=16000,
thinking={"type": "adaptive", "display": "summarized"},
output_config={"effort": effort},
messages=[{"role": "user", "content": prompt}])
summary = [b.thinking for b in r.content if b.type == "thinking"]
answer = "".join(b.text for b in r.content if b.type == "text")
return answer, summary, r.usage.output_tokens
def think_openai(prompt, effort="medium"):
from openai import OpenAI
r = OpenAI().responses.create(model=os.environ["OPENAI_MODEL"], input=prompt,
reasoning={"effort": effort, "summary": "auto"})
return r.output_text, None, r.usage.output_tokens
def think_gemini(prompt, level="low"):
from google import genai
from google.genai import types
client = genai.Client()
r = client.models.generate_content(
model=os.environ["GEMINI_MODEL"], contents=prompt,
config=types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(thinking_level=level)))
return r.text, None, r.usage_metadata.total_token_countUse these in your effort sweep: same questions, several settings, record accuracy, latency and tokens. Reasoning tokens are billed as output tokens, so the token column is your cost signal.
Long outputs and timeouts
High effort on hard problems can run for minutes and produce long outputs. Use streaming for long requests (SDKs may require it for very large maximum token settings), set generous but finite timeouts, and give users progress feedback. For agents, set turn, time and spend limits in code as well.
Thinking in multi-turn and tool use
When a model thinks between tool calls, providers may return thinking blocks (sometimes encrypted or signed) that must be passed back unchanged in the next request for the conversation to continue correctly. Keep your loop append-only: add the model's full content blocks to history rather than editing them. Do not show thinking content to end users as an authoritative explanation, and do not parse it for decisions.
Worked example: an agency's media-plan checker
An agency in Riyadh checks media plans against client rules (budget caps per channel, flighting dates, minimum frequency) before sending them to clients. At low effort the model misses about one in ten violations on their test set (illustrative figure); at high effort it catches nearly all, at several times the tokens and latency. They run high effort only on plans above a spend threshold and low effort for small plans, and they re-run the sweep each quarter because newer models often reach the same accuracy at lower effort.
Hands-on: route effort by difficulty
Once your sweep shows where extra thinking pays, encode the decision instead of leaving it to each developer:
ROUTES = {
"caption_rewrite": {"effort": "low"},
"invoice_extract": {"effort": "low"},
"media_plan_check": {"effort": "medium"},
"contract_review": {"effort": "high"},
}
def effort_for(route: str, spend_aed: float = 0.0) -> str:
effort = ROUTES.get(route, {"effort": "medium"})["effort"]
if route == "media_plan_check" and spend_aed >= 500_000: # high-stakes plans get more thinking
effort = "high"
return effortKeep this table in configuration, log the effort used on every request, and review it when models change.
Common mistakes
- Setting the highest effort globally "to be safe", then discovering latency and cost problems in production.
- Copying an old fixed thinking budget into code for a model that now rejects it.
- Comparing effort levels on a handful of easy questions, where every level looks the same.
- Showing thinking summaries to customers as if they were the reasons for a decision.
How to measure success
Your decision table per route: accuracy at each effort level, median and slowest latency, output tokens and cost per successful outcome. Choose the lowest setting that meets the bar, route exceptions upward, and re-test after model upgrades.
Key takeaways
- Current models use adaptive thinking steered by effort (Claude), reasoning effort (OpenAI) or thinking level (Gemini 3).
- Use more thinking for multi-constraint, costly, checkable tasks; little for rewriting, simple extraction or voice turns.
- Reasoning tokens bill as output; decide effort with sweeps of accuracy, latency and tokens.
- Stream long requests, keep multi-turn loops append-only, and never parse thinking for decisions.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run an effort sweep on 15 real tasks across two settings (and two providers if you can), recording accuracy, latency and output tokens, and write a one-paragraph routing decision.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.