Multimodal & Reasoning Models in PracticeReasoning models and test-time compute · Lesson 13 of 17

Extended and adaptive thinking: controls across APIs

Article · 13 min · 8 min lecture

Video lecture

Extended and adaptive thinking: controls across APIs

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Extended and adaptive thinking

  • What the terms mean now
  • Controls in three APIs
  • When it's worth it
  • Sweeps, timeouts, multi-turn

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What "extended thinking" means now

Extended thinking describes a model spending extra tokens reasoning before (and sometimes between) its visible answers. Early implementations asked you to set a fixed thinking budget in tokens. Current flagship models increasingly use adaptive thinking: the model decides how much to think per request, guided by an effort setting. The practical controls in 2026:

ProviderControlNotes
Anthropic Claudethinking: {"type": "adaptive"} plus output_config.effort (low to max)Fixed budget_tokens is deprecated or rejected on the newest models; thinking output is summarised or omitted, never the raw chain of thought
OpenAIreasoning.effort in the Responses APIReasoning summaries can be requested; values vary by model
Google Geminithinking_level on Gemini 3 modelsOlder numeric thinking_budget kept for backward compatibility; do not use both

Always confirm parameter names and defaults for your exact model; they have changed between generations and will change again.

When extended thinking is worth it

Use more thinking when the task has many interacting constraints or steps, errors are costly, and the answer can be checked:

  • Multi-step financial or pricing analysis, schedules, allocation problems.
  • Code changes, debugging and review.
  • Contract or policy analysis against several rules.
  • Agentic work with many tools, where planning quality compounds.

Use little or none for:

  • Rewriting, formatting, tone changes, translation of short text.
  • Simple classification and extraction at high volume.
  • Real-time voice turns, where latency dominates.

Hands-on: one function, three providers

import os

def think_claude(prompt, effort="medium"):
    import anthropic
    client = anthropic.Anthropic()
    r = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=16000,
        thinking={"type": "adaptive", "display": "summarized"},
        output_config={"effort": effort},
        messages=[{"role": "user", "content": prompt}])
    summary = [b.thinking for b in r.content if b.type == "thinking"]
    answer = "".join(b.text for b in r.content if b.type == "text")
    return answer, summary, r.usage.output_tokens

def think_openai(prompt, effort="medium"):
    from openai import OpenAI
    r = OpenAI().responses.create(model=os.environ["OPENAI_MODEL"], input=prompt,
                                  reasoning={"effort": effort, "summary": "auto"})
    return r.output_text, None, r.usage.output_tokens

def think_gemini(prompt, level="low"):
    from google import genai
    from google.genai import types
    client = genai.Client()
    r = client.models.generate_content(
        model=os.environ["GEMINI_MODEL"], contents=prompt,
        config=types.GenerateContentConfig(
            thinking_config=types.ThinkingConfig(thinking_level=level)))
    return r.text, None, r.usage_metadata.total_token_count

Use these in your effort sweep: same questions, several settings, record accuracy, latency and tokens. Reasoning tokens are billed as output tokens, so the token column is your cost signal.

Long outputs and timeouts

High effort on hard problems can run for minutes and produce long outputs. Use streaming for long requests (SDKs may require it for very large maximum token settings), set generous but finite timeouts, and give users progress feedback. For agents, set turn, time and spend limits in code as well.

Thinking in multi-turn and tool use

When a model thinks between tool calls, providers may return thinking blocks (sometimes encrypted or signed) that must be passed back unchanged in the next request for the conversation to continue correctly. Keep your loop append-only: add the model's full content blocks to history rather than editing them. Do not show thinking content to end users as an authoritative explanation, and do not parse it for decisions.

Worked example: an agency's media-plan checker

An agency in Riyadh checks media plans against client rules (budget caps per channel, flighting dates, minimum frequency) before sending them to clients. At low effort the model misses about one in ten violations on their test set (illustrative figure); at high effort it catches nearly all, at several times the tokens and latency. They run high effort only on plans above a spend threshold and low effort for small plans, and they re-run the sweep each quarter because newer models often reach the same accuracy at lower effort.

Hands-on: route effort by difficulty

Once your sweep shows where extra thinking pays, encode the decision instead of leaving it to each developer:

ROUTES = {
    "caption_rewrite":   {"effort": "low"},
    "invoice_extract":   {"effort": "low"},
    "media_plan_check":  {"effort": "medium"},
    "contract_review":   {"effort": "high"},
}

def effort_for(route: str, spend_aed: float = 0.0) -> str:
    effort = ROUTES.get(route, {"effort": "medium"})["effort"]
    if route == "media_plan_check" and spend_aed >= 500_000:   # high-stakes plans get more thinking
        effort = "high"
    return effort

Keep this table in configuration, log the effort used on every request, and review it when models change.

Common mistakes

  • Setting the highest effort globally "to be safe", then discovering latency and cost problems in production.
  • Copying an old fixed thinking budget into code for a model that now rejects it.
  • Comparing effort levels on a handful of easy questions, where every level looks the same.
  • Showing thinking summaries to customers as if they were the reasons for a decision.

How to measure success

Your decision table per route: accuracy at each effort level, median and slowest latency, output tokens and cost per successful outcome. Choose the lowest setting that meets the bar, route exceptions upward, and re-test after model upgrades.

Key takeaways

  • Current models use adaptive thinking steered by effort (Claude), reasoning effort (OpenAI) or thinking level (Gemini 3).
  • Use more thinking for multi-constraint, costly, checkable tasks; little for rewriting, simple extraction or voice turns.
  • Reasoning tokens bill as output; decide effort with sweeps of accuracy, latency and tokens.
  • Stream long requests, keep multi-turn loops append-only, and never parse thinking for decisions.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. On the newest Claude models, how should you control how much the model thinks?
  2. Which task is least likely to benefit from high thinking effort?
  3. In a multi-turn tool loop with thinking enabled, what should you do with the model's thinking blocks?

Put it into practice

Run an effort sweep on 15 real tasks across two settings (and two providers if you can), recording accuracy, latency and output tokens, and write a one-paragraph routing decision.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.