Skip to content

Advanced Prompt Engineering · Decomposition, reasoning and verification · lesson 6 of 17 · 13 min

Reasoning models: when step-by-step helps and when it hurts

Two generations of "think step by step"

Early prompt engineering discovered that asking a model to reason step by step before answering (chain-of-thought prompting) improved accuracy on maths, logic and multi-step problems. Since then, many providers have released reasoning models (often described as having "thinking", "extended thinking" or adjustable "reasoning effort"). These models are trained to produce internal reasoning before their final answer, and they spend extra computation at answer time, a concept often called test-time compute.

This changes the prompting playbook. Techniques that helped older models can be redundant, or even counterproductive, with reasoning models. This area is moving quickly, so treat the guidance below as durable principles and always check your provider's current documentation.

When explicit step-by-step helps

  • Non-reasoning or fast models on multi-step tasks: asking the model to work through the problem in a tagged section before answering still helps.
  • Auditable reasoning: you want a visible, structured rationale for a human reviewer (for example "list the policy clauses you relied on, then give the decision").
  • Domain procedures: when there is a required procedure, such as a regulatory checklist, spelling out those steps ensures they happen.

A clean pattern for non-reasoning models:

Work through the problem inside <analysis> tags: identify the relevant facts,
check each eligibility rule, and note any missing information.
Then give only the final decision inside <decision> tags.

Your code can then show only the decision to the user and keep the analysis for review.

When it hurts or wastes money

  • Reasoning models already reason. Adding a rigid, hand-written step list can constrain a model whose own reasoning would have been better. Provider guidance for recent reasoning models suggests that high-level instructions (for example "consider this thoroughly and check edge cases") often outperform prescriptive step-by-step scripts.
  • Simple tasks. Classification of obvious cases, reformatting, or short rewrites gain little from reasoning but pay in latency and tokens.
  • Over-verification. Some newer models verify their own work well by default. Prompts carried over from older models that demand repeated self-checks can add tokens and latency without improving accuracy. Re-test inherited instructions when you switch models.
  • Rationalisation. Visible reasoning is not guaranteed to reflect how the model actually reached its answer. Treat it as a useful artefact, not proof.

Controlling reasoning effort

Many APIs now expose a control for how much the model reasons, such as an effort level or a reasoning budget, and some models decide adaptively how much to think based on task difficulty. The practical approach:

  1. Start with a moderate or default setting.
  2. Build a small evaluation set (see the evaluation module) that includes easy and hard cases.
  3. Compare accuracy, latency and cost at two or three effort levels.
  4. Route: low effort for easy, high-volume tasks; higher effort for hard, high-stakes ones.

Worked example: pricing eligibility

A subscription business must decide whether a customer qualifies for a loyalty discount under five rules. With a fast non-reasoning model, the team uses an analysis-then-decision template, which improves consistency. When they switch to a reasoning model, they remove the step script, keep the five rules and a clear definition of done, and ask for a brief justification citing which rules applied. Accuracy stays the same or improves, and the prompt gets shorter. Because the task is high volume, they then test a lower effort level and find it sufficient for most customers, escalating only ambiguous cases.

Few-shot with reasoning models

Examples still help reasoning models, particularly for output format and domain conventions. Some providers note that including example reasoning (for instance in thinking-style tags inside examples) can shape the style of the model's own reasoning. Keep examples diverse so you do not constrain the model to one path.

Evaluating "thinking" choices

Do not decide by intuition. For any change (adding step-by-step, removing it, changing effort), measure on a fixed evaluation set:

  • task accuracy or rubric score
  • median and slow-tail latency
  • tokens and cost per request
  • failure types (did errors change character?)

Hands-on: thinking controls in three APIs

The names differ, but the idea is the same: tell the model how much effort to spend. Check each provider's documentation for current parameters, supported models and defaults, which change between model generations.

Claude. Current Claude models use adaptive thinking: the model decides how much to think, and an effort setting steers the depth and overall token spend. On the newest models the older fixed budget_tokens setting is deprecated or rejected.

import os, anthropic
client = anthropic.Anthropic()

def ask(question: str, effort: str = "medium") -> str:
    resp = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"),
        max_tokens=16000,
        thinking={"type": "adaptive"},
        output_config={"effort": effort},   # low | medium | high | xhigh | max
        messages=[{"role": "user", "content": question}],
    )
    return "".join(b.text for b in resp.content if b.type == "text")

OpenAI (Responses API). Reasoning models accept an effort setting:

from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
    model=os.environ["OPENAI_MODEL"],
    reasoning={"effort": "low"},   # see the docs for the values your model supports
    input="Which of these five loyalty rules apply to customer #88? ...",
)
print(resp.output_text)

Gemini. Gemini 3 models expose a thinking level (for example low or high; some models add minimal or medium), replacing the older numeric thinking budget, which remains for backward compatibility.

Hands-on: an effort sweep

Decide effort with data, not instinct:

import time, statistics
CASES = [...]  # (question, expected) pairs from your eval set, easy and hard mixed

for effort in ["low", "medium", "high"]:
    correct, latencies = 0, []
    for q, expected in CASES:
        t0 = time.perf_counter()
        answer = ask(q, effort)
        latencies.append(time.perf_counter() - t0)
        correct += int(expected.lower() in answer.lower())
    print(effort, f"acc={correct / len(CASES):.0%}",
          f"p50={statistics.median(latencies):.1f}s", f"max={max(latencies):.1f}s")

Log token usage from each response too, since thinking tokens are billed as output. Pick the lowest effort that meets your accuracy bar per task type, and route hard cases upward.

Displaying and storing reasoning

Providers differ in what they return: a summary of the reasoning, an empty placeholder, or encrypted blocks you must pass back unchanged in multi-turn tool use. Do not show raw reasoning to end users as if it were an explanation, and never parse it for business logic.

Going further

If your provider returns reasoning content (or a summary of it), log it for debugging, but do not build product logic that parses it; formats and availability of reasoning output vary between providers and change over time. Build your contracts on the final, structured answer.

Video lecture: Reasoning models: when step-by-step helps and when it hurts

Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.

  1. Reasoning models: when step-by-step helps
  2. Two generations
  3. When explicit steps help
  4. When it hurts
  5. Effort controls today
  6. The effort sweep
  7. Worked example: loyalty rules
  8. Practical notes
  9. Example 1: stacked discounts
  10. Example 2: renewal quote checks (illustrative)
  11. Recap

Lecture transcript

Reasoning models: when step-by-step helps

For years, the most famous prompt trick was: let's think step by step. Then models arrived that think before answering by default, and that trick started, in some cases, to get in the way. In this lecture you will learn when explicit step-by-step reasoning still helps, when it wastes money or even hurts, how reasoning effort is controlled in the Claude, OpenAI and Gemini APIs today, and how to choose effort levels with data instead of instinct.

Two generations

Early prompt engineering discovered that asking a model to reason step by step before answering improved accuracy on maths, logic and multi-step problems. That is chain-of-thought prompting. Since then, providers have released reasoning models, described as having thinking, extended thinking, adaptive thinking or adjustable reasoning effort. They are trained to reason before the final answer, spending extra computation at answer time. That is often called test-time compute. It changes the playbook: techniques that helped older models can be redundant or even counterproductive now.

When explicit steps help

Explicit step-by-step still helps in three situations. With fast or non-reasoning models on multi-step tasks, asking for analysis in a tagged section before a final decision still improves consistency. When you need auditable reasoning, a visible rationale for a human reviewer, such as list the policy clauses you relied on, then give the decision. And when there is a required procedure, like a regulatory checklist, spelling out those steps ensures they happen. Your code can then show the decision to the user and keep the analysis for review.

When it hurts

Now when it hurts or wastes money. Reasoning models already reason, and a rigid hand-written step list can constrain a model whose own approach would have been better. Provider guidance for recent reasoning models suggests high-level instructions, like consider this thoroughly and check edge cases, often beat prescriptive scripts. Simple tasks, like obvious classifications and short rewrites, gain little from reasoning but pay in latency and tokens. Some newer models verify their own work well by default, so inherited instructions demanding repeated self-checks add cost without accuracy. And remember: visible reasoning is not guaranteed to reflect how the model actually reached its answer.

Effort controls today

Let us look at the controls. On current Claude models you use adaptive thinking, where the model decides how much to think, and an effort setting, from low through medium and high up to higher levels on the newest models, steers depth and overall token spend. The older fixed thinking-token budget is deprecated or rejected on the newest models, so check the migration guides. In OpenAI's Responses API, reasoning models take a reasoning effort setting. And Gemini three models expose a thinking level, replacing the older numeric thinking budget, which is kept for backward compatibility.

The effort sweep

How do you pick a level? Run an effort sweep. Take an eval set that mixes easy and hard cases. Run it at low, medium and high effort. Record accuracy, median latency, the slowest latencies, and tokens, because thinking tokens are billed as output. Then pick the lowest effort that meets your accuracy bar for each task type, and route the hard, high-stakes cases upward. The lesson's hands-on code does exactly this in about fifteen lines.

Worked example: loyalty rules

Here is a worked example. A subscription business decides whether a customer qualifies for a loyalty discount under five rules. With a fast model, an analysis-then-decision template improves consistency. When they switch to a reasoning model, they remove the step script, keep the five rules and a clear definition of done, and ask for a brief justification citing which rules applied. Accuracy holds or improves and the prompt gets shorter. Because it is high volume, they test a lower effort level, find it sufficient for most customers, and escalate only ambiguous cases.

Practical notes

A few practical notes. Examples still help reasoning models, especially for format and domain conventions; keep them diverse. Providers differ in what they return from thinking: a readable summary, an empty placeholder, or encrypted blocks that you must pass back unchanged in multi-turn tool use. Do not show raw reasoning to users as an explanation, and never parse it for business logic. Build your contracts on the final, structured answer. And monitor reasoning token usage; spikes often signal ambiguous or contradictory instructions that you can fix in the prompt.

Example 1: stacked discounts

A simple worked example. Ask a fast model: a shirt costs forty pounds, it is twenty-five percent off, and there is a further ten percent off at the till; what do I pay? A rushed answer might say thirty-five percent off, so twenty-six pounds. Wrong. The discounts stack: forty minus twenty-five percent is thirty, minus ten percent is twenty-seven. With a fast model, asking it to work through the calculation in an analysis section before giving the answer fixes this. With a reasoning model, you do not need to ask; it already works it out. And for either, if the number matters, have code do the arithmetic.

Example 2: renewal quote checks (illustrative)

Now a business scenario, with illustrative numbers. An insurance broker in Manchester uses AI to check about four hundred policy renewal quotes a week against eight eligibility and pricing rules. On their evaluation set of one hundred and fifty cases, a fast model with an analysis-then-decision template scores around eighty-eight percent. A reasoning model at medium effort, with the rules and a clear definition of done but no step script, scores around ninety-six percent. At high effort it scores ninety-seven, but average latency more than doubles and cost rises sharply. The team chooses medium effort, and routes the few cases the model flags as ambiguous to a human underwriter. They also remove an old instruction, check every rule three times, which had been making responses longer without improving accuracy. Numbers illustrative; the decision process is what matters.

Recap

To recap. Explicit chain-of-thought still helps fast models, auditable workflows and mandatory procedures. With reasoning models, prefer clear goals and constraints over rigid scripts, and remove redundant self-check demands. Tune effort per task with a sweep that measures accuracy, latency and tokens. Try this now: pick a multi-step task and run fifteen cases with no reasoning instruction, with a step script, and with high-level guidance or a higher effort setting. Record accuracy, latency and cost. Next: self-critique and verification loops.

Key takeaways

  • Chain-of-thought helps non-reasoning models on multi-step problems and when you need auditable rationale.
  • Reasoning models already think; prefer high-level guidance over rigid step scripts, and remove redundant self-check instructions.
  • Tune reasoning effort per task using an evaluation set that measures accuracy, latency and cost.
  • Visible reasoning is useful but is not proof of how the model decided; build contracts on the final answer.

Try it

Pick a multi-step task. Run it on 15 cases with (a) no reasoning instruction, (b) a step script, (c) high-level guidance or a higher effort setting. Record accuracy, latency and cost for each.