Multimodal & Reasoning Models in PracticeReasoning models and test-time compute · Lesson 11 of 17

Reasoning models and test-time compute

Article · 12 min · 7 min lecture

Video lecture

Reasoning models and test-time compute

11 chapters · about 7 min · full transcript

Coming soon

Chapter 1 of 11

Reasoning models and test-time compute

  • Training-time vs test-time compute
  • How reasoning models learn
  • Cost and latency effects
  • Where it pays off

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Two ways to get a better answer

For years, the main route to better AI was training-time compute: bigger models trained on more data with more computation. A second route has become central: test-time compute (also called inference-time compute), meaning the model spends more computation while answering a particular question.

Reasoning models (often marketed with terms like "thinking", "extended thinking" or "reasoning effort") are trained to generate an internal chain of reasoning before their final answer: breaking the problem down, trying approaches, checking intermediate results and backtracking. Research and practice have shown that for many tasks, especially mathematics, coding, logic, planning and multi-step analysis, allowing more reasoning improves accuracy.

How they are trained, conceptually

Details vary by lab and are often not fully published, but a common ingredient is reinforcement learning on tasks where answers can be checked (maths problems with known answers, code with tests). The model is rewarded for reaching correct answers, and in the process learns reasoning strategies: decomposing problems, verifying, and revising. This is different from simply prompting a model to "think step by step", although it builds on that insight.

What changes for users

  • Latency: reasoning takes time. A reasoning model may take noticeably longer before its answer starts, depending on effort settings and task difficulty.
  • Cost: reasoning tokens are typically billed (usually as output tokens) even when you do not see them in full. Hard problems can consume many.
  • Controls: many APIs expose settings such as an effort level or a reasoning budget. Some models adapt their thinking to task difficulty automatically. Some offer a summary of their reasoning rather than the full trace.
  • Visibility: you may see full reasoning, a summary, or nothing, depending on the provider. Visible reasoning is useful for debugging but should not be treated as a guaranteed faithful account of how the answer was produced.

Where reasoning pays off

  • Multi-step maths and quantitative analysis.
  • Coding, debugging and code review.
  • Complex planning and scheduling with constraints.
  • Analysing contracts or policies against multiple rules.
  • Agentic tasks where the model must decide among many tools and steps.
  • Hard judgement questions where considering alternatives matters.

Where it usually does not

  • Simple lookups, rewriting, formatting and short classification.
  • Creative writing where fluency and voice matter more than logic (though planning can help structure long pieces).
  • Real-time conversation where latency dominates the experience.

The practical principle: match compute to difficulty. Paying for deep reasoning on easy tasks wastes money and time; skimping on hard tasks costs accuracy.

A simple illustration

Consider a scheduling request: "Arrange six client meetings next week across London and Dubai time zones, respecting each client's availability, no more than two meetings per day for our director, and travel days on Tuesday." A fast non-reasoning model may produce a plausible schedule that violates a constraint. A reasoning model is more likely to check constraints systematically, though you should still verify the result in code (constraint checks are easy to automate).

Limits and cautions

  • Not infallible: reasoning models still make errors, sometimes after long, confident reasoning.
  • Overthinking: on simple tasks, long reasoning can occasionally lead the model to second-guess a correct answer.
  • Benchmarks vs reality: headline benchmark gains may not transfer to your tasks. Test on your own data (module 6).
  • Fast-moving area: model releases, controls and pricing change frequently; revisit your choices periodically.

The shape of the trade-off today

Across providers, reasoning has become a dial rather than a separate product. Many current flagship models reason by default and let you set effort; some decide adaptively how much to think per request. Three consequences for planning:

  • The same model can be your fast model and your careful model, at different effort settings. Before building a two-model cascade, test the stronger model at low effort; it sometimes matches a cheaper model at similar cost with one fewer thing to maintain.
  • Cost now depends on difficulty. Hard prompts consume more reasoning tokens. Budget with distributions (median and slowest cases), not averages.
  • Latency varies more. A hard request may take far longer than an easy one on the same settings. Design user experiences with streaming, progress indicators or background jobs.

Hands-on: see test-time compute in your own numbers

Run the same set of questions at two effort levels and compare output tokens, latency and accuracy:

import os, time, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")

QUESTIONS = [
    ("A shop sells 3 abayas at 180, 240 and 310 AED. With 15% off the two cheapest and "
     "free shipping over 500 AED, what is the total for one of each?", "..."),  # add expected answers
    # ... 10-20 more, mixing easy and hard
]

for effort in ["low", "high"]:
    tokens = secs = 0
    for q, _expected in QUESTIONS:
        t0 = time.perf_counter()
        r = client.messages.create(model=MODEL, max_tokens=16000,
                                   thinking={"type": "adaptive"},
                                   output_config={"effort": effort},
                                   messages=[{"role": "user", "content": q}])
        secs += time.perf_counter() - t0
        tokens += r.usage.output_tokens   # includes thinking tokens
    print(effort, f"avg output tokens={tokens / len(QUESTIONS):.0f}", f"avg secs={secs / len(QUESTIONS):.1f}")

Add an accuracy check against expected answers, and you have a small test-time-compute curve for your own workload.

Test-time compute beyond one model call

Self-consistency (sample several answers and take the majority), best-of-n with a verifier, and tool-assisted checking (running code or tests) are all ways to buy accuracy with compute. They multiply cost, so reserve them for high-value, verifiable tasks such as financial calculations, code changes and compliance checks.

Going further

Test-time compute also includes techniques outside a single model's reasoning: generating several candidate answers and selecting by majority vote (self-consistency), using a verifier model or tests to pick the best candidate, or searching over solution paths. These can be combined with reasoning models for high-stakes tasks, at multiplied cost. Evaluate whether the accuracy gain justifies the spend for each use case.

Key takeaways

  • Test-time compute means spending more computation while answering; reasoning models are trained to reason before answering.
  • Reasoning improves accuracy on maths, code, planning, rule-heavy analysis and agentic tasks.
  • It costs latency and tokens; controls like effort levels let you match compute to difficulty.
  • Reasoning models still err; verify results and test on your own tasks.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is test-time compute?
  2. Which task benefits most from a reasoning model?
  3. How are reasoning tokens typically billed?

Put it into practice

List five tasks you use AI for. Classify each as 'needs reasoning', 'maybe' or 'no', and test one 'maybe' task at two effort levels.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.