Multimodal & Reasoning Models in PracticeReasoning models and test-time compute · Lesson 11 of 17
Reasoning models and test-time compute
Video lecture
Reasoning models and test-time compute
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Reasoning models and test-time compute
For years, the main route to better AI was bigger models trained on more data. Then a second route became central: letting a model spend more computation while it answers a particular question. That is test-time compute, and it is why many models now think before they speak. In this lecture you will learn what test-time compute is, how reasoning models are trained, what changes for cost and latency, where extra reasoning pays off and where it does not, and how to see the trade-off in your own numbers.
0:39 How reasoning models work
Reasoning models, marketed with words like thinking, extended thinking or reasoning effort, generate an internal chain of reasoning before the final answer: breaking the problem down, trying approaches, checking intermediate results and backtracking. Research and practice show that for many tasks, especially maths, coding, logic, planning and multi-step analysis, allowing more reasoning improves accuracy. How are they trained? Details vary and are often unpublished, but a common ingredient is reinforcement learning on tasks where answers can be checked, like maths with known answers or code with tests. The model is rewarded for correct results, and learns strategies like decomposing, verifying and revising.
1:23 A dial, not a product
Reasoning has become a dial rather than a separate product. Many flagship models reason by default and let you set effort, and some decide adaptively how much to think per request. That has three consequences. The same model can be your fast model and your careful model, at different effort settings; before building a two-model cascade, test the stronger model at low effort. Cost now depends on difficulty, because hard prompts consume more reasoning tokens, which are usually billed as output even when you do not see them. And latency varies more, so design with streaming, progress indicators or background jobs.
2:07 What changes for users
What changes for users? Latency, because reasoning takes time. Cost, because reasoning tokens are billed. Controls, like effort levels, and on some models adaptive thinking. And visibility: depending on the provider you may see a summary of the reasoning, nothing, or encrypted blocks. Visible reasoning is useful for debugging, but it is not guaranteed to be a faithful account of how the answer was produced. Never treat it as proof.
2:37 Match compute to difficulty
Where does extra reasoning pay off? Multi-step maths and quantitative analysis. Coding, debugging and code review. Complex planning and scheduling with constraints. Analysing contracts or policies against multiple rules. Agentic tasks where the model chooses among many tools and steps. And hard judgement calls where considering alternatives matters. Where does it usually not? Simple lookups, rewriting, formatting and short classification. Creative writing where voice matters more than logic. And real-time conversation, where latency dominates. The principle: match compute to difficulty.
3:12 Illustration: constrained scheduling
Here is an illustration. Arrange six client meetings next week across London and Dubai time zones, respecting each client's availability, with no more than two meetings a day for our director and travel days on Tuesday. A fast model may produce a plausible schedule that quietly violates a constraint. A reasoning model is more likely to check constraints systematically. But you should still verify the result in code, because constraint checks are easy to automate. Reasoning improves the odds; verification closes the gap.
3:48 See it in your numbers
You can see test-time compute in your own numbers. The lesson's hands-on code runs the same set of questions, mixing easy and hard, at low and high effort on a Claude model with adaptive thinking. It records the average output tokens, which include thinking tokens, and the average time per question. Add an accuracy check against expected answers, and you have your own small test-time-compute curve. You will often find that easy questions barely change, while hard ones improve, at a clear cost in tokens and seconds.
4:26 Beyond one call
Test-time compute also includes techniques beyond one model's reasoning. Self-consistency samples several answers and takes the majority. Best-of-n uses a verifier or tests to pick the best candidate. Tool-assisted checking runs code or tests. These multiply cost, so reserve them for high-value, verifiable tasks such as financial calculations, code changes and compliance checks. And keep your cautions: reasoning models still make errors, sometimes after long confident reasoning; they can overthink simple tasks; and benchmark gains may not transfer to your data.
5:01 Example 1: the two-jug puzzle
A simple worked example. Ask a model: I have a three-litre jug and a five-litre jug; how do I measure exactly four litres? A fast answer sometimes skips steps or gets the order wrong. A reasoning model tends to try a sequence, check the amounts after each pour, and fix mistakes before answering: fill the five, pour into the three, leaving two; empty the three; pour the two in; fill the five again; top up the three, leaving four. That checking-as-you-go is exactly what test-time compute buys you on multi-step problems.
5:41 Example 2: delivery route planning (illustrative)
Now a business scenario, with illustrative numbers. A logistics firm in Jeddah plans daily delivery routes for twenty vans with time windows and weight limits. They test a fast model and a reasoning model on fifty historical days, asking for a plan and then checking every constraint in code. The fast model produces plans that break at least one constraint on about thirty percent of days. The reasoning model at medium effort breaks constraints on about six percent, at roughly four times the tokens and three times the latency. Because planning runs once a night, latency does not matter, so they choose the reasoning model. They keep the code checker, and any plan that still breaks a constraint is sent back once with the violations listed, then escalated to a planner. Nightly planning time for staff falls from two hours to about twenty minutes of review. Illustrative figures.
6:45 Recap
To recap. Test-time compute means spending more computation while answering. Reasoning models are trained to think first, often with reinforcement learning on checkable tasks. Reasoning is now a dial, so cost and latency depend on difficulty and settings. Use more reasoning for maths, code, constrained planning, rule-heavy analysis and agents, and verify results. Try this now: list five tasks you use AI for, classify each as needs reasoning, maybe, or no, and test one maybe task at two effort levels. Next: working with reasoning models in practice.
Two ways to get a better answer
For years, the main route to better AI was training-time compute: bigger models trained on more data with more computation. A second route has become central: test-time compute (also called inference-time compute), meaning the model spends more computation while answering a particular question.
Reasoning models (often marketed with terms like "thinking", "extended thinking" or "reasoning effort") are trained to generate an internal chain of reasoning before their final answer: breaking the problem down, trying approaches, checking intermediate results and backtracking. Research and practice have shown that for many tasks, especially mathematics, coding, logic, planning and multi-step analysis, allowing more reasoning improves accuracy.
How they are trained, conceptually
Details vary by lab and are often not fully published, but a common ingredient is reinforcement learning on tasks where answers can be checked (maths problems with known answers, code with tests). The model is rewarded for reaching correct answers, and in the process learns reasoning strategies: decomposing problems, verifying, and revising. This is different from simply prompting a model to "think step by step", although it builds on that insight.
What changes for users
- Latency: reasoning takes time. A reasoning model may take noticeably longer before its answer starts, depending on effort settings and task difficulty.
- Cost: reasoning tokens are typically billed (usually as output tokens) even when you do not see them in full. Hard problems can consume many.
- Controls: many APIs expose settings such as an effort level or a reasoning budget. Some models adapt their thinking to task difficulty automatically. Some offer a summary of their reasoning rather than the full trace.
- Visibility: you may see full reasoning, a summary, or nothing, depending on the provider. Visible reasoning is useful for debugging but should not be treated as a guaranteed faithful account of how the answer was produced.
Where reasoning pays off
- Multi-step maths and quantitative analysis.
- Coding, debugging and code review.
- Complex planning and scheduling with constraints.
- Analysing contracts or policies against multiple rules.
- Agentic tasks where the model must decide among many tools and steps.
- Hard judgement questions where considering alternatives matters.
Where it usually does not
- Simple lookups, rewriting, formatting and short classification.
- Creative writing where fluency and voice matter more than logic (though planning can help structure long pieces).
- Real-time conversation where latency dominates the experience.
The practical principle: match compute to difficulty. Paying for deep reasoning on easy tasks wastes money and time; skimping on hard tasks costs accuracy.
A simple illustration
Consider a scheduling request: "Arrange six client meetings next week across London and Dubai time zones, respecting each client's availability, no more than two meetings per day for our director, and travel days on Tuesday." A fast non-reasoning model may produce a plausible schedule that violates a constraint. A reasoning model is more likely to check constraints systematically, though you should still verify the result in code (constraint checks are easy to automate).
Limits and cautions
- Not infallible: reasoning models still make errors, sometimes after long, confident reasoning.
- Overthinking: on simple tasks, long reasoning can occasionally lead the model to second-guess a correct answer.
- Benchmarks vs reality: headline benchmark gains may not transfer to your tasks. Test on your own data (module 6).
- Fast-moving area: model releases, controls and pricing change frequently; revisit your choices periodically.
The shape of the trade-off today
Across providers, reasoning has become a dial rather than a separate product. Many current flagship models reason by default and let you set effort; some decide adaptively how much to think per request. Three consequences for planning:
- The same model can be your fast model and your careful model, at different effort settings. Before building a two-model cascade, test the stronger model at low effort; it sometimes matches a cheaper model at similar cost with one fewer thing to maintain.
- Cost now depends on difficulty. Hard prompts consume more reasoning tokens. Budget with distributions (median and slowest cases), not averages.
- Latency varies more. A hard request may take far longer than an easy one on the same settings. Design user experiences with streaming, progress indicators or background jobs.
Hands-on: see test-time compute in your own numbers
Run the same set of questions at two effort levels and compare output tokens, latency and accuracy:
import os, time, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")
QUESTIONS = [
("A shop sells 3 abayas at 180, 240 and 310 AED. With 15% off the two cheapest and "
"free shipping over 500 AED, what is the total for one of each?", "..."), # add expected answers
# ... 10-20 more, mixing easy and hard
]
for effort in ["low", "high"]:
tokens = secs = 0
for q, _expected in QUESTIONS:
t0 = time.perf_counter()
r = client.messages.create(model=MODEL, max_tokens=16000,
thinking={"type": "adaptive"},
output_config={"effort": effort},
messages=[{"role": "user", "content": q}])
secs += time.perf_counter() - t0
tokens += r.usage.output_tokens # includes thinking tokens
print(effort, f"avg output tokens={tokens / len(QUESTIONS):.0f}", f"avg secs={secs / len(QUESTIONS):.1f}")Add an accuracy check against expected answers, and you have a small test-time-compute curve for your own workload.
Test-time compute beyond one model call
Self-consistency (sample several answers and take the majority), best-of-n with a verifier, and tool-assisted checking (running code or tests) are all ways to buy accuracy with compute. They multiply cost, so reserve them for high-value, verifiable tasks such as financial calculations, code changes and compliance checks.
Going further
Test-time compute also includes techniques outside a single model's reasoning: generating several candidate answers and selecting by majority vote (self-consistency), using a verifier model or tests to pick the best candidate, or searching over solution paths. These can be combined with reasoning models for high-stakes tasks, at multiplied cost. Evaluate whether the accuracy gain justifies the spend for each use case.
Key takeaways
- Test-time compute means spending more computation while answering; reasoning models are trained to reason before answering.
- Reasoning improves accuracy on maths, code, planning, rule-heavy analysis and agentic tasks.
- It costs latency and tokens; controls like effort levels let you match compute to difficulty.
- Reasoning models still err; verify results and test on your own tasks.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
List five tasks you use AI for. Classify each as 'needs reasoning', 'maybe' or 'no', and test one 'maybe' task at two effort levels.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.