Multimodal & Reasoning Models in PracticeReasoning models and test-time compute · Lesson 12 of 17

Working with reasoning models in practice

Article · 11 min · 8 min lecture

Video lecture

Working with reasoning models in practice

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Working with reasoning models

  • Prompting patterns
  • Setting effort
  • Reasoning with tools
  • Verification and communication

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Prompting reasoning models

Reasoning models change some prompting habits (our Advanced Prompt Engineering course covers this in depth). The essentials:

  • State the goal, constraints and definition of done clearly. This matters more than prescribing steps.
  • Prefer high-level guidance ("consider edge cases and check each constraint") over rigid step-by-step scripts; provider guidance for recent models suggests their own reasoning often outperforms hand-written procedures.
  • Provide all needed information up front. Reasoning cannot recover facts that are not present; it can only reason over what it has (or fetch with tools).
  • Ask for a concise final answer in a defined format; you do not need to request the reasoning in the output.
  • Remove legacy instructions written to compensate for older models (for example repeated "double-check everything" demands), then re-test.

Setting effort sensibly

A practical approach to effort or budget settings:

  1. Build a small evaluation set of real tasks at varying difficulty.
  2. Run at low, medium and high settings.
  3. Record accuracy, median latency, slow-tail latency and cost.
  4. Choose the lowest setting that meets your accuracy bar, per task type.

Often the result is a routing design: a fast model or low effort for easy, high-volume tasks, and higher effort for a small share of hard or high-stakes ones.

Combining reasoning with tools

Reasoning models are especially strong in agentic setups where they decide which tools to call and interpret results. Some systems allow reasoning between tool calls (interleaved), which helps the model reflect on each result before the next step. Design tools so results are clear and verifiable, and still enforce permissions and approvals in code.

Verifying reasoning outputs

Longer reasoning does not guarantee correctness. Keep verification proportional to stakes:

  • Code checks for constraints (schedules, totals, rules).
  • Tests for generated code.
  • Citations for factual claims.
  • Human review for decisions with legal, financial or safety impact.

Worked example: pricing analysis

An analyst asks: "Given these three pricing tiers and last quarter's customer usage data, what happens to revenue and churn risk if we raise the middle tier by a set percentage?"

  • With a fast model: a fluent answer with an arithmetic slip in the revenue projection.
  • With a reasoning model at medium effort: correct arithmetic and a clearer separation of assumptions, taking longer.
  • Best practice: the analyst asks the model to write the calculation as a spreadsheet formula or a short Python script, runs it on the actual data, and uses the model's reasoning for the qualitative churn discussion, labelled as assumptions.
# Model-proposed, human-run projection (illustrative)
new_rev = sum(c.users * price_new[c.tier] for c in customers
              if not (c.tier == "mid" and c.at_risk))

The reasoning model is valuable, but the numbers come from executed code on real data.

Choosing between reasoning and non-reasoning models

SituationTypical choice
Real-time chat, simple Q&AFast model or low effort
Classification at high volumeFast model; escalate uncertain cases
Complex analysis, planning, codeReasoning model, medium or high effort
Agent with many toolsReasoning-capable model, with limits
High-stakes decision supportReasoning model plus verification and human review

These are starting points, not rules. Your evaluation decides.

Communicating to stakeholders

Stakeholders often assume "the thinking model" is always better. Explain the trade-off in business terms: better accuracy on hard problems, higher cost and slower responses, and no guarantee of correctness. Present the evaluation table (accuracy, latency, cost per task) when proposing a model choice.

Prompt patterns that suit reasoning models

<goal>
Recommend whether to raise the middle pricing tier by 10%, for the
leadership meeting on Thursday.
</goal>

<context>
Pricing tiers, last quarter's usage by customer, churn history by tier
(attached). Our churn tolerance is 3% per quarter.
</context>

<constraints>
- Do all arithmetic in the Python tool; do not estimate.
- Separate facts from assumptions, and label each assumption.
- If the data cannot support a conclusion, say so.
</constraints>

<output>
A one-paragraph recommendation, a table of projected revenue and churn
under three scenarios, and the list of assumptions.
</output>

Notice what is absent: no "think step by step", no prescribed sequence. The goal, context, constraints and output contract do the work, and the model plans its own route.

Interleaved thinking with tools

Several current models can reason between tool calls: look at a result, think about it, then decide the next call. This is especially valuable in agents and data analysis. Two practical implications: pass the model's content blocks back unchanged in multi-turn tool use (some providers require returning reasoning blocks exactly as received), and give tools clear, compact outputs so the reasoning has something clean to work with.

Showing progress to users

Long reasoning can look like a frozen screen. Options depending on provider: stream the final answer as soon as it starts, display reasoning summaries where the API provides them, or show your own progress indicators based on tool calls. Never present raw or summarised reasoning as a guaranteed explanation of the decision.

Measuring whether reasoning is paying off

For each route, track accuracy (or rubric score), reasoning tokens per request, latency at the median and slowest few percent, and cost per successful outcome. Review monthly: models improve, and a route that needed high effort last quarter may be fine at low effort today.

Going further

Monitor reasoning token usage in production. Unexpected spikes can indicate prompts that trigger unnecessary deliberation (for example ambiguous or contradictory instructions), which you can often fix by clarifying the prompt rather than lowering effort.

Key takeaways

  • Give reasoning models clear goals, constraints and complete information; prefer high-level guidance to rigid scripts.
  • Choose effort per task type via evaluation of accuracy, latency and cost; route easy tasks to cheaper settings.
  • Reasoning models excel in tool-using agents but still need permissions, approvals and verification.
  • Have numbers computed by executed code on real data; use reasoning for analysis and assumptions.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the best way to choose a reasoning effort setting?
  2. A reasoning model produces a revenue projection. What is the most reliable way to get correct numbers?
  3. Reasoning token usage spikes on a support workflow. What is a likely cause worth checking first?

Put it into practice

Pick one analytical task. Run it with a fast model and a reasoning model, have both output any calculations as code, execute it, and compare correctness, time and cost.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.