---
title: "Working with reasoning models in practice"
description: "Prompting reasoning models Reasoning models change some prompting habits (our Advanced Prompt Engineering course covers this in depth). The essentials: -…"
url: https://optimizeall.com/learn/multimodal-and-reasoning-models/working-with-reasoning-models
updated: 2026-10-05
---

Multimodal & Reasoning Models in Practice · Reasoning models and test-time compute · lesson 12 of 17 · 11 min

# Working with reasoning models in practice

## Prompting reasoning models

Reasoning models change some prompting habits (our Advanced Prompt Engineering course covers this in depth). The essentials:

- **State the goal, constraints and definition of done clearly.** This matters more than prescribing steps.
- **Prefer high-level guidance** ("consider edge cases and check each constraint") over rigid step-by-step scripts; provider guidance for recent models suggests their own reasoning often outperforms hand-written procedures.
- **Provide all needed information up front.** Reasoning cannot recover facts that are not present; it can only reason over what it has (or fetch with tools).
- **Ask for a concise final answer** in a defined format; you do not need to request the reasoning in the output.
- **Remove legacy instructions** written to compensate for older models (for example repeated "double-check everything" demands), then re-test.

## Setting effort sensibly

A practical approach to effort or budget settings:

1. Build a small evaluation set of real tasks at varying difficulty.
2. Run at low, medium and high settings.
3. Record accuracy, median latency, slow-tail latency and cost.
4. Choose the lowest setting that meets your accuracy bar, per task type.

Often the result is a **routing** design: a fast model or low effort for easy, high-volume tasks, and higher effort for a small share of hard or high-stakes ones.

## Combining reasoning with tools

Reasoning models are especially strong in agentic setups where they decide which tools to call and interpret results. Some systems allow reasoning between tool calls (interleaved), which helps the model reflect on each result before the next step. Design tools so results are clear and verifiable, and still enforce permissions and approvals in code.

## Verifying reasoning outputs

Longer reasoning does not guarantee correctness. Keep verification proportional to stakes:

- **Code checks** for constraints (schedules, totals, rules).
- **Tests** for generated code.
- **Citations** for factual claims.
- **Human review** for decisions with legal, financial or safety impact.

## Worked example: pricing analysis

An analyst asks: "Given these three pricing tiers and last quarter's customer usage data, what happens to revenue and churn risk if we raise the middle tier by a set percentage?"

- With a fast model: a fluent answer with an arithmetic slip in the revenue projection.
- With a reasoning model at medium effort: correct arithmetic and a clearer separation of assumptions, taking longer.
- Best practice: the analyst asks the model to write the calculation as a spreadsheet formula or a short Python script, runs it on the actual data, and uses the model's reasoning for the qualitative churn discussion, labelled as assumptions.

```python
# Model-proposed, human-run projection (illustrative)
new_rev = sum(c.users * price_new[c.tier] for c in customers
              if not (c.tier == "mid" and c.at_risk))
```

The reasoning model is valuable, but the numbers come from executed code on real data.

## Choosing between reasoning and non-reasoning models

| Situation | Typical choice |
|---|---|
| Real-time chat, simple Q&A | Fast model or low effort |
| Classification at high volume | Fast model; escalate uncertain cases |
| Complex analysis, planning, code | Reasoning model, medium or high effort |
| Agent with many tools | Reasoning-capable model, with limits |
| High-stakes decision support | Reasoning model plus verification and human review |

These are starting points, not rules. Your evaluation decides.

## Communicating to stakeholders

Stakeholders often assume "the thinking model" is always better. Explain the trade-off in business terms: better accuracy on hard problems, higher cost and slower responses, and no guarantee of correctness. Present the evaluation table (accuracy, latency, cost per task) when proposing a model choice.

## Prompt patterns that suit reasoning models

```text
<goal>
Recommend whether to raise the middle pricing tier by 10%, for the
leadership meeting on Thursday.
</goal>

<context>
Pricing tiers, last quarter's usage by customer, churn history by tier
(attached). Our churn tolerance is 3% per quarter.
</context>

<constraints>
- Do all arithmetic in the Python tool; do not estimate.
- Separate facts from assumptions, and label each assumption.
- If the data cannot support a conclusion, say so.
</constraints>

<output>
A one-paragraph recommendation, a table of projected revenue and churn
under three scenarios, and the list of assumptions.
</output>
```

Notice what is absent: no "think step by step", no prescribed sequence. The goal, context, constraints and output contract do the work, and the model plans its own route.

## Interleaved thinking with tools

Several current models can reason **between** tool calls: look at a result, think about it, then decide the next call. This is especially valuable in agents and data analysis. Two practical implications: pass the model's content blocks back unchanged in multi-turn tool use (some providers require returning reasoning blocks exactly as received), and give tools clear, compact outputs so the reasoning has something clean to work with.

## Showing progress to users

Long reasoning can look like a frozen screen. Options depending on provider: stream the final answer as soon as it starts, display reasoning summaries where the API provides them, or show your own progress indicators based on tool calls. Never present raw or summarised reasoning as a guaranteed explanation of the decision.

## Measuring whether reasoning is paying off

For each route, track accuracy (or rubric score), reasoning tokens per request, latency at the median and slowest few percent, and cost per successful outcome. Review monthly: models improve, and a route that needed high effort last quarter may be fine at low effort today.

## Going further

Monitor reasoning token usage in production. Unexpected spikes can indicate prompts that trigger unnecessary deliberation (for example ambiguous or contradictory instructions), which you can often fix by clarifying the prompt rather than lowering effort.

## Video lecture: Working with reasoning models in practice

Lecture coming soon · 11 chapters · about 8 minutes. Read the full transcript below.

1. Working with reasoning models
2. Prompting principles
3. A reasoning-friendly prompt
4. Setting effort
5. Reasoning with tools
6. Verify in proportion
7. Progress and monitoring
8. Communicating trade-offs
9. Example 1: a three-day trip plan
10. Example 2: supplier bid comparison (illustrative)
11. Recap

## Lecture transcript

### Working with reasoning models

An analyst asks two models the same question: what happens to revenue and churn if we raise our middle pricing tier? The fast model writes a fluent answer with an arithmetic slip. The reasoning model gets the arithmetic right and separates assumptions more clearly, but takes longer. The best answer came from neither alone. In this lecture you will learn how to prompt reasoning models, how to set effort sensibly, how reasoning works with tools, how to verify outputs, how to show progress to users, and how to explain the trade-offs to stakeholders.

### Prompting principles

Prompting reasoning models is mostly about clarity, not choreography. State the goal, constraints and definition of done clearly. Prefer high-level guidance, like consider edge cases and check each constraint, over rigid step-by-step scripts; provider guidance for recent models suggests their own reasoning often beats hand-written procedures. Provide all the information up front, because reasoning cannot recover facts that are not present, though it can fetch them with tools. Ask for a concise final answer in a defined format. And remove legacy instructions, like repeated double-check everything, then re-test.

### A reasoning-friendly prompt

Here is a pattern from the lesson. A goal tag: recommend whether to raise the middle pricing tier by ten percent, for Thursday's leadership meeting. A context tag: the pricing tiers, last quarter's usage by customer and churn history, plus a churn tolerance of three percent per quarter. A constraints tag: do all arithmetic in the Python tool, do not estimate; separate facts from assumptions and label each assumption; and if the data cannot support a conclusion, say so. An output tag: a one-paragraph recommendation, a table of three scenarios, and the list of assumptions. Notice what is absent: no think step by step, no prescribed sequence.

### Setting effort

Set effort with a simple routine. Build a small evaluation set of real tasks at varying difficulty. Run it at low, medium and high settings. Record accuracy, median latency, slowest latency and cost. Then choose the lowest setting that meets your accuracy bar for each task type. The usual result is a routing design: low effort or a fast model for easy, high-volume work, and higher effort for a small share of hard or high-stakes cases.

### Reasoning with tools

Reasoning models are especially strong in agentic setups, deciding which tools to call and interpreting the results. Several current models can reason between tool calls, which is called interleaved thinking: look at a result, think, then choose the next call. Two practical implications. Pass the model's content blocks back unchanged in multi-turn tool use, because some providers require reasoning blocks to be returned exactly as received. And give tools clear, compact outputs, so the reasoning has something clean to work with. Permissions and approvals still belong in code.

### Verify in proportion

Longer reasoning does not guarantee correctness, so keep verification proportional to stakes. Use code checks for constraints like schedules, totals and rules. Use tests for generated code. Require citations for factual claims. And keep human review for decisions with legal, financial or safety impact. Back to the analyst: the best practice was to ask the reasoning model to write the projection as a spreadsheet formula or a short Python script, run it on the actual data, and use the model's reasoning for the qualitative churn discussion, labelled as assumptions. The numbers come from executed code on real data.

### Progress and monitoring

Long reasoning can look like a frozen screen. Depending on the provider, you can stream the final answer as soon as it starts, display reasoning summaries where the API provides them, or show your own progress indicators based on tool calls. Just never present raw or summarised reasoning as a guaranteed explanation of the decision. And monitor reasoning tokens in production. Unexpected spikes often point to ambiguous or contradictory instructions, which you can fix by clarifying the prompt rather than simply lowering effort.

### Communicating trade-offs

Stakeholders often assume the thinking model is always better. Explain the trade-off in business terms: better accuracy on hard problems, higher cost and slower responses, and no guarantee of correctness. Show the evaluation table, accuracy, latency and cost per task, when proposing a model choice. As a starting point: real-time chat and simple questions suit a fast model or low effort; high-volume classification suits a fast model with escalation; complex analysis, planning and code suit reasoning at medium or high effort; and high-stakes decision support needs reasoning plus verification and human review. Then review monthly, because newer models often reach the same accuracy at lower effort.

### Example 1: a three-day trip plan

A simple worked example. You ask a reasoning model: plan a three-day trip to Istanbul for two people, budget one thousand euros including flights from London, and my partner does not like early mornings. Version one of your prompt adds: first list the attractions, then sort by district, then calculate costs, then write the plan. The result is rigid and over the budget. Version two states only the goal, the constraints and the output: a day-by-day plan, a cost table that must stay under budget with assumptions labelled, and nothing before ten in the morning. The model plans its own route, fits the budget, and lists its price assumptions so you can check them. Goals and constraints beat choreography.

### Example 2: supplier bid comparison (illustrative)

Now a business scenario, with illustrative numbers. A procurement analyst at a manufacturing company in Sialkot compares bids from six suppliers across price, delivery time, quality history and payment terms, about thirty bids a quarter. She uses a reasoning model with a prompt that states the goal, the weighting of each criterion, and a constraint: all scoring arithmetic must be done in the code tool. The model writes and runs the scoring code, then writes a recommendation that separates facts from assumptions, such as assuming last year's quality record predicts this year's. She reviews the code, checks two scores by hand, and presents the table plus the reasoning to the purchasing committee. Preparation time falls from about a day to about two hours per round, and the committee now sees exactly how scores were calculated. Illustrative figures.

### Recap

To recap. Give reasoning models clear goals, constraints, complete information and an output contract, and remove legacy scripts. Choose effort by evaluation and route by difficulty. Use interleaved thinking with clean tools, return blocks unchanged, and keep approvals in code. Verify in proportion to stakes, with numbers from executed code. Try this now: pick one analytical task, run it with a fast model and a reasoning model, have both output their calculations as code, execute it, and compare correctness, time and cost. Next: extended and adaptive thinking controls across APIs.

## Key takeaways

- Give reasoning models clear goals, constraints and complete information; prefer high-level guidance to rigid scripts.
- Choose effort per task type via evaluation of accuracy, latency and cost; route easy tasks to cheaper settings.
- Reasoning models excel in tool-using agents but still need permissions, approvals and verification.
- Have numbers computed by executed code on real data; use reasoning for analysis and assumptions.

## Try it

Pick one analytical task. Run it with a fast model and a reasoning model, have both output any calculations as code, execute it, and compare correctness, time and cost.

- [Previous: Reasoning models and test-time compute](https://optimizeall.com/learn/multimodal-and-reasoning-models/test-time-compute-concepts)
- [Next: Extended and adaptive thinking: controls across APIs](https://optimizeall.com/learn/multimodal-and-reasoning-models/thinking-controls-across-apis)
- [All lessons of Multimodal & Reasoning Models in Practice](https://optimizeall.com/learn/multimodal-and-reasoning-models)
