---
title: "Cost and latency trade-offs — Advanced Prompt Engineering"
description: "The three-way trade-off Every production AI feature balances quality, cost and latency . Advanced prompt engineering includes knowing which lever to…"
url: https://optimizeall.com/learn/advanced-prompt-engineering/cost-latency-tradeoffs
updated: 2026-10-05
---

Advanced Prompt Engineering · Prompt operations: versioning, cost and latency · lesson 16 of 17 · 12 min

# Cost and latency trade-offs

## The three-way trade-off

Every production AI feature balances **quality, cost and latency**. Advanced prompt engineering includes knowing which lever to pull. Model prices, speeds and limits change frequently and differ between providers, so this lesson focuses on durable mechanics rather than numbers; always check current pricing pages.

## Where cost comes from

Most APIs charge per token, typically with different rates for input and output tokens, and output tokens usually cost more. Reasoning or "thinking" tokens are generally billed as output. So cost is driven by:

- **Input size:** system prompt, examples, retrieved documents, conversation history, tool definitions.
- **Output size:** answer length plus any reasoning tokens.
- **Number of calls:** chains, retries, verification loops, agent steps.
- **Model tier:** larger, more capable models cost more per token.

A quick back-of-envelope model (illustrative, with made-up rates):

```text
cost per request = input_tokens x input_rate + output_tokens x output_rate
monthly cost     = cost per request x requests per day x 30

Example (hypothetical rates): 3,000 input tokens, 500 output tokens,
20,000 requests/day. If output costs 4x input, output is ~40% of spend
even though it is only ~14% of tokens.
```

## Where latency comes from

- **Time to first token:** affected by input length, model size and provider load.
- **Generation time:** roughly proportional to output tokens, including reasoning.
- **Sequential calls:** a five-step chain waits for five round trips.
- **Tool calls:** each external lookup adds its own latency.

Users perceive latency differently depending on the interface: streaming the answer token by token makes long responses feel faster, while background jobs can tolerate much more.

## The levers, from least to most effort

1. **Shorten outputs.** Ask for the length you need. Remove preambles ("Certainly! Here is..."), which cost tokens and time.
2. **Trim inputs.** Remove redundant instructions, examples that do not earn their place, and stale conversation history.
3. **Prompt caching.** Many providers offer caching of a repeated prompt prefix (for example a long system prompt or reference document), which can substantially cut cost and time-to-first-token for repeated requests. Structure prompts so the stable parts come first and variable parts last to maximise cache hits. Check your provider's rules for how caching is enabled and priced.
4. **Right-size the model.** Use smaller, faster models for simple steps (classification, extraction, routing) and reserve the largest models for hard reasoning.
5. **Tune reasoning effort.** Lower effort for easy cases; higher only where evals show it pays.
6. **Parallelise.** Run independent chain steps or map-reduce chunks concurrently.
7. **Batch processing.** For non-urgent work (overnight enrichment, bulk classification), batch APIs are often cheaper, in exchange for slower turnaround.
8. **Route by difficulty.** A cheap first pass handles easy cases; uncertain ones escalate to a stronger model or a human.

## Worked example: routing

An e-commerce brand classifies 50,000 product reviews a week into themes. Their first version uses the most capable model for everything. The revised design:

1. A small, fast model classifies each review and returns a confidence field.
2. Reviews with low confidence or safety-related themes go to the larger model.
3. The weekly job runs through a batch endpoint overnight.
4. The long taxonomy and examples sit in a cached prefix.

Evaluation shows quality within their tolerance on the eval set, with a large reduction in cost. The key was not a clever prompt, but measuring where the expensive model actually mattered.

## Don't optimise blind

Every cost cut is a potential quality cut. For each change, rerun the eval set and compare quality, cost and latency together, ideally in one table. Watch the slow tail (the slowest few percent of requests) as well as the average, because users remember the worst waits.

## Failure modes

- **Penny-wise prompts.** Removing the examples that prevented expensive human escalations.
- **Hidden retries.** Validation failures silently triple calls. Track retries as a metric.
- **Cache busting.** Putting a timestamp or user name at the top of the prompt, which invalidates the cached prefix on every request.

## Hands-on: measure cost per request from real usage data

Every major API returns token usage with each response. Log it, and compute cost with rates you keep in config (never hard-code prices in logic; they change).

```python
import os, time, anthropic
client = anthropic.Anthropic()

RATES = {  # USD per million tokens: fill from the provider's current pricing page
    "input": float(os.environ["RATE_INPUT"]),
    "output": float(os.environ["RATE_OUTPUT"]),
    "cache_read": float(os.environ["RATE_CACHE_READ"]),
    "cache_write": float(os.environ["RATE_CACHE_WRITE"]),
}

def call_and_cost(**kwargs):
    t0 = time.perf_counter()
    resp = client.messages.create(**kwargs)
    u = resp.usage
    cost = (u.input_tokens * RATES["input"] + u.output_tokens * RATES["output"]
            + (u.cache_read_input_tokens or 0) * RATES["cache_read"]
            + (u.cache_creation_input_tokens or 0) * RATES["cache_write"]) / 1e6
    return resp, {"latency_s": round(time.perf_counter() - t0, 2), "cost_usd": round(cost, 6),
                  "in": u.input_tokens, "out": u.output_tokens,
                  "cache_read": u.cache_read_input_tokens}
```

Output tokens include thinking tokens on reasoning models, which is why effort settings show up directly in your bill.

## Newer levers worth knowing

- **Effort settings** on reasoning models trade depth for tokens and latency; lower effort often suffices for routine routes.
- **Batch APIs** from the major providers process non-urgent requests asynchronously at a discount.
- **Faster serving tiers or modes** exist on some platforms for latency-critical routes, usually at a premium price; check availability for your model.
- **Streaming** does not reduce cost but greatly improves perceived latency; always stream user-facing long answers.
- **Prompt caching** is covered in depth in the next lesson.

## Measure cost per outcome

Join your request logs with business outcomes (ticket resolved, lead qualified, document approved). Divide total AI spend, plus human review time, by successful outcomes. Decisions based on this number routinely differ from decisions based on price per token.

## Going further

Build a cost dashboard per feature: tokens in and out, calls per user action, cache hit rate, retry rate and cost per successful outcome (for example cost per resolved ticket). Cost per outcome is the number your business actually cares about, and it often reveals that a more expensive model is cheaper overall because it resolves more on the first attempt.

## Video lecture: Cost and latency trade-offs

Lecture coming soon · 12 chapters · about 7 minutes. Read the full transcript below.

1. Cost and latency trade-offs
2. Why balance all three?
3. Where cost comes from
4. Illustrative maths
5. Where latency comes from
6. The levers
7. Worked example: review classification
8. Example 1: removing preambles (illustrative)
9. Example 2: a travel assistant (illustrative)
10. Measure and avoid traps
11. Cost per outcome
12. Recap

## Lecture transcript

### Cost and latency trade-offs

An e-commerce brand was classifying fifty thousand product reviews a week with its most capable, most expensive model. The results were great. The bill was not. After a redesign, quality stayed within tolerance and costs fell dramatically. The trick was not a clever prompt. It was measuring where the expensive model actually mattered. In this lecture you will learn where cost and latency come from, the levers that reduce them, how to log real cost per request, and why cost per outcome is the number that really matters.

### Why balance all three?

Why does this matter? Because AI features that users love can still be cancelled when the bill arrives, and features that are cheap but slow get abandoned by users. The teams that succeed treat quality, cost and latency as one design problem, and they measure all three every time they change something. Think of it like running a restaurant kitchen: the dish must taste great, arrive quickly and still make a profit. Optimising only one of those closes the restaurant.

### Where cost comes from

Every production AI feature balances quality, cost and latency. Prices, speeds and limits change often, so learn the mechanics and check current pricing pages. Most APIs charge per token, with different rates for input and output, and output usually costs more. Reasoning or thinking tokens are generally billed as output. So cost comes from input size, meaning system prompt, examples, documents, history and tool definitions; output size, including reasoning; the number of calls in chains, retries, verification and agent steps; and the model tier.

### Illustrative maths

Here is an illustrative calculation with made-up rates. Say each request has three thousand input tokens and five hundred output tokens, at twenty thousand requests a day. If output costs four times as much as input per token, output is about forty percent of your spend, even though it is only about fourteen percent of your tokens. That is why shortening outputs and removing preambles like certainly, here is, often pays off immediately.

### Where latency comes from

Latency has its own sources. Time to first token depends on input length, model size and load. Generation time grows with output tokens, including reasoning. Sequential calls add round trips; a five-step chain waits five times. And tool calls add their own latency. Users perceive latency differently by interface: streaming makes long answers feel faster, while background jobs can tolerate far more. Always stream user-facing long answers.

### The levers

Now the levers, from least to most effort. Shorten outputs and remove preambles. Trim inputs: redundant instructions, examples that do not earn their place, stale history. Use prompt caching for repeated prefixes, which the next lesson covers in depth. Right-size the model: small, fast models for classification, extraction and routing; large ones for hard reasoning. Tune reasoning effort, because lower effort often suffices for routine routes. Parallelise independent steps. Use batch APIs for non-urgent work; the major providers process batches asynchronously at a discount. And route by difficulty, escalating only uncertain cases.

### Worked example: review classification

Back to the review classifier. The redesign: a small, fast model classifies each review and returns a confidence field. Low-confidence reviews and anything safety-related go to the larger model. The weekly job runs through a batch endpoint overnight. And the long taxonomy and examples sit in a cached prefix. Their evaluation shows quality within tolerance, with a large reduction in cost. They measured before and after on the same eval set, together, in one table.

### Example 1: removing preambles (illustrative)

A simple worked example, with illustrative numbers. Your assistant's answers average six hundred output tokens, and about one hundred of those are preamble and closing pleasantries, like certainly, here is your answer, and I hope this helps. Remove them with one instruction: start directly with the answer; no preamble or closing lines. That is roughly one sixth of your output tokens gone, and because output tokens usually cost more than input, it can be a meaningful share of your bill. Users also see the real answer sooner. Small instruction, measurable saving. Measure it on your eval set to confirm quality holds.

### Example 2: a travel assistant (illustrative)

Now a business scenario, with illustrative numbers. A travel agency in Dubai runs an AI assistant with about thirty thousand requests a day. Each request carries a ten-thousand-token system prompt and destination guide, and uses the largest model at high effort. Median response time is nine seconds, and the monthly bill is growing fast. The team measures before changing anything. Then they cache the stable prompt and guide. They route simple requests, like visa document checklists and opening hours, to a smaller model at low effort, about seventy percent of traffic. They keep the large model for complex itineraries. They stream all user-facing answers. Median response time drops to about three seconds, cost per request falls by roughly two thirds, and quality on their eval set stays within tolerance. The deciding metric was cost per completed booking enquiry, not cost per token. Illustrative figures.

### Measure and avoid traps

Measure real cost from real usage. Every major API returns token usage with each response. The lesson's code logs input, output, cache read and cache write tokens per request, multiplies by rates you keep in configuration, never hard-coded in logic, and records latency. Then watch four failure modes. Penny-wise prompts that remove the examples preventing expensive human escalations. Hidden retries, where validation failures silently triple calls. Cache busting, like a timestamp at the top of the prompt. And optimising averages while ignoring the slow tail that users remember.

### Cost per outcome

The most important metric is cost per successful outcome. Join your request logs with business outcomes, a resolved ticket, a qualified lead, an approved document, and divide total AI spend plus human review time by the number of successes. A cheaper model that cuts cost per request in half but increases escalations to humans may be more expensive overall. Decisions based on cost per outcome routinely differ from decisions based on price per token.

### Recap

To recap. Cost is driven by input, output including thinking, number of calls and model tier; latency by first token, generation, sequential steps and tools. Pull the levers, shorter outputs, trimmed inputs, caching, right-sized models, effort tuning, parallelism, batching and routing, but measure each against your eval set. Try this now: estimate tokens in and out, calls per user action and daily requests for one feature, identify the two biggest cost drivers, and test one lever for each. Next: prompt caching in practice.

## Key takeaways

- Balance quality, cost and latency; prices change, so learn the mechanics and check current pricing.
- Cost is driven by input size, output (including reasoning) tokens, number of calls and model tier.
- Levers: shorter outputs, trimmed inputs, prompt caching, right-sized models, effort tuning, parallelism, batching and routing.
- Measure every optimisation against the eval set and track cost per successful outcome.

## Try it

For one AI feature, estimate tokens in and out per request, calls per user action and requests per day. Identify the two largest cost drivers and one lever for each, then test its quality impact.

- [Previous: Versioning and managing prompts in production](https://optimizeall.com/learn/advanced-prompt-engineering/versioning-prompts)
- [Next: Prompt caching in practice](https://optimizeall.com/learn/advanced-prompt-engineering/prompt-caching-in-practice)
- [All lessons of Advanced Prompt Engineering](https://optimizeall.com/learn/advanced-prompt-engineering)
