---
title: "Cost and latency engineering for agents"
description: "Where agent money goes An agent's bill is dominated by input tokens resent on every step (system prompt, tool definitions, growing history), then output…"
url: https://optimizeall.com/learn/ai-agents-engineering/cost-and-latency-engineering
updated: 2026-10-05
---

Building Production AI Agents · Production engineering: cost, deployment and frameworks · lesson 14 of 18 · 16 min

# Cost and latency engineering for agents

## Where agent money goes

An agent's bill is dominated by **input tokens resent on every step** (system prompt, tool definitions, growing history), then output tokens (including reasoning), then tool costs (search APIs, compute). Latency is dominated by the number of sequential model calls, output length and slow tools. So the levers are: fewer steps, smaller and more stable context, the right model and effort per step, caching, parallelism and batching.

## Lever 1: prompt caching

Providers cache a **prefix** of your request so repeated content is billed and processed more cheaply.

- **Anthropic**: explicit `cache_control` breakpoints (or top-level automatic caching). Cache reads are billed at roughly a tenth of the base input price; writes carry a premium (about 1.25× for the default 5-minute TTL, 2× for the 1-hour TTL). Check current pricing pages.
- **OpenAI**: automatic caching of repeated prefixes above a minimum length, with discounted cached-input pricing.
- **Google Gemini**: implicit caching on supported models plus explicit cached-content objects you create and reference.

Caching is a **prefix match**: any change early in the request invalidates everything after it. Put stable content first (system prompt, tool list in deterministic order), volatile content last. Never put timestamps or request IDs in the system prompt. Verify with the usage fields (for Claude, `cache_read_input_tokens`); if reads stay at zero, something is changing the prefix.

## Lever 2: fewer, cheaper steps

- Better tools (lesson 4) cut calls per task more than any model tweak.
- Consolidate chatty tools; return concise outputs; paginate.
- Clear stale tool results and compact long histories (lesson 7).
- Put hard caps on steps and tokens; escalate rather than loop.

## Lever 3: right model and effort per step

Use small, fast models for routing, extraction and sub-agent reading; larger models for planning and synthesis; and tune reasoning effort (lesson 5). Before building an elaborate cascade, measure a simpler alternative: the strongest model at lower effort may match a cascade while keeping one cache namespace (caches are model-specific). Always compare **cost per completed task**.

## Lever 4: parallelism

Models can request several tools in one turn; execute them concurrently. Fan out independent sub-tasks with a concurrency limit. Parallelism cuts wall-clock time but not tokens, and it multiplies rate-limit usage.

## Lever 5: batch and background

Non-urgent work (nightly enrichment, bulk classification, eval runs) can use **batch APIs**, which process asynchronously at a discount (Anthropic's Message Batches and OpenAI's Batch API are both priced at 50% of standard rates at the time of writing; Gemini also offers batch mode). Results arrive within hours, not seconds, and arrive in any order: key them by your own IDs.

## Lever 6: streaming and perceived latency

Stream tokens and progress events ("Searching CRM… Found 3 contacts… Drafting") so users see progress. Perceived latency often matters more than total time for interactive agents.

## Worked example: cutting a UK agency's reporting-agent bill

Baseline: a monthly client-report agent averaged 14 steps, a 9k-token system prompt plus tool list, full raw analytics JSON in tool results, and one large model at high effort for everything.

Changes and illustrative effect:

1. Cached the stable prefix (system + tools): input cost on steps 2–14 dropped sharply.
2. Analytics tool returns a 40-row summary instead of raw JSON with 2,000 rows: steps fell to 8 as the model stopped paging.
3. Data-pull steps moved to a small model at low effort; synthesis stayed on the large model.
4. The 60 monthly reports moved to the batch API overnight.

Result (illustrative, measured on their eval set): cost per completed report fell by well over half with no drop in reviewer acceptance. The biggest single win was the tool output change, not the model change.

## Hands-on: a cost meter with caching

```python
import os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5")
# Example prices per million tokens - load real values from config; check the pricing page.
PRICE = {"in": float(os.environ.get("PRICE_IN", "0")), "out": float(os.environ.get("PRICE_OUT", "0")),
         "cache_read_mult": 0.1, "cache_write_mult": 1.25}

SYSTEM = [{"type": "text", "text": open("system_prompt.md").read(),
           "cache_control": {"type": "ephemeral"}}]          # stable prefix, cached

def step(messages, tools):
    resp = client.messages.create(model=MODEL, max_tokens=4000, system=SYSTEM,
                                  tools=tools, messages=messages)
    u = resp.usage
    cost = ((u.input_tokens * PRICE["in"])
            + ((u.cache_read_input_tokens or 0) * PRICE["in"] * PRICE["cache_read_mult"])
            + ((u.cache_creation_input_tokens or 0) * PRICE["in"] * PRICE["cache_write_mult"])
            + (u.output_tokens * PRICE["out"])) / 1_000_000
    print(f"in={u.input_tokens} cache_read={u.cache_read_input_tokens} "
          f"cache_write={u.cache_creation_input_tokens} out={u.output_tokens} cost=${cost:.4f}")
    return resp, cost
```

Keep the tools list identical and in the same order on every step, or the cache will miss. Aggregate `cost` per run, per tenant and per feature in your traces (lesson 13).

## Latency checklist

- Stream everything interactive.
- Run independent tool calls concurrently.
- Cache prefixes (cached tokens are also faster to process).
- Keep outputs short; ask for structured, compact responses.
- Put slow tools behind caches or async jobs.
- Consider region and platform choice for network latency.

## Pitfalls

- Optimizing price per token while calls per task balloon.
- Breaking caching with dynamic system prompts.
- Parallel fan-out that trips rate limits and triggers retries (net slower).
- Using batch for anything a user is waiting on.

## Measuring success

Cost per completed task (p50/p95), cache hit ratio, steps per task, time to first token, total run time, and quality on your eval set. Every optimization must hold quality constant.

## Video lecture: Cost and latency engineering for agents

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Cost and latency engineering
2. Why it matters
3. Anatomy of agent cost
4. Lever 1: prompt caching
5. Levers 2 and 3
6. Simple example (illustrative)
7. Levers 4–6
8. Example: agency reporting agent
9. Hands-on: cost meter
10. Traps and metrics
11. Deeper: the reporting agent (illustrative)
12. Watch me do it: the cost meter
13. Try this now
14. Recap

## Lecture transcript

### Cost and latency engineering

Agents can get expensive fast, and slow too. The good news is that most of the cost comes from a few predictable places, and there are six levers that fix them. In this lesson you'll learn where agent money and time actually go, and how to cut both without hurting quality.

### Why it matters

Why does cost engineering matter? Because a proof of concept that costs a few cents per run can become a serious monthly bill at scale, and slow agents lose users. Think of a delivery business. The cost isn't only fuel per kilometer. It's the number of trips, the route, whether trucks leave half empty, and whether urgent parcels go by courier while bulk goes by freight. Agent costs work the same way: steps are trips, context is cargo, caching is a warehouse near the customer, and batch is freight.

### Anatomy of agent cost

First, the anatomy of the bill. On every step, you resend the system prompt, the tool definitions and the growing history. Those input tokens dominate. Then come output tokens, including reasoning, then tool costs like search APIs and compute. Latency is driven by how many sequential model calls you make, how long the outputs are, and how slow your tools are. So the levers are fewer steps, smaller and more stable context, the right model and effort for each step, caching, parallelism and batching.

### Lever 1: prompt caching

Lever one is prompt caching. Providers can cache the start of your request so repeated content is cheaper and faster. With Claude you mark cache breakpoints, and cached reads cost roughly a tenth of the normal input price, with a small premium to write the cache. OpenAI caches long repeated prefixes automatically. Gemini has implicit caching plus explicit caches you create. The golden rule: caching is a prefix match. Put stable things first, like the system prompt and tools in a fixed order, and put changing things last. A timestamp in your system prompt silently breaks everything.

### Levers 2 and 3

Lever two is fewer, cheaper steps. Better tools cut calls per task more than any model tweak. Consolidate chatty tools, return concise outputs, clear stale results, compact long histories, and put hard caps on steps and tokens. Lever three is the right model and effort for each step. Small models for routing and reading, bigger ones for planning and synthesis. But before building a complex cascade, test the strongest model at lower effort. It may match the cascade and keep one cache, since caches are specific to each model.

### Simple example (illustrative)

A simple example with illustrative numbers. Your agent's system prompt and tools total eight thousand tokens, and a typical run takes ten steps. Without caching, you send those eight thousand tokens ten times: eighty thousand input tokens per run just for the fixed part. With prefix caching, the first step writes the cache, and the next nine steps read it at a fraction of the price. Then you trim a verbose tool so runs take seven steps instead of ten. Two small changes, and the fixed part of the bill drops dramatically.

### Levers 4–6

Lever four is parallelism. When the model asks for several tools at once, run them concurrently, and fan out independent sub tasks with a concurrency limit. That cuts waiting time, but not tokens, and it multiplies your rate limit usage. Lever five is batch. For work nobody is waiting on, like overnight enrichment or eval runs, batch APIs process requests asynchronously at a discount. At the time of writing, Anthropic and OpenAI both price batch at half the standard rate. Results arrive in any order, so key them by your own ids. Lever six is streaming: show tokens and progress events so users see something happening.

### Example: agency reporting agent

A worked example. A UK agency's monthly reporting agent averaged fourteen steps, with a nine thousand token system prompt, raw analytics JSON in tool results, and one large model at high effort for everything. They cached the stable prefix. They made the analytics tool return a forty row summary instead of two thousand raw rows, and steps fell to eight. Data pulls moved to a small model at low effort. And sixty monthly reports moved to overnight batch. On their eval set, cost per completed report fell by well over half with no drop in reviewer acceptance. The biggest single win was the tool output change.

### Hands-on: cost meter

The lesson's code adds a cost meter to each step. It marks the system prompt for caching, reads the usage fields, including cache reads and cache writes, and computes cost from prices you load from configuration. Never hard code prices from memory; check the provider's pricing page. Keep the tools list identical and in the same order on every step, or the cache will miss. Then aggregate cost per run, per tenant and per feature in your traces.

### Traps and metrics

Watch for these traps: optimizing price per token while calls per task balloon, breaking caching with a dynamic system prompt, fanning out so hard you trip rate limits and end up slower after retries, and using batch for anything a user is waiting on. Measure cost per completed task at the median and ninety fifth percentile, cache hit ratio, steps per task, time to first token, total run time, and quality on your eval set.

### Deeper: the reporting agent (illustrative)

Let's deepen the UK agency's reporting agent with illustrative numbers. Sixty monthly client reports. Before: fourteen steps each, a nine thousand token fixed prefix resent every step, raw analytics JSON, one large model at high effort, all run during office hours. After: the cached prefix, a forty row summary tool, a small model for data pulls, and an overnight batch. Their dashboard showed cost per completed report falling well below half, time to first token faster on interactive runs thanks to caching, and reviewer acceptance unchanged on their evaluation set. The finance director's favorite chart was simple: cost per report by month, with each change annotated where it happened.

### Watch me do it: the cost meter

Watch me do it. Let's walk through the cost meter. Prices come from environment variables, input and output per million tokens, with cache read and cache write multipliers. The system prompt is a list with one text block, loaded from a file, and it carries cache control, so everything up to that point can be cached. The step function calls messages create with the same system and tools every time, then reads four usage fields: normal input tokens, cache read tokens, cache creation tokens and output tokens. It computes cost by multiplying each by its price and divides by a million. I run two steps. Step one: cache write shows about nine thousand tokens, cache read zero. Step two: cache read shows about nine thousand, cache write zero, and the cost line is visibly smaller. If step two showed zero cache reads, I'd go hunting for a changing prefix.

### Try this now

Try this now. Add the cost meter from the lesson to your agent and run your eval set once as a baseline. Record cost per completed task at the median and ninety fifth percentile. Then make exactly one change: turn on prefix caching, or shrink your largest tool output, or move one step to a smaller model. Rerun the eval. Did quality hold? Did cost per completed task drop? One change at a time is how you learn which lever actually matters for your agent.

### Recap

To recap: input tokens resent every step dominate cost, so cut steps and stabilize prefixes. Use caching, the right model and effort per step, parallelism, batch and streaming, and always judge by cost per completed task at constant quality. Your next step: add the cost meter to your agent, turn on prefix caching, and compare the numbers before and after on your eval set.

## Key takeaways

- Input tokens resent every step dominate agent cost; fewer steps and stable prefixes matter most.
- Prompt caching is a prefix match: stable content first, volatile content last, verify with usage fields.
- Tool output design often beats model changes as a cost lever.
- Use batch APIs for non-urgent bulk work and streaming for perceived latency.
- Optimize cost per completed task while holding eval quality constant.

## Try it

Add the cost meter to your agent, enable prefix caching, and compare cost per completed task and cache read tokens before and after on your eval set.

- [Previous: Observability: tracing, logging and dashboards for agents](https://optimizeall.com/learn/ai-agents-engineering/observability-and-tracing)
- [Next: Deploying agents: architectures, sandboxes and safe rollout](https://optimizeall.com/learn/ai-agents-engineering/deployment-patterns)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
