Building Production AI AgentsProduction engineering: cost, deployment and frameworks · Lesson 14 of 18

Cost and latency engineering for agents

Article · 16 min · 9 min lecture

Video lecture

Cost and latency engineering for agents

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Cost and latency engineering

  • Where the money goes
  • Six levers
  • Measuring per completed task

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Where agent money goes

An agent's bill is dominated by input tokens resent on every step (system prompt, tool definitions, growing history), then output tokens (including reasoning), then tool costs (search APIs, compute). Latency is dominated by the number of sequential model calls, output length and slow tools. So the levers are: fewer steps, smaller and more stable context, the right model and effort per step, caching, parallelism and batching.

Lever 1: prompt caching

Providers cache a prefix of your request so repeated content is billed and processed more cheaply.

  • Anthropic: explicit cache_control breakpoints (or top-level automatic caching). Cache reads are billed at roughly a tenth of the base input price; writes carry a premium (about 1.25× for the default 5-minute TTL, 2× for the 1-hour TTL). Check current pricing pages.
  • OpenAI: automatic caching of repeated prefixes above a minimum length, with discounted cached-input pricing.
  • Google Gemini: implicit caching on supported models plus explicit cached-content objects you create and reference.

Caching is a prefix match: any change early in the request invalidates everything after it. Put stable content first (system prompt, tool list in deterministic order), volatile content last. Never put timestamps or request IDs in the system prompt. Verify with the usage fields (for Claude, cache_read_input_tokens); if reads stay at zero, something is changing the prefix.

Lever 2: fewer, cheaper steps

  • Better tools (lesson 4) cut calls per task more than any model tweak.
  • Consolidate chatty tools; return concise outputs; paginate.
  • Clear stale tool results and compact long histories (lesson 7).
  • Put hard caps on steps and tokens; escalate rather than loop.

Lever 3: right model and effort per step

Use small, fast models for routing, extraction and sub-agent reading; larger models for planning and synthesis; and tune reasoning effort (lesson 5). Before building an elaborate cascade, measure a simpler alternative: the strongest model at lower effort may match a cascade while keeping one cache namespace (caches are model-specific). Always compare cost per completed task.

Lever 4: parallelism

Models can request several tools in one turn; execute them concurrently. Fan out independent sub-tasks with a concurrency limit. Parallelism cuts wall-clock time but not tokens, and it multiplies rate-limit usage.

Lever 5: batch and background

Non-urgent work (nightly enrichment, bulk classification, eval runs) can use batch APIs, which process asynchronously at a discount (Anthropic's Message Batches and OpenAI's Batch API are both priced at 50% of standard rates at the time of writing; Gemini also offers batch mode). Results arrive within hours, not seconds, and arrive in any order: key them by your own IDs.

Lever 6: streaming and perceived latency

Stream tokens and progress events ("Searching CRM… Found 3 contacts… Drafting") so users see progress. Perceived latency often matters more than total time for interactive agents.

Worked example: cutting a UK agency's reporting-agent bill

Baseline: a monthly client-report agent averaged 14 steps, a 9k-token system prompt plus tool list, full raw analytics JSON in tool results, and one large model at high effort for everything.

Changes and illustrative effect:

  1. Cached the stable prefix (system + tools): input cost on steps 2–14 dropped sharply.
  2. Analytics tool returns a 40-row summary instead of raw JSON with 2,000 rows: steps fell to 8 as the model stopped paging.
  3. Data-pull steps moved to a small model at low effort; synthesis stayed on the large model.
  4. The 60 monthly reports moved to the batch API overnight.

Result (illustrative, measured on their eval set): cost per completed report fell by well over half with no drop in reviewer acceptance. The biggest single win was the tool output change, not the model change.

Hands-on: a cost meter with caching

import os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5")
# Example prices per million tokens - load real values from config; check the pricing page.
PRICE = {"in": float(os.environ.get("PRICE_IN", "0")), "out": float(os.environ.get("PRICE_OUT", "0")),
         "cache_read_mult": 0.1, "cache_write_mult": 1.25}

SYSTEM = [{"type": "text", "text": open("system_prompt.md").read(),
           "cache_control": {"type": "ephemeral"}}]          # stable prefix, cached

def step(messages, tools):
    resp = client.messages.create(model=MODEL, max_tokens=4000, system=SYSTEM,
                                  tools=tools, messages=messages)
    u = resp.usage
    cost = ((u.input_tokens * PRICE["in"])
            + ((u.cache_read_input_tokens or 0) * PRICE["in"] * PRICE["cache_read_mult"])
            + ((u.cache_creation_input_tokens or 0) * PRICE["in"] * PRICE["cache_write_mult"])
            + (u.output_tokens * PRICE["out"])) / 1_000_000
    print(f"in={u.input_tokens} cache_read={u.cache_read_input_tokens} "
          f"cache_write={u.cache_creation_input_tokens} out={u.output_tokens} cost=${cost:.4f}")
    return resp, cost

Keep the tools list identical and in the same order on every step, or the cache will miss. Aggregate cost per run, per tenant and per feature in your traces (lesson 13).

Latency checklist

  • Stream everything interactive.
  • Run independent tool calls concurrently.
  • Cache prefixes (cached tokens are also faster to process).
  • Keep outputs short; ask for structured, compact responses.
  • Put slow tools behind caches or async jobs.
  • Consider region and platform choice for network latency.

Pitfalls

  • Optimizing price per token while calls per task balloon.
  • Breaking caching with dynamic system prompts.
  • Parallel fan-out that trips rate limits and triggers retries (net slower).
  • Using batch for anything a user is waiting on.

Measuring success

Cost per completed task (p50/p95), cache hit ratio, steps per task, time to first token, total run time, and quality on your eval set. Every optimization must hold quality constant.

Key takeaways

  • Input tokens resent every step dominate agent cost; fewer steps and stable prefixes matter most.
  • Prompt caching is a prefix match: stable content first, volatile content last, verify with usage fields.
  • Tool output design often beats model changes as a cost lever.
  • Use batch APIs for non-urgent bulk work and streaming for perceived latency.
  • Optimize cost per completed task while holding eval quality constant.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. cache_read_input_tokens stays at zero across repeated agent steps. What is the most likely cause?
  2. Which workload is best suited to a batch API?
  3. A team switched to a cheaper model but cost per completed task rose. Why might that happen?

Put it into practice

Add the cost meter to your agent, enable prefix caching, and compare cost per completed task and cache read tokens before and after on your eval set.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.