Advanced Prompt EngineeringPrompt operations: versioning, cost and latency · Lesson 17 of 17

Prompt caching in practice

Article · 12 min · 8 min lecture

Video lecture

Prompt caching in practice

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Prompt caching in practice

  • What it is
  • Provider approaches
  • Ordering for hits
  • Cache killers and verification

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What prompt caching is

Many requests share a long, identical beginning: the same system prompt, tool definitions, reference documents or conversation history. Prompt caching lets the provider store the processed form of that prefix and reuse it on later requests. Cached input is billed at a fraction of the normal input rate and reduces time to first token, especially for long prefixes.

The one rule to remember: caching is a prefix match. The cache can only be reused up to the first byte that differs. Any change early in the prompt invalidates everything after it.

How the major providers handle it

  • Anthropic (Claude). You mark cache breakpoints with cache_control on content blocks (or enable automatic caching for the request). The default cache lifetime is short (minutes) and refreshed on each hit, with a longer lifetime option at a higher write price. Writing to the cache costs slightly more than normal input; reading costs much less. There is a minimum prefix length that varies by model, below which nothing is cached.
  • OpenAI. Caching is automatic for sufficiently long prompts; cached tokens are reported in the usage details and billed at a discount.
  • Google (Gemini). Offers implicit caching on supported models and explicit context caching, where you create a cache object and reference it.

Exact discounts, minimum lengths and lifetimes vary by model and change over time; check the current pricing and caching documentation.

Designing prompts for cache hits

Order content from most stable to most variable:

1. Tool definitions (stable, deterministic order)
2. System prompt (frozen text, no timestamps or user names)
3. Reference documents / few-shot examples (stable per workload)
4. Conversation history (grows, but earlier turns stay identical)
5. The new user message (always changes)

Common cache killers:

  • A timestamp, request ID or user name at the top of the system prompt.
  • Tool lists assembled in a different order on each request, or JSON serialised without sorted keys.
  • Editing or truncating earlier conversation turns instead of appending.
  • Switching models or certain request settings mid-conversation, since caches are specific to a model and configuration.

If you need to give the model the current date or user context, put it after the cached prefix, in the latest user message.

Hands-on: caching a long reference document with Claude

import os, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")
HANDBOOK = open("docs/staff-handbook-2026.md", encoding="utf-8").read()  # long and stable

SYSTEM = [
    {"type": "text", "text": "You answer staff questions using only the handbook. "
                             "Quote the relevant section. If it is not covered, say so."},
    {"type": "text", "text": f"<handbook>\n{HANDBOOK}\n</handbook>",
     "cache_control": {"type": "ephemeral"}},  # cache everything up to and including this block
]

def ask(question: str):
    resp = client.messages.create(model=MODEL, max_tokens=1024, system=SYSTEM,
                                  messages=[{"role": "user", "content": question}])
    u = resp.usage
    print(f"cache write={u.cache_creation_input_tokens} read={u.cache_read_input_tokens} "
          f"uncached={u.input_tokens}")
    return "".join(b.text for b in resp.content if b.type == "text")

ask("How many days of annual leave do new staff get?")   # first call writes the cache
ask("Can I carry leave over to next year?")              # later calls should read it

If cache_read_input_tokens stays at zero across repeated calls, something in the prefix is changing, or the prefix is shorter than the model's minimum. Diff the exact request bodies to find the culprit.

For OpenAI, inspect the cached-token count in the response usage details; for Gemini, check the cached content token count in usage metadata.

Caching in multi-turn chats and agents

In a chat or agent loop, the conversation grows by appending. Place a cache breakpoint near the end of the stable history so each new turn reuses everything before it. Append-only histories cache well; histories that are rewritten each turn (for example re-summarised or reordered) do not. When you must compact, do it deliberately and accept one cache miss, rather than rewriting history every turn.

Worked example: an agency's content assistant

A Karachi agency runs a content assistant with a 20,000-token brand and style pack per client (illustrative size). Each writer sends many short requests per hour. Before caching, every request paid full input price for the brand pack. After moving the brand pack into a cached system block, keeping it byte-identical per client, and moving the date and writer name into the user message, most requests read from cache. Cost per request drops sharply and responses start faster. Their monitoring alerts if the cache-hit rate for any client falls below a threshold, which once caught a template change that inserted the current date into the system prompt.

Economics: when caching pays

Caching pays when the same long prefix is reused within the cache lifetime. It helps less for one-off requests, very short prompts, or traffic spread so thinly that entries expire between requests. For bursty workloads, some teams pre-warm the cache before a batch of requests; for steady traffic, the default lifetime is usually enough.

How to measure success

Track cache-hit ratio (cached input tokens divided by total input tokens), cost per request and time to first token, per route. A healthy long-prefix route should show most input tokens served from cache after the first request.

Key takeaways

  • Prompt caching reuses a processed identical prefix, cutting input cost and time to first token.
  • Order content from stable to variable: tools, system prompt, reference material, history, new message.
  • Timestamps, reordered tools, unsorted JSON and rewritten history silently break caching.
  • Verify with cache usage fields and track cache-hit ratio per route.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your cache read tokens stay at zero across identical-looking requests. Which cause is most likely?
  2. Where should the current date and user name go in a cache-friendly prompt?
  3. When does prompt caching help least?

Put it into practice

Pick one route with a long stable prefix, restructure it from stable to variable, add caching, and log cache read, write and uncached tokens for 20 requests before and after.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.