Advanced Prompt EngineeringPrompt operations: versioning, cost and latency · Lesson 17 of 17
Prompt caching in practice
Video lecture
Prompt caching in practice
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Prompt caching in practice
A Karachi agency's content assistant sends the same twenty-thousand-token brand pack with every single request, hundreds of times an hour. Every time, they pay full price to process it again. Prompt caching fixes that, but only if you understand one rule. In this lecture you will learn what prompt caching is, how the major providers implement it, how to order prompts for cache hits, the silent cache killers, how to verify hits in code, and when caching pays.
0:34 The one rule
Many requests share a long, identical beginning: the same system prompt, tool definitions, reference documents or conversation history. Prompt caching lets the provider store the processed form of that prefix and reuse it on later requests. Cached input is billed at a fraction of the normal input rate and reduces time to first token. Now the one rule: caching is a prefix match. The cache can only be reused up to the first byte that differs. Any change early in the prompt invalidates everything after it.
1:11 Provider approaches
Providers implement it differently. With Claude, you mark cache breakpoints with cache control on content blocks, or enable automatic caching for the request. The default lifetime is short, minutes, refreshed on each hit, with a longer option at a higher write price. Writing to the cache costs slightly more than normal input, and reading costs much less. There is a minimum prefix length that varies by model. OpenAI caches automatically for sufficiently long prompts and reports cached tokens in usage details. Google's Gemini offers implicit caching on supported models, plus explicit context caching where you create and reference a cache object. Check current docs for discounts, minimums and lifetimes.
1:58 Order: stable → variable
Order content from most stable to most variable. First, tool definitions, in a deterministic order. Second, the system prompt, frozen, with no timestamps or user names. Third, reference documents and few-shot examples that are stable for the workload. Fourth, the conversation history, which grows but whose earlier turns stay identical. And last, the new user message, which always changes. If the model needs today's date or the user's name, put them after the cached prefix, in the latest user message.
2:33 Silent cache killers
Now the silent cache killers. A timestamp, request ID or user name at the top of the system prompt. Tool lists assembled in a different order on each request, or JSON serialised without sorted keys. Editing or truncating earlier conversation turns instead of appending. And switching models or certain request settings mid-conversation, because caches are specific to a model and configuration. None of these raise an error. The cache just quietly stops working, and your bill quietly goes up.
3:07 Verify hits
Verify in code. In the lesson's example, a long staff handbook sits in a system block with cache control set. After each call, the code prints three usage numbers: tokens written to the cache, tokens read from the cache, and uncached input tokens. The first question writes the cache; later questions should read from it. If cache reads stay at zero across repeated calls, something in the prefix is changing, or the prefix is below the model's minimum. Diff the exact request bodies to find the culprit. For OpenAI, check the cached-token count in usage details; for Gemini, the cached content token count in usage metadata.
3:53 Caching in chats and agents
In chats and agent loops, the history grows by appending. Place a cache breakpoint near the end of the stable history so each new turn reuses everything before it. Append-only histories cache well. Histories that are re-summarised or reordered every turn do not. When you must compact, do it deliberately and accept one cache miss, rather than rewriting history every turn.
4:20 Worked example + economics
Back to the agency. They moved the brand pack into a cached system block, kept it byte-identical per client, and moved the date and writer's name into the user message. Most requests now read from the cache, cost per request dropped sharply, and responses start faster. They also added monitoring that alerts if any client's cache-hit rate falls below a threshold. It once caught a template change that inserted the current date into the system prompt. Caching pays when the same long prefix is reused within the cache lifetime. It helps little for one-off, short or sparse requests; for bursty workloads, some teams pre-warm the cache.
5:06 Example 1: one line breaks the cache
A simple worked example. You have a chat assistant whose system prompt starts with: today is Tuesday, the twenty-third of September, at fourteen oh five. Then comes fifteen thousand tokens of product documentation. Every request has a different time in the first line, so nothing after it ever caches. Move the date and time into the latest user message, keep the documentation byte-identical at the top, and add a cache breakpoint after it. The first request writes the cache. The next requests within the cache lifetime read it. Your usage logs now show large cache reads and small uncached input. One line moved, most of the prefix saved.
5:53 Example 2: contract Q&A (illustrative)
Now a business scenario, with illustrative numbers. A legal-tech startup in London offers a contract Q&A assistant. Lawyers upload a contract, typically about sixty thousand tokens, and ask ten to twenty questions per session. Without caching, every question pays full input price for the whole contract. The team puts the contract in a cached block after a frozen system prompt, and appends questions and answers to the history without rewriting it. Across a session of fifteen questions, the contract is written to cache once and read fourteen times. With cache reads priced at a small fraction of normal input, the input cost per session falls by well over half, and time to first token drops noticeably from the second question onward. Their dashboard tracks cache-hit ratio per session, and once caught a bug where a trailing space was appended to the contract text on every turn. Illustrative figures.
6:57 Recap
To recap. Caching reuses an identical processed prefix, cutting input cost and time to first token. Order content from stable to variable, keep the prefix byte-identical, append rather than rewrite history, and verify with usage fields. Track cache-hit ratio, cost per request and time to first token per route. Try this now: pick one route with a long stable prefix, restructure it, add caching, and log cache reads, writes and uncached tokens for twenty requests before and after. That completes the course. Keep your evals running, and keep measuring.
What prompt caching is
Many requests share a long, identical beginning: the same system prompt, tool definitions, reference documents or conversation history. Prompt caching lets the provider store the processed form of that prefix and reuse it on later requests. Cached input is billed at a fraction of the normal input rate and reduces time to first token, especially for long prefixes.
The one rule to remember: caching is a prefix match. The cache can only be reused up to the first byte that differs. Any change early in the prompt invalidates everything after it.
How the major providers handle it
- Anthropic (Claude). You mark cache breakpoints with
cache_controlon content blocks (or enable automatic caching for the request). The default cache lifetime is short (minutes) and refreshed on each hit, with a longer lifetime option at a higher write price. Writing to the cache costs slightly more than normal input; reading costs much less. There is a minimum prefix length that varies by model, below which nothing is cached. - OpenAI. Caching is automatic for sufficiently long prompts; cached tokens are reported in the usage details and billed at a discount.
- Google (Gemini). Offers implicit caching on supported models and explicit context caching, where you create a cache object and reference it.
Exact discounts, minimum lengths and lifetimes vary by model and change over time; check the current pricing and caching documentation.
Designing prompts for cache hits
Order content from most stable to most variable:
1. Tool definitions (stable, deterministic order)
2. System prompt (frozen text, no timestamps or user names)
3. Reference documents / few-shot examples (stable per workload)
4. Conversation history (grows, but earlier turns stay identical)
5. The new user message (always changes)Common cache killers:
- A timestamp, request ID or user name at the top of the system prompt.
- Tool lists assembled in a different order on each request, or JSON serialised without sorted keys.
- Editing or truncating earlier conversation turns instead of appending.
- Switching models or certain request settings mid-conversation, since caches are specific to a model and configuration.
If you need to give the model the current date or user context, put it after the cached prefix, in the latest user message.
Hands-on: caching a long reference document with Claude
import os, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")
HANDBOOK = open("docs/staff-handbook-2026.md", encoding="utf-8").read() # long and stable
SYSTEM = [
{"type": "text", "text": "You answer staff questions using only the handbook. "
"Quote the relevant section. If it is not covered, say so."},
{"type": "text", "text": f"<handbook>\n{HANDBOOK}\n</handbook>",
"cache_control": {"type": "ephemeral"}}, # cache everything up to and including this block
]
def ask(question: str):
resp = client.messages.create(model=MODEL, max_tokens=1024, system=SYSTEM,
messages=[{"role": "user", "content": question}])
u = resp.usage
print(f"cache write={u.cache_creation_input_tokens} read={u.cache_read_input_tokens} "
f"uncached={u.input_tokens}")
return "".join(b.text for b in resp.content if b.type == "text")
ask("How many days of annual leave do new staff get?") # first call writes the cache
ask("Can I carry leave over to next year?") # later calls should read itIf cache_read_input_tokens stays at zero across repeated calls, something in the prefix is changing, or the prefix is shorter than the model's minimum. Diff the exact request bodies to find the culprit.
For OpenAI, inspect the cached-token count in the response usage details; for Gemini, check the cached content token count in usage metadata.
Caching in multi-turn chats and agents
In a chat or agent loop, the conversation grows by appending. Place a cache breakpoint near the end of the stable history so each new turn reuses everything before it. Append-only histories cache well; histories that are rewritten each turn (for example re-summarised or reordered) do not. When you must compact, do it deliberately and accept one cache miss, rather than rewriting history every turn.
Worked example: an agency's content assistant
A Karachi agency runs a content assistant with a 20,000-token brand and style pack per client (illustrative size). Each writer sends many short requests per hour. Before caching, every request paid full input price for the brand pack. After moving the brand pack into a cached system block, keeping it byte-identical per client, and moving the date and writer name into the user message, most requests read from cache. Cost per request drops sharply and responses start faster. Their monitoring alerts if the cache-hit rate for any client falls below a threshold, which once caught a template change that inserted the current date into the system prompt.
Economics: when caching pays
Caching pays when the same long prefix is reused within the cache lifetime. It helps less for one-off requests, very short prompts, or traffic spread so thinly that entries expire between requests. For bursty workloads, some teams pre-warm the cache before a batch of requests; for steady traffic, the default lifetime is usually enough.
How to measure success
Track cache-hit ratio (cached input tokens divided by total input tokens), cost per request and time to first token, per route. A healthy long-prefix route should show most input tokens served from cache after the first request.
Key takeaways
- Prompt caching reuses a processed identical prefix, cutting input cost and time to first token.
- Order content from stable to variable: tools, system prompt, reference material, history, new message.
- Timestamps, reordered tools, unsorted JSON and rewritten history silently break caching.
- Verify with cache usage fields and track cache-hit ratio per route.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick one route with a long stable prefix, restructure it from stable to variable, add caching, and log cache read, write and uncached tokens for 20 requests before and after.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.