---
title: "Prompt caching across providers | Optimize All Academy"
description: "The idea Many requests share a long, identical beginning: a system prompt, tool definitions, a policy document, a product catalog, few-shot examples, or…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/prompt-caching-across-providers
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Cost, scale and reliability · lesson 10 of 19 · 14 min

# Prompt caching across providers

## The idea

Many requests share a long, identical beginning: a system prompt, tool definitions, a policy document, a product catalog, few-shot examples, or the conversation so far. **Prompt caching** lets the provider reuse the processing of that shared **prefix**, making repeated requests cheaper and faster. Caching is a prefix match: any change early in the request invalidates everything after it.

## How the three providers do it (verify current details)

| | Claude | OpenAI | Gemini |
|---|---|---|---|
| How to enable | `cache_control: {"type": "ephemeral"}` on specific blocks (up to 4 breakpoints), or top-level automatic caching | Automatic for sufficiently long prompts; optional `prompt_cache_key` to improve routing; configurable retention on supported models | Implicit caching on supported models; explicit caches via `client.caches.create(...)` referenced with `cached_content` |
| Lifetime | 5-minute default TTL, optional 1-hour TTL; reads refresh the timer | Provider-managed (in-memory by default; extended retention options on some models) | Implicit: provider-managed; explicit: you set a TTL (storage is billed) |
| Pricing shape | Cache reads ≈ 0.1× base input; writes 1.25× (5 min) or 2× (1 h) | Discounted cached input tokens | Discounted cached tokens; explicit caches add storage cost |
| How to verify | `usage.cache_read_input_tokens`, `usage.cache_creation_input_tokens` | `usage.input_tokens_details.cached_tokens` | `usage_metadata.cached_content_token_count` |

Minimum cacheable lengths exist on every provider (short prompts silently don't cache). Always read the usage fields to confirm.

## Designing for cache hits

1. **Order content from most stable to least stable**: tools → system prompt → reference documents → few-shot examples → conversation history → the new user message.
2. **Freeze the stable part byte-for-byte**: no timestamps, request IDs, user names or random ordering in it. Sort tool definitions deterministically. Serialize JSON with stable key order.
3. **Put volatile data at the end**: today's date, the user's question, retrieved snippets.
4. **Keep the same model**: caches are specific to a model; routing between models forfeits cache reuse.
5. **Watch timing**: with short TTLs, traffic gaps longer than the TTL mean cold caches; consider longer TTL options for bursty traffic, and compare the write premium against expected reads.

## Hands-on: Claude cache breakpoints and verification

```python
import os
import anthropic

client = anthropic.Anthropic()
POLICY = open("returns_policy_uk.md", encoding="utf-8").read()     # long, stable document

def answer(question: str):
    resp = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=600,
        system=[
            {"type": "text", "text": "You answer customer questions using only the policy below."},
            {"type": "text", "text": POLICY, "cache_control": {"type": "ephemeral"}},   # cache up to here
        ],
        messages=[{"role": "user", "content": question}])                             # volatile part last
    u = resp.usage
    print(f"cache_write={u.cache_creation_input_tokens} cache_read={u.cache_read_input_tokens} "
          f"uncached_in={u.input_tokens} out={u.output_tokens}")
    return "".join(b.text for b in resp.content if b.type == "text")

answer("Can I return sale items?")        # first call writes the cache
answer("How long do refunds take?")       # second call within the TTL should read it
```

For OpenAI, keep the long stable prefix identical and check `resp.usage.input_tokens_details.cached_tokens`; set a consistent `prompt_cache_key` for requests sharing a prefix. For Gemini, either rely on implicit caching (check `usage_metadata.cached_content_token_count`) or create an explicit cache with a TTL for a large document used by many requests.

## Worked example: a UK retailer's support assistant (illustrative numbers)

Each request carried a 12,000-token returns-and-delivery policy plus 3,000 tokens of tool definitions, and a 150-token question. After moving the date stamp out of the system prompt, sorting tools deterministically and adding a cache breakpoint after the policy, most requests read the prefix from cache during business hours. Input cost per conversation dropped sharply, and time to first token improved because cached tokens are processed faster.

## Explicit caching on Gemini, briefly

For a large document that many requests will reference over a known window (for example a 200-page product manual used all day by a support team), Gemini's explicit caches let you upload the content once with `client.caches.create(model=..., config=types.CreateCachedContentConfig(contents=[...], ttl=...))` and then pass `cached_content=cache.name` in each request's config. You pay a storage cost while the cache lives, so size the TTL to the working day and delete caches you no longer need.

## Common silent invalidators

- A timestamp or "current date" in the system prompt.
- Tool lists built from a dictionary with non-deterministic order.
- Per-user personalization inserted before the shared document.
- Switching models mid-conversation (caches are model-scoped).
- JSON serialized with varying key order or whitespace.

## Measuring success

Cache hit ratio (cached input tokens ÷ total input tokens), cost per conversation before and after, time to first token, and the number of cache writes per hour (unexpected writes reveal invalidators).

## Video lecture: Prompt caching across providers

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Prompt caching
2. Why it matters
3. Caching is a prefix match
4. Provider approaches
5. Verify with usage fields
6. Simple example: the date at the top
7. Business example: UK retail support
8. Gemini explicit caches
9. Silent invalidators
10. Timing and lifetimes
11. Common mistakes
12. Keep caching healthy
13. Deeper: UK retail caching (illustrative)
14. Watch me do it: cache breakpoints
15. Recap + try this now

## Lecture transcript

### Prompt caching

Every time your support assistant answers a question, it might resend the same twelve thousand token policy document. Imagine paying full price to have a colleague re read the entire handbook before every single customer email. Prompt caching lets the provider remember the part that doesn't change, so repeated requests are cheaper and faster. In this lesson you'll learn how caching works on Claude, OpenAI and Gemini, and how to design prompts that actually hit the cache.

### Why it matters

Why does this matter? Because in many applications the stable part of the request is far bigger than the new part. System prompts, tool definitions, policy documents and conversation history often dwarf the user's latest question. Caching can cut the cost of that stable part dramatically and make responses start faster. Think of a coffee shop that pre grinds beans for the morning rush. The first cup takes the full time; every cup after that is quicker because the prep is done.

### Caching is a prefix match

Here's the key idea: caching is a prefix match. The provider can reuse processing only for the beginning of the request, up to the point where something changes. Change one character early on, and everything after it has to be processed again. So the order of your request matters: tools and system prompt first, then reference documents, then examples, then history, and finally the new user message.

### Provider approaches

How do the providers differ? With Claude, you mark cache breakpoints on specific content blocks, up to four, or use automatic caching. The default lifetime is five minutes, with an optional one hour option, and reading the cache refreshes the timer. Cache reads cost roughly a tenth of normal input, and writes cost a little more than normal. OpenAI caches long prompts automatically, and you can pass a prompt cache key to improve reuse. Gemini offers implicit caching on supported models, plus explicit caches you create with a time to live, which are billed for storage. Check current prices before planning.

### Verify with usage fields

How do you know it's working? Read the usage fields. Claude reports cache read and cache creation input tokens. OpenAI reports cached tokens inside the input token details. Gemini reports a cached content token count. If the read number stays at zero across repeated requests, something is changing your prefix. Also remember each provider has a minimum length, so short prompts simply won't cache.

### Simple example: the date at the top

A simple example. A recipe app sends a two thousand word style guide with every request, then a user's question: make this dish vegetarian. At first, the developer adds the current date at the very top of the style guide for freshness. Cache reads stay at zero, because the date changes the prefix every day, and within the day the timestamp includes the time. Moving the date to the end, next to the question, fixes it immediately.

### Business example: UK retail support

Now a realistic business example, with illustrative numbers. A UK retailer's support assistant sends a twelve thousand token returns and delivery policy, plus three thousand tokens of tool definitions, with every question of about a hundred and fifty tokens. The team removed a date stamp from the system prompt, sorted tools in a fixed order, and added a cache breakpoint after the policy. During business hours most requests now read the prefix from cache, input cost per conversation dropped sharply, and time to first token improved.

### Gemini explicit caches

What about Gemini's explicit caches? They suit a specific situation: a large document that many requests will use over a known period, like a two hundred page product manual that the support team queries all day. You upload the content once as a cache with a lifetime, and each request references the cache by name instead of resending the manual. You pay a storage cost while it lives, so size the lifetime to the working day, and delete caches you no longer need.

### Silent invalidators

Let's list the silent invalidators, because they're responsible for most caching failures. A timestamp or current date in the system prompt. Tool lists built from a dictionary with unpredictable order. Personalization, like the user's name, inserted before the shared document. Switching models mid conversation, since caches belong to each model. And JSON serialized with varying key order or spacing. Each one silently turns your cache into an expensive no op.

### Timing and lifetimes

Timing matters too. With a five minute lifetime, a gap of more than five minutes between requests means the next one starts cold. For bursty traffic with longer gaps, a longer lifetime option can pay off, but it usually costs more to write, so compare the premium against how many reads you expect. For Gemini's explicit caches, you pay for storage while they live, so set the lifetime to match your usage window.

### Common mistakes

Common mistakes. Assuming caching works without checking usage fields. Caching tiny prompts that fall below the minimum. Putting per request data at the top. And building elaborate multi model routing that fragments your caches, when one model at a lower effort setting might be cheaper overall.

### Keep caching healthy

Here's a habit that keeps caching healthy long term: add a cache hit ratio chart to your standard dashboard, next to cost and latency. Caching often breaks silently during an unrelated change, like someone adding a greeting with the user's first name to the top of the system prompt. If the hit ratio drops the day that change ships, you'll catch it in minutes instead of discovering it on next month's invoice.

### Deeper: UK retail caching (illustrative)

Let's deepen the UK retailer's support assistant with illustrative numbers. Twelve thousand tokens of policy, three thousand of tools, and a hundred and fifty token question. Before the fixes, every request paid for fifteen thousand input tokens at full price. After removing the date stamp, sorting tools and adding the breakpoint, during business hours the vast majority of those tokens were cache reads at roughly a tenth of the price. Input cost per conversation fell dramatically, and time to first token improved. Overnight, traffic gaps meant the cache went cold, so the first morning requests paid for a cache write; the team accepted that rather than paying for a longer lifetime.

### Watch me do it: cache breakpoints

Watch me do it. Let's walk through the caching example. The policy file is read once at start up. The system parameter is a list of two text blocks: a short instruction, then the full policy with cache control set to ephemeral. That marker says: cache everything up to and including this block. The customer's question goes in messages, after the cached part. After each call, I print four usage numbers: cache creation tokens, cache read tokens, uncached input tokens and output tokens. I ask, can I return sale items. The output shows a large cache write and zero reads. Then, within five minutes, how long do refunds take: now cache write is zero and cache read is large, while uncached input is just the question. If I add today's date to the first system block and run again, reads drop back to zero, and I've just reproduced the most common caching bug.

### Recap + try this now

Quick recap. Caching reuses an identical prefix. Order content from stable to volatile, freeze the stable part, verify with usage fields, and watch for silent invalidators and timing gaps. Try this now: add caching to one workload that repeats a long prefix, log the cache usage fields for fifty requests, remove any invalidators you find, and report your hit ratio before and after.

## Key takeaways

- Prompt caching reuses processing of an identical request prefix to cut cost and latency.
- Claude uses explicit cache_control breakpoints or automatic caching; OpenAI caches automatically; Gemini has implicit and explicit caches.
- Order content from stable to volatile and freeze the stable part byte-for-byte.
- Verify with each provider's usage fields; short prompts and model switches don't cache.
- Hunt silent invalidators: timestamps, unsorted tools, early personalization.

## Try it

Add caching to one repeated-prefix workload, log the provider's cache usage fields for 50 requests, remove any invalidators you find, and report the hit ratio before and after.

- [Previous: Batch APIs: bulk processing at a discount](https://optimizeall.com/learn/ai-platform-apis-integration/batch-apis-for-bulk-work)
- [Next: Rate limits, errors, retries and backoff](https://optimizeall.com/learn/ai-platform-apis-integration/rate-limits-retries-and-backoff)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
