Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICost, scale and reliability · Lesson 10 of 19

Prompt caching across providers

Article · 14 min · 9 min lecture

Video lecture

Prompt caching across providers

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Prompt caching

  • Pay once for the stable part
  • Three providers' approaches
  • Designing for cache hits

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The idea

Many requests share a long, identical beginning: a system prompt, tool definitions, a policy document, a product catalog, few-shot examples, or the conversation so far. Prompt caching lets the provider reuse the processing of that shared prefix, making repeated requests cheaper and faster. Caching is a prefix match: any change early in the request invalidates everything after it.

How the three providers do it (verify current details)

ClaudeOpenAIGemini
How to enablecache_control: {"type": "ephemeral"} on specific blocks (up to 4 breakpoints), or top-level automatic cachingAutomatic for sufficiently long prompts; optional prompt_cache_key to improve routing; configurable retention on supported modelsImplicit caching on supported models; explicit caches via client.caches.create(...) referenced with cached_content
Lifetime5-minute default TTL, optional 1-hour TTL; reads refresh the timerProvider-managed (in-memory by default; extended retention options on some models)Implicit: provider-managed; explicit: you set a TTL (storage is billed)
Pricing shapeCache reads ≈ 0.1× base input; writes 1.25× (5 min) or 2× (1 h)Discounted cached input tokensDiscounted cached tokens; explicit caches add storage cost
How to verifyusage.cache_read_input_tokens, usage.cache_creation_input_tokensusage.input_tokens_details.cached_tokensusage_metadata.cached_content_token_count

Minimum cacheable lengths exist on every provider (short prompts silently don't cache). Always read the usage fields to confirm.

Designing for cache hits

  1. Order content from most stable to least stable: tools → system prompt → reference documents → few-shot examples → conversation history → the new user message.
  2. Freeze the stable part byte-for-byte: no timestamps, request IDs, user names or random ordering in it. Sort tool definitions deterministically. Serialize JSON with stable key order.
  3. Put volatile data at the end: today's date, the user's question, retrieved snippets.
  4. Keep the same model: caches are specific to a model; routing between models forfeits cache reuse.
  5. Watch timing: with short TTLs, traffic gaps longer than the TTL mean cold caches; consider longer TTL options for bursty traffic, and compare the write premium against expected reads.

Hands-on: Claude cache breakpoints and verification

import os
import anthropic

client = anthropic.Anthropic()
POLICY = open("returns_policy_uk.md", encoding="utf-8").read()     # long, stable document

def answer(question: str):
    resp = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=600,
        system=[
            {"type": "text", "text": "You answer customer questions using only the policy below."},
            {"type": "text", "text": POLICY, "cache_control": {"type": "ephemeral"}},   # cache up to here
        ],
        messages=[{"role": "user", "content": question}])                             # volatile part last
    u = resp.usage
    print(f"cache_write={u.cache_creation_input_tokens} cache_read={u.cache_read_input_tokens} "
          f"uncached_in={u.input_tokens} out={u.output_tokens}")
    return "".join(b.text for b in resp.content if b.type == "text")

answer("Can I return sale items?")        # first call writes the cache
answer("How long do refunds take?")       # second call within the TTL should read it

For OpenAI, keep the long stable prefix identical and check resp.usage.input_tokens_details.cached_tokens; set a consistent prompt_cache_key for requests sharing a prefix. For Gemini, either rely on implicit caching (check usage_metadata.cached_content_token_count) or create an explicit cache with a TTL for a large document used by many requests.

Worked example: a UK retailer's support assistant (illustrative numbers)

Each request carried a 12,000-token returns-and-delivery policy plus 3,000 tokens of tool definitions, and a 150-token question. After moving the date stamp out of the system prompt, sorting tools deterministically and adding a cache breakpoint after the policy, most requests read the prefix from cache during business hours. Input cost per conversation dropped sharply, and time to first token improved because cached tokens are processed faster.

Explicit caching on Gemini, briefly

For a large document that many requests will reference over a known window (for example a 200-page product manual used all day by a support team), Gemini's explicit caches let you upload the content once with client.caches.create(model=..., config=types.CreateCachedContentConfig(contents=[...], ttl=...)) and then pass cached_content=cache.name in each request's config. You pay a storage cost while the cache lives, so size the TTL to the working day and delete caches you no longer need.

Common silent invalidators

  • A timestamp or "current date" in the system prompt.
  • Tool lists built from a dictionary with non-deterministic order.
  • Per-user personalization inserted before the shared document.
  • Switching models mid-conversation (caches are model-scoped).
  • JSON serialized with varying key order or whitespace.

Measuring success

Cache hit ratio (cached input tokens ÷ total input tokens), cost per conversation before and after, time to first token, and the number of cache writes per hour (unexpected writes reveal invalidators).

Key takeaways

  • Prompt caching reuses processing of an identical request prefix to cut cost and latency.
  • Claude uses explicit cache_control breakpoints or automatic caching; OpenAI caches automatically; Gemini has implicit and explicit caches.
  • Order content from stable to volatile and freeze the stable part byte-for-byte.
  • Verify with each provider's usage fields; short prompts and model switches don't cache.
  • Hunt silent invalidators: timestamps, unsorted tools, early personalization.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. cache_read_input_tokens stays at zero across repeated Claude requests. What is a likely cause?
  2. Where should the user's new question go for best caching?
  3. Why can routing each request to a different model reduce savings?

Put it into practice

Add caching to one repeated-prefix workload, log the provider's cache usage fields for 50 requests, remove any invalidators you find, and report the hit ratio before and after.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.