Claude API Prompt Caching: Real Cost Savings Explained
Prompt caching on the Claude API lets you reuse large, repeated chunks of a prompt — system instructions, documents, few-shot examples — without paying full price for them on every request. Anthropic charges a premium to write content into the cache (roughly 25% more than a normal input token) but then charges only a fraction of the standard input rate — around 10% — for every subsequent read within the cache's lifetime (typically 5 minutes, refreshed on each hit). For workloads with long static context and frequent repeat calls, this routinely cuts input token costs by 70–90%.
The actual savings depend entirely on your traffic pattern. If you send a 10,000-token system prompt with every request and call the API dozens of times per minute, caching pays for itself almost immediately. If your prompts are short and mostly unique, caching adds overhead without meaningful benefit. This article breaks down how the economics work, where caching helps most, and how to measure whether it's actually saving you money.
How prompt caching pricing actually works
Claude's API splits cached content into two cost events:
- Cache write — the first time a block of content is sent with a cache control marker, it's written to the cache at a premium over the base input rate (roughly 1.25x).
- Cache read — every subsequent request that hits the same cached block within the TTL window pays a small fraction of the base input rate (roughly 0.1x, i.e. 90% cheaper).
- Cache miss — if the TTL expires or the content changes, you pay the write price again on the next call.
So the break-even point is simple: if you re-send the same cached block more than roughly 1–2 times within the cache window, you're already saving money compared to sending it uncached every time. The savings compound fast after that. A block reused 20 times in a session costs roughly (1 write + 19 reads) instead of 20 full-price reads — a massive reduction in total input tokens billed.
What's worth caching
Not every prompt benefits. Good caching candidates share these traits:
- Large and static — system prompts, tool definitions, style guides, long reference documents that don't change between calls.
- Reused frequently — if the same context is sent again within minutes, not hours.
- Positioned early in the prompt — Claude caches content as a prefix, so cacheable material needs to come before the variable/user-specific parts of the message.
Typical examples:
- A customer support bot with a 5,000-token knowledge base injected into every system prompt.
- A coding assistant that sends the same large file or codebase context on every turn of a conversation.
- A document Q&A tool where one PDF is cached and queried dozens of times.
- Multi-turn agents that re-send tool definitions and instructions on every step.
Poor candidates: one-off requests, prompts under a few hundred tokens, or content that changes on every call (user messages, live data, timestamps).
Measuring the real savings
Don't guess — calculate. For a given workload, compare:
uncached_cost = requests * input_tokens * base_rate
cached_cost = (1 * input_tokens * write_rate)
+ ((requests - 1) * input_tokens * read_rate)
With base_rate normalized to 1, write_rate ≈ 1.25, read_rate ≈ 0.1:
- 5 requests reusing the same context: uncached = 5.0, cached ≈ 1.25 + 0.4 = 1.65 → 67% savings
- 50 requests: uncached = 50.0, cached ≈ 1.25 + 4.9 = 6.15 → 88% savings
- 2 requests: uncached = 2.0, cached ≈ 1.25 + 0.1 = 1.35 → still 32% savings
The savings curve flattens as reuse count grows, but it's positive after the very first repeat call. The main risk is cache expiry — if your request rate is too low (gaps longer than the TTL), you keep paying write prices and never benefit from the cheap reads.
Caching in practice
Implementing caching directly against the Claude API means managing cache control blocks in your request structure, tracking which content blocks are marked cacheable, and monitoring usage metadata to confirm hits versus misses. That metadata — cache creation tokens vs. cache read tokens — is the only reliable way to know if caching is actually working, because a silent miss (expired TTL, slightly altered prefix) will bill you full write price without you noticing unless you check the response usage object on every call.
This is one of the areas where a managed layer helps. If you're running Claude through SubToAPI, usage metadata — including cache read and write token counts — is surfaced per request in your dashboard, so you can see exactly how much a given application key is saving without building your own tracking. Combined with per-key usage limits, it's a quick way to catch prompts that should be cached but aren't hitting consistently. See /docs/messages for request structure and /docs/streaming if you're caching context on long-running streamed sessions.
Practical checklist
- Put static, reusable content (system prompts, docs, tool definitions) before dynamic user content in the message structure.
- Keep cached blocks identical byte-for-byte between calls — even whitespace changes trigger a cache miss.
- Watch request frequency relative to the cache TTL; sparse traffic won't benefit.
- Check usage metadata on every response to confirm cache reads are actually happening, not just writes.
- Re-evaluate periodically — as your prompts evolve, what was cacheable may no longer be static.
Prompt caching is one of the few optimizations with almost no downside: it doesn't change output quality, doesn't add latency in practice, and the pricing model guarantees savings once you cross the break-even reuse count. The only real work is structuring prompts correctly and verifying hits in production.
Questions
Does prompt caching reduce output token costs too? No. Caching only affects input tokens — the content you send to the model. Output tokens (what Claude generates) are billed at the normal rate regardless of caching.
What happens if my cached content changes slightly between requests? Any change to the cached prefix — including whitespace — invalidates the cache for that block, triggering a fresh write at the premium rate on the next call instead of a cheap read.
Is prompt caching worth it for low-traffic apps? Only if the same large context gets reused at least twice within the cache TTL (typically 5 minutes). For infrequent, unique requests, caching adds a write premium with no read savings to offset it.