Claude API Caching Strategy for Cost Savings
If you're sending the same system prompt, documents, or tool definitions on every request, you're paying full price for tokens that never change. Claude's prompt caching feature lets you store those static chunks of context on Anthropic's side and reuse them across requests at a fraction of the cost — typically around 90% cheaper for cache reads versus fresh input tokens. The strategy that actually saves money isn't "turn caching on," it's deciding what to cache, where to place it in the prompt, and how to structure requests so cache hits stay high.
This article walks through the mechanics of Claude API caching, the structural decisions that determine whether you get real savings, and common mistakes that quietly break your cache hit rate.
How Claude's prompt caching actually works
Caching works by marking specific blocks of your prompt with a cache_control breakpoint. On the first request, Claude processes and stores that block. On subsequent requests within the cache's time-to-live window (typically 5 minutes, with an extended option available), if the same block appears identically up to the breakpoint, Claude reads it from cache instead of reprocessing it from scratch.
Three things determine whether a cache hit happens:
- Exact prefix match. Everything from the start of the prompt up to the cache breakpoint must be byte-identical to a previous request.
- Timing. The cache entry must still be within its TTL. Idle periods longer than the TTL force a full reprocess.
- Placement order. Cached content must come before non-cached, variable content in the request structure.
A typical request layout looks like this:
{
"model": "claude-sonnet-4-5",
"system": [
{
"type": "text",
"text": "You are a support agent. Here is the full product documentation: ...",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [
{ "role": "user", "content": "How do I reset my password?" }
]
}
The large, static documentation block gets cached. The user's question, which changes every request, stays outside the cached region.
What to cache (and what not to)
Caching pays off when a block of content is large, reused across many requests, and stable for at least a few minutes. Good candidates:
- Long system prompts with instructions, persona definitions, or style guides
- Reference documents, knowledge base excerpts, or product manuals used for RAG-style answering
- Tool definitions in multi-tool agent setups, which can run to thousands of tokens on their own
- Few-shot examples used to steer output format
Poor candidates:
- Content that changes on every call, like the current user's message or live data
- Small blocks under a few hundred tokens, where the overhead isn't worth it
- Content shared by only one or two requests before it changes
A common mistake is putting the cache breakpoint after variable content instead of before it. Since caching only matches from the start of the prompt forward, any dynamic text placed before your static block invalidates the cache on every single request. Always order your prompt as: static/cacheable content first, dynamic content last.
Structuring multi-turn conversations for cache reuse
In chat applications, the conversation history grows with every turn. If you cache the entire growing history, you get diminishing returns because each new turn still invalidates the previous cache (the suffix changes). A better pattern is to cache in layers:
- Cache the system prompt and any fixed reference material as one breakpoint (long TTL, rarely changes)
- Cache the conversation history up to the second-to-last turn as a separate breakpoint
- Leave only the newest user message and the model's next response uncached
This way, each new message only reprocesses the small, new portion of the conversation rather than the entire history, and you can stack up to multiple cache breakpoints per request depending on the model.
Measuring whether your caching strategy works
Every response includes usage fields that break down cache behavior: cache creation tokens (first write), cache read tokens (hits), and regular input tokens. Track the ratio of cache read tokens to total input tokens over time. If that ratio is low, your cache breakpoints are probably misplaced, your TTL is too short for your traffic pattern, or your "static" content isn't actually static (timestamps, request IDs, or formatting differences sneaking into the cached block are common culprits).
If you're running Claude through a proxy or API layer, make sure it passes cache_control fields through unmodified and surfaces the usage metadata back to you — stripping or rewriting those fields silently kills your savings. SubToAPI exposes this usage breakdown per request in the dashboard, which makes it straightforward to see which endpoints or prompt templates are actually hitting cache versus reprocessing every time. See /docs/messages for the request format and /docs for the full API reference.
A practical caching strategy checklist
- Put static, reusable content (system prompts, docs, tool specs) at the start of the request, before any variable content
- Use
cache_controlbreakpoints on blocks over roughly 1,000 tokens — smaller blocks rarely justify the overhead - For chat apps, cache history up to the second-to-last turn, not the whole conversation
- Keep request traffic frequent enough to stay inside the TTL window, or use the extended TTL option for lower-traffic endpoints
- Monitor cache read vs. creation token ratios and adjust breakpoint placement when the ratio drops
- Avoid injecting timestamps, request IDs, or random ordering into content you intend to cache
If you're prototyping this and want to see cache metrics without building your own usage dashboard, you can start a free trial at /signup and compare costs across Solo, Team, and Scale plans on /pricing.
Questions
Does prompt caching reduce latency as well as cost? Yes. Cache reads skip the processing step for that portion of the prompt, which typically makes responses with large cached context noticeably faster than an equivalent fully-processed request.
How long does a Claude API cache entry last? The default ephemeral cache TTL is short, around 5 minutes of inactivity, with an extended TTL option available for content you reuse less frequently but still want cached.
Can I cache more than one block in a single request? Yes, you can set multiple cache_control breakpoints in one request — for example, one for a system prompt and another for conversation history — as long as each cached segment is placed before the dynamic content that follows it.