Claude API Context Caching Explained
Context caching lets you reuse a large chunk of prompt content — like a long system prompt, a document, or a set of tool definitions — across multiple API calls without paying full price to reprocess it every time. Instead of sending the same 20,000-token knowledge base on every request and getting billed for all of it at standard input rates, you cache it once and subsequent calls read from the cache at a much lower cost.
This matters because Claude's pricing is token-based, and most real applications send a lot of repeated context: system instructions, few-shot examples, retrieved documents, tool schemas. Without caching, that repeated content is the single biggest driver of API cost in production apps that make many calls with a stable context. Caching exists specifically to fix that.
How Context Caching Actually Works
When you send a request, you can mark a portion of the prompt as cacheable (typically the front part — system prompt, static instructions, long reference documents). On the first request, Claude processes that content normally and writes it to a cache. On following requests within the cache's lifetime, if the same content appears in the same position, Claude reads from the cache instead of reprocessing it from scratch.
A few mechanics to understand:
- Caching requires an exact prefix match. The cached segment has to be identical, token-for-token, and in the same position in the prompt. If you change even the system prompt's wording slightly, you get a cache miss and it's reprocessed (and re-cached).
- Cache writes cost more than normal input tokens, but cache reads cost significantly less. The economics work in your favor once you're reusing the same context across more than a couple of requests.
- Caches expire. They live for a short time window after each use (minutes, not days), so caching is most valuable for bursts of related calls — a chat session, a batch job, a pipeline processing many documents against the same instructions — not for content you touch once a day.
- Order matters. Put static, reusable content first in your prompt (system instructions, reference docs, tool definitions), and put the variable, per-request content (user's actual question, the specific input) last. Only the stable prefix benefits from caching.
When Caching Actually Pays Off
Context caching makes sense when you have a large, stable chunk of context reused across calls:
- A customer support bot with a long system prompt and product documentation, answering many user questions per session.
- A document Q&A tool where the same document is queried repeatedly.
- A coding assistant with a large codebase context reused across multiple follow-up questions.
- Batch classification or extraction jobs applying the same instructions/schema to hundreds of inputs in a tight time window.
It does not help much for one-off requests, highly varied prompts, or short system prompts where the static portion is small relative to the per-request content. In those cases the cache write overhead isn't worth it.
A Simplified Example
Here's the general shape of a cached request (field names vary by client, but the pattern is consistent across cache-aware APIs):
{
"model": "claude-...",
"system": [
{
"type": "text",
"text": "You are a support agent. Here is the full product manual: ...",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [
{ "role": "user", "content": "How do I reset my password?" }
]
}
The system block (your long, static manual) is marked for caching. The messages array, which changes every request, is not. On the second call with the same system block, you pay the cache-read rate for that content instead of the full input rate.
How This Shows Up in Your Bill
Token usage responses typically break out:
- Input tokens processed normally
- Tokens written to cache (slightly more expensive than standard input)
- Tokens read from cache (cheaper than standard input)
- Output tokens (unchanged)
If you're building on top of Claude through a proxy or gateway, check that the usage metadata actually surfaces this breakdown — otherwise you're flying blind on whether caching is working or just adding write overhead with no reads. If you're running Claude through SubToAPI, usage metadata for each request is visible in the dashboard, which makes it straightforward to confirm your cache hit rate is actually reducing cost rather than just adding overhead. See /docs/messages for request/response shape details.
Practical Tips for Using It Well
- Structure prompts with static content first. This is the single most important rule — caching only works on a matching prefix, so variable content at the top breaks it.
- Batch related calls together in time. Since caches expire quickly, spreading a job out over hours defeats the purpose. Run batch jobs back-to-back.
- Don't cache small content. If your static prefix is a few hundred tokens, the savings aren't meaningful. Caching earns its keep on prompts in the thousands-of-tokens range.
- Watch for accidental cache misses. Timestamps, request IDs, or dynamic values injected into the "static" part of your prompt will silently break the match and you'll pay full write cost every time without realizing it.
- Combine with streaming for latency wins too. A cache hit reduces time-to-first-token as well as cost, since the model skips reprocessing the cached segment. See /docs/streaming for details on handling streamed responses.
If you're integrating Claude into an app and want a simpler path to get started — generating an application key, hitting a standard HTTPS endpoint, and tracking usage per key or per team seat — /docs/quickstart walks through the setup, and /pricing has the plan breakdown (Solo, Team, Scale) if you're evaluating options beyond a direct account.
Questions
Does context caching reduce latency, not just cost? Yes. A cache hit means Claude doesn't reprocess the cached tokens, which lowers time-to-first-token in addition to lowering the bill. The latency benefit is often as valuable as the cost savings for chat-style apps.
How long does a cache last before it expires? Caches are short-lived — typically on the order of minutes after last use, not hours or days. Design your workflow so related calls happen close together in time to actually benefit from it.
Do I need to change my application logic to use caching? Only the prompt structure matters: put static, reusable content in a fixed position (usually first) and mark it as cacheable, and keep variable per-request content separate. No changes are needed to how you parse or handle the response.