Claude API Prompt Caching: A Complete Feature Guide
Prompt caching is a Claude API feature that lets you store large, reusable chunks of a prompt — system instructions, reference documents, tool definitions, few-shot examples — so subsequent requests don't have to reprocess them from scratch. If your app sends the same long context on every call (a common pattern for RAG apps, coding assistants, and chatbots with big system prompts), caching can cut both latency and cost significantly.
The short answer: you mark a block of your prompt with a cache_control breakpoint, Claude stores the processed state of everything up to that point, and future requests that reuse the same prefix get a much cheaper, faster "cache hit" instead of paying full price to reprocess those tokens. This guide covers how it actually works, what you can and can't cache, pricing mechanics, and a working implementation.
How prompt caching actually works
When you send a request to the Messages API, Claude processes your input tokens before generating a response. Normally, every token in every request is processed fresh, even if 90% of the prompt is identical to the last call. Prompt caching changes that by letting you designate a cache breakpoint — a point in the prompt after which everything before it can be reused.
On the first request with a given prefix, Claude writes that prefix to cache. This "cache write" costs slightly more than a normal input token. On every subsequent request that sends the identical prefix, Claude reads from cache instead of reprocessing it — a "cache hit," which is dramatically cheaper and faster than a full pass.
The cache is scoped to the exact token sequence, so even a single character change before the breakpoint invalidates it for that segment.
What you can cache
Prompt caching works on any content block you place before a cache_control marker, including:
- System prompts — especially long ones with detailed instructions, persona definitions, or formatting rules
- Tool definitions — if you're using tool use with many tools or verbose schemas
- Reference documents — a knowledge base excerpt, codebase context, or PDF content included in every turn
- Few-shot examples — multiple example input/output pairs used to steer output format
- Conversation history — in multi-turn chats, you can cache everything except the latest user message
You can set multiple cache breakpoints in a single request (up to four), which is useful when you have both a static system prompt and a semi-static document that changes less often than the live conversation.
Cache lifetime and pricing mechanics
Caches aren't permanent. By default, a cache entry lives for 5 minutes from the last time it was accessed — each hit refreshes the TTL. There's also an extended 1-hour cache option for workloads with sparser traffic, at a higher write cost.
Pricing works like this, relative to standard input token cost:
- Cache write: roughly 1.25x the normal input price (5-minute cache) or around 2x (1-hour cache)
- Cache hit (read): roughly 0.1x the normal input price — a 90% discount
- Cache miss / no caching used: standard input pricing applies
This means caching pays off when you reuse the same prefix multiple times within the cache window. A single request that never repeats gains nothing from caching — you'd just eat the write premium. But for a chatbot with a 2,000-token system prompt handling dozens of requests per minute, the savings compound fast.
Implementing it
Here's a minimal example using the Messages API directly, caching a long system prompt:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are a support agent for Acme Corp. [... several thousand tokens of policy and tone guidelines ...]",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{"role": "user", "content": "How do I reset my password?"}
]
}'
The response includes usage fields showing how many tokens were written to cache versus read from it, so you can confirm caching is actually kicking in:
{
"usage": {
"input_tokens": 12,
"cache_creation_input_tokens": 1840,
"cache_read_input_tokens": 0,
"output_tokens": 58
}
}
On the next call with the identical system block, you'd expect cache_read_input_tokens to be populated instead, confirming a hit.
Caching through SubToAPI
If you're accessing Claude through SubToAPI — which turns your existing Claude access into a standard HTTPS API with sub_live_... keys — prompt caching works the same way, since requests pass through to the same Messages format. You still set cache_control breakpoints in your request body; SubToAPI adds usage metadata per key so you can see cache read/write token counts in your dashboard, which is useful for tracking whether your team's caching strategy is actually reducing cost across seats. See the quickstart for setup, or check streaming docs if you're combining caching with streamed responses — caching works fine alongside streaming, it only affects the input side.
Best practices
- Put static content first, dynamic content last. Cache breakpoints only cover the prefix, so structure your prompt with system instructions and reference docs at the top, user-specific content at the bottom.
- Don't cache things that change every request. Timestamps, user IDs, or session-specific data inside a cached block will break the cache on every call.
- Watch the TTL for low-traffic endpoints. If your app gets a request every 10 minutes, the 5-minute cache won't help — consider the 1-hour option instead.
- Check usage metadata. Always verify
cache_read_input_tokensin responses; a misplaced breakpoint or a subtle prompt change can silently disable caching without erroring out. - Cache tool definitions separately from documents if both are large, using multiple breakpoints so a document update doesn't invalidate your tool cache.
questions
Does prompt caching change Claude's output quality? No. Caching only affects how input tokens are processed internally — it has no effect on the model's reasoning or the content of its responses.
Is prompt caching worth it for low-volume apps? Usually not. If a given prefix is only sent once, you pay the cache write premium with no read discount to offset it. It pays off when the same prefix is reused multiple times within the TTL window.
Can I combine prompt caching with tool use and streaming? Yes. Caching applies to input tokens regardless of whether you're using tool use or streaming responses — set cache_control on the relevant blocks and the rest of the request works as normal.