Claude API Prompt Caching: Implementation Guide
What Prompt Caching Actually Does
If you're searching for "claude api prompt caching implementation," you're probably trying to cut latency and token costs on requests that reuse the same large context — a system prompt, a document, a tool schema — across many calls. Claude's prompt caching lets you mark specific blocks of your request as cacheable so Anthropic stores a processed version of them server-side and reuses it on subsequent calls instead of re-running the full prefill every time.
The short answer: you implement it by adding a cache_control object to the content blocks you want cached, choosing a TTL (5 minutes or 1 hour depending on availability), and structuring your requests so the cached portion stays identical byte-for-byte between calls. Get the structure right and you'll see meaningful latency drops and lower cost on the cached tokens for every call after the first one. Get it wrong — reordering blocks, changing whitespace, injecting a timestamp into the cached section — and you silently pay full price with none of the benefit.
How Caching Works Under the Hood
Claude processes your prompt in order: system prompt, then messages, then tools, then the current turn. When you add a cache_control breakpoint after a block of content, Claude hashes everything up to that point and checks if a matching cache entry already exists. If it does, it skips reprocessing that prefix and starts generating from where the cache ends. If it doesn't, it processes everything normally and writes a new cache entry for next time.
This means order and exact content matter. You can have multiple cache breakpoints in a single request (up to four), but each one only helps if the prefix up to that point is identical to a previous call. A common mistake is caching a block that contains something that changes every request — a current date, a request ID, a user's live session data — which invalidates the cache on every single call.
Where to Put Cache Breakpoints
The highest-value places to cache are:
- Long system prompts — instructions, persona definitions, formatting rules that don't change between requests
- Large reference documents — a PDF's extracted text, a knowledge base excerpt, an API spec you're asking Claude to reason about
- Tool definitions — if you're passing 10+ tools with detailed JSON schemas, those schemas rarely change between calls
- Few-shot examples — a fixed set of example conversations you prepend for consistency
What you should not cache: the user's current message, anything with a timestamp, anything that includes session-specific IDs or variables.
Implementation: Raw Anthropic API
Here's a minimal example showing a cached system prompt and a cached document, with the user's actual question left uncached:
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-opus-4-20250514",
"max_tokens": 1024,
"system": [
{
"type": "text",
"text": "You are a technical support assistant for Acme Cloud. Follow the style guide exactly...",
"cache_control": {"type": "ephemeral"}
}
],
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "<full product documentation, 15,000 tokens>",
"cache_control": {"type": "ephemeral"}
},
{
"type": "text",
"text": "How do I rotate an API key for a team account?"
}
]
}
]
}'
The first call pays full price to write the cache. Every subsequent call within the TTL window that sends the identical system prompt and document text reuses the cached prefix and only pays full price for the new question and the response.
In JavaScript, the structure is the same — you're just adding cache_control to whichever content blocks are stable:
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({
model: "claude-opus-4-20250514",
max_tokens: 1024,
system: [
{ type: "text", text: SYSTEM_PROMPT, cache_control: { type: "ephemeral" } },
],
messages: [
{
role: "user",
content: [
{ type: "text", text: DOCUMENT_TEXT, cache_control: { type: "ephemeral" } },
{ type: "text", text: userQuestion },
],
},
],
}),
});
Checking Whether It's Actually Working
Don't assume caching is active — verify it from the response usage metadata. Anthropic returns fields like cache_creation_input_tokens and cache_read_input_tokens alongside the normal input_tokens and output_tokens. On your first call you should see tokens counted under cache creation; on repeat calls within the TTL, those same tokens should show up under cache read instead.
If you're building this on top of SubToAPI, the same metadata is exposed in the usage data returned with each response and in the dashboard, so you can confirm cache hits are happening per application key without instrumenting your own logging. See /docs/messages for the exact response shape.
Common Implementation Mistakes
- Rebuilding the document string with minor differences each run (extra whitespace, re-serialized JSON with different key order) — this breaks the hash match silently.
- Putting cache_control on the wrong block — caching only helps the content before and including the breakpoint, not content after it.
- Expecting cross-session reuse without re-sending the cached content — you still send the full cached text every request; caching skips the processing, not the transmission.
- Ignoring TTL expiry — if your traffic pattern has long gaps between calls, a 5-minute cache won't help; check whether the longer TTL option is available for your use case before architecting around it.
- Caching highly dynamic tool schemas — if you regenerate tool definitions per request (e.g., dynamic permission-based tool lists), caching won't trigger. See /docs/tools for tool definition patterns that stay stable.
Where SubToAPI Fits
If you're already proxying Claude access through SubToAPI with an application key (sub_live_...), prompt caching works the same way — you pass cache_control blocks exactly as shown above through the /v1/messages endpoint. The benefit of going through SubToAPI is that cache hit/miss data, token counts, and cost per application key are already aggregated in the dashboard, which is useful once you have multiple apps or team members sharing the same underlying Claude access. Check /docs/quickstart for key setup and /pricing for plan details, or /signup to try it on the free trial.
questions
Does prompt caching change the output quality or determinism? No. Caching only skips reprocessing of the cached prefix — the model still generates from the same effective context, so output quality is unaffected.
How long does a cache entry last before it expires? Cache entries are ephemeral and expire after a short TTL (minutes, not hours) unless refreshed by another request that hits the same cache within the window — check current Anthropic documentation for exact TTL options available to your account.
Can I cache part of a conversation and still add new messages after it? Yes. Place cache_control after the stable prefix (system prompt, documents, early turns) and append new, uncached messages after it — only the prefix up to the breakpoint needs to match for the cache to hit.