Claude API Caching Strategies to Save Money
If you're spending more on Claude API calls than expected, caching is almost always the fastest lever to pull. There are two distinct caching strategies that save money on the Claude API, and most teams only implement one of them: prompt caching (reusing large, repeated context so Claude doesn't re-process it every call) and response caching (storing and reusing full completions for identical or near-identical requests so you skip the model call entirely).
This guide covers both, how to combine them, how to avoid the bugs that cause stale or wrong answers, and how to measure whether caching is actually saving you money versus just adding complexity.
Why caching matters for Claude API costs
Claude API pricing is per-token, and most real-world costs come from two places: large system prompts or context sent on every request, and repeated or near-duplicate user queries. Caching attacks both.
- Prompt caching cuts input token costs by letting Claude reuse a previously processed prefix (system prompt, tool definitions, long reference documents) instead of re-reading it from scratch.
- Response caching cuts both input and output token costs by avoiding the model call entirely when you already have the answer.
Used together, teams commonly cut 30-70% off their Claude bill without changing the product experience at all.
Strategy 1: Prompt caching for repeated context
If your app sends the same system prompt, the same tool schema, or the same large document on every request — think RAG pipelines, coding assistants, or customer support bots with a long knowledge base — prompt caching is the highest-leverage optimization available.
How it works: you mark a stable prefix of your prompt (system instructions, a document, few-shot examples) as cacheable. On subsequent calls within the cache's TTL, Claude reuses the already-processed version of that prefix instead of reprocessing every token.
What to cache:
- System prompts longer than a few hundred tokens
- Tool/function definitions that don't change per request
- Reference documents, style guides, or knowledge base chunks reused across users
- Few-shot examples embedded in the prompt
What NOT to cache:
- The actual user message (it changes every call, caching it does nothing)
- Content with timestamps, user IDs, or session-specific data mixed into the "stable" prefix — this breaks cache hits silently
A common mistake: interleaving dynamic content inside the cacheable block. If your system prompt includes Current user: {user_id} at the top, every request has a technically different prefix and you get zero cache hits while thinking you're saving money. Keep dynamic fields at the very end of the prompt, after everything cacheable.
// Structure: stable content first, dynamic content last
const systemPrompt = `
${LARGE_STABLE_INSTRUCTIONS}
${TOOL_DEFINITIONS_JSON}
`.trim();
const userMessage = `User ID: ${userId}\n\nQuery: ${query}`;
Strategy 2: Response caching for repeated questions
Prompt caching saves on input tokens. Response caching saves on everything, because you skip the API call completely. This works well for:
- FAQ-style assistants where many users ask the same questions
- Classification or extraction tasks with a bounded input space
- Any workflow where the same input (or a normalized version of it) recurs
Basic approach: hash the normalized request (model, system prompt, user input, temperature, tool config) and use it as a cache key in Redis, Memcached, or even an in-process LRU cache for low-traffic apps.
import crypto from "crypto";
function cacheKey(request) {
const normalized = JSON.stringify({
model: request.model,
system: request.system,
messages: request.messages,
temperature: request.temperature ?? 1,
});
return crypto.createHash("sha256").update(normalized).digest("hex");
}
async function getCachedOrCall(request, redisClient, callFn) {
const key = `claude:${cacheKey(request)}`;
const cached = await redisClient.get(key);
if (cached) return JSON.parse(cached);
const response = await callFn(request);
await redisClient.set(key, JSON.stringify(response), "EX", 3600); // 1 hour TTL
return response;
}
Important caveat: response caching only works for deterministic or near-deterministic use cases. If temperature is above 0 and creative variation matters to users, caching identical responses can make your product feel robotic or repetitive. Reserve it for extraction, classification, structured output, and factual Q&A — not creative writing or open-ended chat.
Choosing TTLs without breaking correctness
TTL (time-to-live) decisions are where caching strategies quietly go wrong. Set it too long and users get stale answers after your knowledge base updates; set it too short and you lose most of the savings.
A practical framework:
| Data type | Suggested TTL | Reason | |---|---|---| | Static reference docs, tool schemas | 24h+ | Rarely change | | Product/FAQ knowledge base | 1-6h | Updates occasionally | | User-specific session context | Minutes | Changes per conversation | | Time-sensitive data (prices, availability) | Do not cache responses | Correctness risk outweighs savings |
When in doubt, cache the prompt prefix aggressively and the response conservatively — the cost of a stale system prompt is low, the cost of a stale factual answer is a support ticket.
Monitoring whether caching is actually saving money
Caching without measurement is guesswork. Track three numbers:
- Cache hit rate — what percentage of requests are served from cache or hit the cached prompt prefix.
- Token cost per request, before and after — compare average tokens billed with caching enabled versus disabled on a sample.
- Staleness incidents — how often a cached response was wrong because the underlying data changed.
If you're routing Claude API traffic through SubToAPI, every request already returns usage metadata (input/output tokens, cache status) alongside the response, so you can log and aggregate these numbers without building separate telemetry. Combined with per-key usage tracking across a team, it's straightforward to see which endpoints or features are burning the most tokens and would benefit most from caching — see /docs/messages for the response format and /pricing for plan details if you're evaluating a managed layer on top of raw API access.
A simple decision framework
- Same large context on every call, varying user question → prompt caching
- Same or near-identical full requests recurring → response caching
- Both conditions true → use both, prompt caching for the shared context, response caching for the final answer
- Highly dynamic, creative, or time-sensitive output → cache nothing, optimize elsewhere (shorter prompts, smaller models for simple tasks)
Questions
Does prompt caching change the model's output? No. Prompt caching only affects how previously-seen context is processed internally; it does not alter what the model generates. Response caching, by contrast, literally returns a stored prior output, so it's only appropriate when identical output is desired.
How do I know if my cache hit rate is good enough? It depends on traffic patterns, but below 20% hit rate the operational complexity of caching (invalidation logic, storage, monitoring) often isn't worth it. Above 50% hit rate on either prompt or response caching usually translates into meaningful, visible cost reduction.
Can I combine caching with a lower-cost model for simple requests? Yes, and this compounds well: cache repeated heavy context, and route simple classification or short-answer tasks to a cheaper, faster model tier while reserving full Claude capability for complex requests. Start with /docs/quickstart to see how model selection and request structure work together in practice.