← Blog

Claude API Caching Strategy for Cost Savings

2026-10-11 · 5 min read · SubToAPI Team

If you're sending the same system prompt, documents, or tool definitions on every request, you're paying full price for tokens that never change. Claude's prompt caching feature lets you store those static chunks of context on Anthropic's side and reuse them across requests at a fraction of the cost — typically around 90% cheaper for cache reads versus fresh input tokens. The strategy that actually saves money isn't "turn caching on," it's deciding what to cache, where to place it in the prompt, and how to structure requests so cache hits stay high.

This article walks through the mechanics of Claude API caching, the structural decisions that determine whether you get real savings, and common mistakes that quietly break your cache hit rate.

How Claude's prompt caching actually works

Caching works by marking specific blocks of your prompt with a cache_control breakpoint. On the first request, Claude processes and stores that block. On subsequent requests within the cache's time-to-live window (typically 5 minutes, with an extended option available), if the same block appears identically up to the breakpoint, Claude reads it from cache instead of reprocessing it from scratch.

Three things determine whether a cache hit happens:

  1. Exact prefix match. Everything from the start of the prompt up to the cache breakpoint must be byte-identical to a previous request.
  2. Timing. The cache entry must still be within its TTL. Idle periods longer than the TTL force a full reprocess.
  3. Placement order. Cached content must come before non-cached, variable content in the request structure.

A typical request layout looks like this:

{
  "model": "claude-sonnet-4-5",
  "system": [
    {
      "type": "text",
      "text": "You are a support agent. Here is the full product documentation: ...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [
    { "role": "user", "content": "How do I reset my password?" }
  ]
}

The large, static documentation block gets cached. The user's question, which changes every request, stays outside the cached region.

What to cache (and what not to)

Caching pays off when a block of content is large, reused across many requests, and stable for at least a few minutes. Good candidates:

Poor candidates:

A common mistake is putting the cache breakpoint after variable content instead of before it. Since caching only matches from the start of the prompt forward, any dynamic text placed before your static block invalidates the cache on every single request. Always order your prompt as: static/cacheable content first, dynamic content last.

Structuring multi-turn conversations for cache reuse

In chat applications, the conversation history grows with every turn. If you cache the entire growing history, you get diminishing returns because each new turn still invalidates the previous cache (the suffix changes). A better pattern is to cache in layers:

This way, each new message only reprocesses the small, new portion of the conversation rather than the entire history, and you can stack up to multiple cache breakpoints per request depending on the model.

Measuring whether your caching strategy works

Every response includes usage fields that break down cache behavior: cache creation tokens (first write), cache read tokens (hits), and regular input tokens. Track the ratio of cache read tokens to total input tokens over time. If that ratio is low, your cache breakpoints are probably misplaced, your TTL is too short for your traffic pattern, or your "static" content isn't actually static (timestamps, request IDs, or formatting differences sneaking into the cached block are common culprits).

If you're running Claude through a proxy or API layer, make sure it passes cache_control fields through unmodified and surfaces the usage metadata back to you — stripping or rewriting those fields silently kills your savings. SubToAPI exposes this usage breakdown per request in the dashboard, which makes it straightforward to see which endpoints or prompt templates are actually hitting cache versus reprocessing every time. See /docs/messages for the request format and /docs for the full API reference.

A practical caching strategy checklist

If you're prototyping this and want to see cache metrics without building your own usage dashboard, you can start a free trial at /signup and compare costs across Solo, Team, and Scale plans on /pricing.

Questions

Does prompt caching reduce latency as well as cost? Yes. Cache reads skip the processing step for that portion of the prompt, which typically makes responses with large cached context noticeably faster than an equivalent fully-processed request.

How long does a Claude API cache entry last? The default ephemeral cache TTL is short, around 5 minutes of inactivity, with an extended TTL option available for content you reuse less frequently but still want cached.

Can I cache more than one block in a single request? Yes, you can set multiple cache_control breakpoints in one request — for example, one for a system prompt and another for conversation history — as long as each cached segment is placed before the dynamic content that follows it.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →