← Blog

Claude API Prompt Caching: Real Cost Savings Explained

2026-10-07 · 5 min read · SubToAPI Team

Prompt caching on the Claude API lets you reuse large, repeated chunks of a prompt — system instructions, documents, few-shot examples — without paying full price for them on every request. Anthropic charges a premium to write content into the cache (roughly 25% more than a normal input token) but then charges only a fraction of the standard input rate — around 10% — for every subsequent read within the cache's lifetime (typically 5 minutes, refreshed on each hit). For workloads with long static context and frequent repeat calls, this routinely cuts input token costs by 70–90%.

The actual savings depend entirely on your traffic pattern. If you send a 10,000-token system prompt with every request and call the API dozens of times per minute, caching pays for itself almost immediately. If your prompts are short and mostly unique, caching adds overhead without meaningful benefit. This article breaks down how the economics work, where caching helps most, and how to measure whether it's actually saving you money.

How prompt caching pricing actually works

Claude's API splits cached content into two cost events:

So the break-even point is simple: if you re-send the same cached block more than roughly 1–2 times within the cache window, you're already saving money compared to sending it uncached every time. The savings compound fast after that. A block reused 20 times in a session costs roughly (1 write + 19 reads) instead of 20 full-price reads — a massive reduction in total input tokens billed.

What's worth caching

Not every prompt benefits. Good caching candidates share these traits:

Typical examples:

Poor candidates: one-off requests, prompts under a few hundred tokens, or content that changes on every call (user messages, live data, timestamps).

Measuring the real savings

Don't guess — calculate. For a given workload, compare:

uncached_cost = requests * input_tokens * base_rate
cached_cost   = (1 * input_tokens * write_rate)
              + ((requests - 1) * input_tokens * read_rate)

With base_rate normalized to 1, write_rate ≈ 1.25, read_rate ≈ 0.1:

The savings curve flattens as reuse count grows, but it's positive after the very first repeat call. The main risk is cache expiry — if your request rate is too low (gaps longer than the TTL), you keep paying write prices and never benefit from the cheap reads.

Caching in practice

Implementing caching directly against the Claude API means managing cache control blocks in your request structure, tracking which content blocks are marked cacheable, and monitoring usage metadata to confirm hits versus misses. That metadata — cache creation tokens vs. cache read tokens — is the only reliable way to know if caching is actually working, because a silent miss (expired TTL, slightly altered prefix) will bill you full write price without you noticing unless you check the response usage object on every call.

This is one of the areas where a managed layer helps. If you're running Claude through SubToAPI, usage metadata — including cache read and write token counts — is surfaced per request in your dashboard, so you can see exactly how much a given application key is saving without building your own tracking. Combined with per-key usage limits, it's a quick way to catch prompts that should be cached but aren't hitting consistently. See /docs/messages for request structure and /docs/streaming if you're caching context on long-running streamed sessions.

Practical checklist

Prompt caching is one of the few optimizations with almost no downside: it doesn't change output quality, doesn't add latency in practice, and the pricing model guarantees savings once you cross the break-even reuse count. The only real work is structuring prompts correctly and verifying hits in production.

Questions

Does prompt caching reduce output token costs too? No. Caching only affects input tokens — the content you send to the model. Output tokens (what Claude generates) are billed at the normal rate regardless of caching.

What happens if my cached content changes slightly between requests? Any change to the cached prefix — including whitespace — invalidates the cache for that block, triggering a fresh write at the premium rate on the next call instead of a cheap read.

Is prompt caching worth it for low-traffic apps? Only if the same large context gets reused at least twice within the cache TTL (typically 5 minutes). For infrequent, unique requests, caching adds a write premium with no read savings to offset it.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →