← Blog

Claude API Context Caching Explained

2026-10-09 · 5 min read · SubToAPI Team

Context caching lets you reuse a large chunk of prompt content — like a long system prompt, a document, or a set of tool definitions — across multiple API calls without paying full price to reprocess it every time. Instead of sending the same 20,000-token knowledge base on every request and getting billed for all of it at standard input rates, you cache it once and subsequent calls read from the cache at a much lower cost.

This matters because Claude's pricing is token-based, and most real applications send a lot of repeated context: system instructions, few-shot examples, retrieved documents, tool schemas. Without caching, that repeated content is the single biggest driver of API cost in production apps that make many calls with a stable context. Caching exists specifically to fix that.

How Context Caching Actually Works

When you send a request, you can mark a portion of the prompt as cacheable (typically the front part — system prompt, static instructions, long reference documents). On the first request, Claude processes that content normally and writes it to a cache. On following requests within the cache's lifetime, if the same content appears in the same position, Claude reads from the cache instead of reprocessing it from scratch.

A few mechanics to understand:

When Caching Actually Pays Off

Context caching makes sense when you have a large, stable chunk of context reused across calls:

It does not help much for one-off requests, highly varied prompts, or short system prompts where the static portion is small relative to the per-request content. In those cases the cache write overhead isn't worth it.

A Simplified Example

Here's the general shape of a cached request (field names vary by client, but the pattern is consistent across cache-aware APIs):

{
  "model": "claude-...",
  "system": [
    {
      "type": "text",
      "text": "You are a support agent. Here is the full product manual: ...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [
    { "role": "user", "content": "How do I reset my password?" }
  ]
}

The system block (your long, static manual) is marked for caching. The messages array, which changes every request, is not. On the second call with the same system block, you pay the cache-read rate for that content instead of the full input rate.

How This Shows Up in Your Bill

Token usage responses typically break out:

If you're building on top of Claude through a proxy or gateway, check that the usage metadata actually surfaces this breakdown — otherwise you're flying blind on whether caching is working or just adding write overhead with no reads. If you're running Claude through SubToAPI, usage metadata for each request is visible in the dashboard, which makes it straightforward to confirm your cache hit rate is actually reducing cost rather than just adding overhead. See /docs/messages for request/response shape details.

Practical Tips for Using It Well

  1. Structure prompts with static content first. This is the single most important rule — caching only works on a matching prefix, so variable content at the top breaks it.
  2. Batch related calls together in time. Since caches expire quickly, spreading a job out over hours defeats the purpose. Run batch jobs back-to-back.
  3. Don't cache small content. If your static prefix is a few hundred tokens, the savings aren't meaningful. Caching earns its keep on prompts in the thousands-of-tokens range.
  4. Watch for accidental cache misses. Timestamps, request IDs, or dynamic values injected into the "static" part of your prompt will silently break the match and you'll pay full write cost every time without realizing it.
  5. Combine with streaming for latency wins too. A cache hit reduces time-to-first-token as well as cost, since the model skips reprocessing the cached segment. See /docs/streaming for details on handling streamed responses.

If you're integrating Claude into an app and want a simpler path to get started — generating an application key, hitting a standard HTTPS endpoint, and tracking usage per key or per team seat — /docs/quickstart walks through the setup, and /pricing has the plan breakdown (Solo, Team, Scale) if you're evaluating options beyond a direct account.

Questions

Does context caching reduce latency, not just cost? Yes. A cache hit means Claude doesn't reprocess the cached tokens, which lowers time-to-first-token in addition to lowering the bill. The latency benefit is often as valuable as the cost savings for chat-style apps.

How long does a cache last before it expires? Caches are short-lived — typically on the order of minutes after last use, not hours or days. Design your workflow so related calls happen close together in time to actually benefit from it.

Do I need to change my application logic to use caching? Only the prompt structure matters: put static, reusable content in a fixed position (usually first) and mark it as cacheable, and keep variable per-request content separate. No changes are needed to how you parse or handle the response.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →