← Blog

Claude API Prompt Caching: A Complete Feature Guide

2026-09-28 · 5 min read · SubToAPI Team

Prompt caching is a Claude API feature that lets you store large, reusable chunks of a prompt — system instructions, reference documents, tool definitions, few-shot examples — so subsequent requests don't have to reprocess them from scratch. If your app sends the same long context on every call (a common pattern for RAG apps, coding assistants, and chatbots with big system prompts), caching can cut both latency and cost significantly.

The short answer: you mark a block of your prompt with a cache_control breakpoint, Claude stores the processed state of everything up to that point, and future requests that reuse the same prefix get a much cheaper, faster "cache hit" instead of paying full price to reprocess those tokens. This guide covers how it actually works, what you can and can't cache, pricing mechanics, and a working implementation.

How prompt caching actually works

When you send a request to the Messages API, Claude processes your input tokens before generating a response. Normally, every token in every request is processed fresh, even if 90% of the prompt is identical to the last call. Prompt caching changes that by letting you designate a cache breakpoint — a point in the prompt after which everything before it can be reused.

On the first request with a given prefix, Claude writes that prefix to cache. This "cache write" costs slightly more than a normal input token. On every subsequent request that sends the identical prefix, Claude reads from cache instead of reprocessing it — a "cache hit," which is dramatically cheaper and faster than a full pass.

The cache is scoped to the exact token sequence, so even a single character change before the breakpoint invalidates it for that segment.

What you can cache

Prompt caching works on any content block you place before a cache_control marker, including:

You can set multiple cache breakpoints in a single request (up to four), which is useful when you have both a static system prompt and a semi-static document that changes less often than the live conversation.

Cache lifetime and pricing mechanics

Caches aren't permanent. By default, a cache entry lives for 5 minutes from the last time it was accessed — each hit refreshes the TTL. There's also an extended 1-hour cache option for workloads with sparser traffic, at a higher write cost.

Pricing works like this, relative to standard input token cost:

This means caching pays off when you reuse the same prefix multiple times within the cache window. A single request that never repeats gains nothing from caching — you'd just eat the write premium. But for a chatbot with a 2,000-token system prompt handling dozens of requests per minute, the savings compound fast.

Implementing it

Here's a minimal example using the Messages API directly, caching a long system prompt:

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "system": [
      {
        "type": "text",
        "text": "You are a support agent for Acme Corp. [... several thousand tokens of policy and tone guidelines ...]",
        "cache_control": {"type": "ephemeral"}
      }
    ],
    "messages": [
      {"role": "user", "content": "How do I reset my password?"}
    ]
  }'

The response includes usage fields showing how many tokens were written to cache versus read from it, so you can confirm caching is actually kicking in:

{
  "usage": {
    "input_tokens": 12,
    "cache_creation_input_tokens": 1840,
    "cache_read_input_tokens": 0,
    "output_tokens": 58
  }
}

On the next call with the identical system block, you'd expect cache_read_input_tokens to be populated instead, confirming a hit.

Caching through SubToAPI

If you're accessing Claude through SubToAPI — which turns your existing Claude access into a standard HTTPS API with sub_live_... keys — prompt caching works the same way, since requests pass through to the same Messages format. You still set cache_control breakpoints in your request body; SubToAPI adds usage metadata per key so you can see cache read/write token counts in your dashboard, which is useful for tracking whether your team's caching strategy is actually reducing cost across seats. See the quickstart for setup, or check streaming docs if you're combining caching with streamed responses — caching works fine alongside streaming, it only affects the input side.

Best practices

questions

Does prompt caching change Claude's output quality? No. Caching only affects how input tokens are processed internally — it has no effect on the model's reasoning or the content of its responses.

Is prompt caching worth it for low-volume apps? Usually not. If a given prefix is only sent once, you pay the cache write premium with no read discount to offset it. It pays off when the same prefix is reused multiple times within the TTL window.

Can I combine prompt caching with tool use and streaming? Yes. Caching applies to input tokens regardless of whether you're using tool use or streaming responses — set cache_control on the relevant blocks and the rest of the request works as normal.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →