← Blog

Claude API Prompt Caching: Implementation Guide

2026-10-01 · 5 min read · SubToAPI Team

What this guide covers

If you're searching for a Claude API prompt caching implementation guide, you're trying to answer one of two questions: how do I actually add caching to my requests, or why isn't my caching working the way I expected. This guide walks through the mechanics — the cache_control block, where to place it, how TTLs work, and how to verify the cache is actually hitting — with working code.

Prompt caching lets Claude skip reprocessing the parts of your prompt that don't change between requests: system instructions, tool definitions, long reference documents, few-shot examples. You mark a point in your prompt as a cache breakpoint, Claude stores the processed state of everything up to that point, and subsequent requests that match the cached prefix skip straight to the new content. The result is lower latency and lower cost on the cached tokens, but only if you implement the breakpoints correctly.

How Claude's prompt caching actually works

Caching is controlled by adding a cache_control object to a content block in your request. It's not a separate endpoint or a flag on the whole request — it's attached to specific blocks, which is what lets you cache part of a prompt while leaving the rest dynamic.

{
  "model": "claude-opus-4-20250514",
  "system": [
    {
      "type": "text",
      "text": "You are a legal document analysis assistant. Here is the full statute reference: ... (10,000 tokens) ...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [
    { "role": "user", "content": "Does clause 4.2 apply to subcontractors?" }
  ]
}

Everything in the request up to and including the block marked with cache_control becomes the cached prefix. On the next request, if that exact prefix matches (same text, same order, same system prompt, same tool definitions), Claude reuses the cached version instead of reprocessing it.

Two things matter here: exact prefix match and minimum token count. Caching only kicks in above a certain token threshold (it varies slightly by model), so caching a 200-token system prompt won't do much — you need a genuinely large, stable chunk of content.

Implementing cache breakpoints step by step

1. Identify your stable content. This is usually: system prompts, tool schemas, long documents (contracts, codebases, knowledge base chunks), or a fixed set of few-shot examples. If it changes every request, don't cache it.

2. Put stable content first, dynamic content last. Caching works on prefixes, so structure your request so static material comes before the per-request user input.

3. Add cache_control at each breakpoint. You can set up to four cache breakpoints per request (system prompt, tools, and up to two in the message list), letting you cache multiple independent segments — for example, a long document plus a separate set of few-shot examples that change less often than the user's actual query.

const response = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
    "content-type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-opus-4-20250514",
    max_tokens: 1024,
    system: [
      {
        type: "text",
        text: longKnowledgeBaseText,
        cache_control: { type: "ephemeral" }
      }
    ],
    tools: [
      {
        name: "lookup_clause",
        description: "...",
        input_schema: { /* ... */ },
        cache_control: { type: "ephemeral" }
      }
    ],
    messages: [
      { role: "user", content: userQuestion }
    ]
  })
});

4. Check the response usage metadata. The response includes cache_creation_input_tokens and cache_read_input_tokens in the usage object. The first request writes the cache (slightly more expensive than a normal request), and every subsequent matching request within the TTL window shows up as cache_read_input_tokens — billed at a fraction of standard input token cost.

5. Respect the TTL. Cached content expires after a short window (minutes, not hours) from last use. If your traffic pattern has long gaps between requests to the same prompt prefix, the cache will cool down and the next request pays the full write cost again. High-frequency, repeated-prefix workloads (chatbots with a shared system prompt, RAG pipelines reusing the same document) benefit the most.

Verifying caching is actually working

The most common implementation mistake is assuming caching is active because the code runs without errors. Always check the usage object on every response:

const data = await response.json();
console.log(data.usage);
// { input_tokens: 12, cache_creation_input_tokens: 9400, cache_read_input_tokens: 0 }
// next request with the same prefix:
// { input_tokens: 12, cache_creation_input_tokens: 0, cache_read_input_tokens: 9400 }

If cache_read_input_tokens stays at zero across repeated calls, something is breaking the prefix match — usually a timestamp, request ID, or dynamically injected value that got placed before the cache breakpoint instead of after it.

Where this gets harder in production

Implementing caching correctly in a prototype is straightforward. Keeping it correct across a team, multiple services, and changing prompts is where things slip — someone edits the system prompt and breaks the cache silently, or two services send slightly different tool schemas and neither gets cache hits. Tracking cache_read_input_tokens per endpoint over time is the only reliable way to catch this before it shows up as a cost spike.

If you're running Claude behind an internal API for your team, SubToAPI turns your existing Claude access into a standard HTTPS API with per-key usage metadata, so you can see cache hit rates per application key rather than digging through raw logs. It doesn't change how caching itself works — that's still governed by your cache_control blocks exactly as described above — but it gives you one dashboard to confirm it's working across every team member's requests. Setup takes a few minutes; see the quickstart or the messages API reference for request formatting that includes cache blocks.

Common pitfalls

questions

Does prompt caching change Claude's output quality? No. It only affects how the input prompt is processed internally. The model's response is identical whether a prefix is served from cache or processed fresh.

Can I cache the conversation history in a multi-turn chat? Yes, by placing cache_control on the last message block you want cached. New turns get appended after the breakpoint, so the growing history can reuse the cached prefix as long as earlier turns stay unchanged.

How long does a cached prompt stay active? The cache has a short TTL that refreshes on each use, typically a few minutes. It's designed for bursty, repeated-prefix traffic rather than long-lived storage — check the cache_read_input_tokens field to confirm hits are occurring.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →