← Blog

Claude API Message History Management Tips

2026-10-07 · 5 min read · SubToAPI Team

Claude's API is stateless — every request you send must include the full conversation history in the messages array, because the model has no memory of previous calls. This means message history management is entirely your application's responsibility: how you store it, trim it, and send it back determines both the quality of responses and your token costs.

The core challenge is balancing two things: giving Claude enough context to respond well, and staying under your context window and cost budget as conversations grow. Below are concrete, practical strategies for handling this well in production.

Understand what "history" actually means to the API

Every call to /v1/messages takes a messages array of alternating user and assistant turns. There's no session ID, no server-side thread, no automatic memory. If you want Claude to "remember" turn 3 while answering turn 7, you must include turn 3 in the request payload yourself.

{
  "model": "claude-3-5-sonnet-20241022",
  "max_tokens": 1024,
  "messages": [
    { "role": "user", "content": "What's the capital of France?" },
    { "role": "assistant", "content": "Paris." },
    { "role": "user", "content": "And its population?" }
  ]
}

This design is simple and predictable, but it means your app's message store is the real source of truth. Get the storage and trimming logic wrong and you'll either blow your context budget or lose conversational coherence.

Store raw turns, not just the latest exchange

Keep a durable, append-only log of every user and assistant message per conversation, tagged with timestamps and token counts. Don't overwrite old turns — you'll want them for:

A simple schema works: conversation_id, role, content, created_at, token_count. Compute token counts at write time so you don't have to re-tokenize on every request just to decide what to trim.

Trim with a token budget, not a message count

A common mistake is capping history at "last 20 messages." Message length varies wildly — 20 short messages might be 500 tokens, or 20 long ones might be 15,000. Instead, work backward from a token budget:

  1. Reserve tokens for the system prompt and max_tokens output.
  2. Subtract that from your model's context window to get your available input budget.
  3. Walk the conversation from most recent to oldest, accumulating token counts, and stop including turns once you hit the budget.
function buildHistory(turns, tokenBudget) {
  const included = [];
  let used = 0;
  for (let i = turns.length - 1; i >= 0; i--) {
    const t = turns[i];
    if (used + t.token_count > tokenBudget) break;
    included.unshift(t);
    used += t.token_count;
  }
  return included;
}

This keeps requests predictable regardless of message length and avoids silently truncating mid-conversation.

Summarize instead of dropping old turns entirely

Hard trimming loses information — if a user mentioned their account ID in turn 2 and you drop it by turn 15, Claude will ask again or hallucinate. A better pattern for long-running conversations:

This is more expensive (an extra API call) but much cheaper than sending full history forever, and it preserves continuity far better than pure truncation.

Separate system instructions from conversation history

Don't bury persistent instructions ("respond in JSON," "you are a support agent for X") inside the message array where they compete with trimming logic. Use the dedicated system parameter — it's not part of messages, won't get trimmed by your history logic, and is sent on every request independent of how much conversation history you include.

Cache stable context separately from dynamic turns

If part of your context rarely changes — a product catalog, a long system prompt, reference documentation — keep it structurally separate from the rolling conversation history. Some pipelines store it as a fixed prefix that's reconstructed identically on every call, rather than interleaving it with user turns. This makes your trimming logic simpler because you only need to manage the dynamic part of the conversation.

Track token usage per conversation

The API returns usage metadata (input and output token counts) with every response. Log this per conversation so you can see which conversations are ballooning in size and tune your trimming thresholds based on real data rather than guesses. If you're running this across multiple users or a product with many concurrent conversations, having per-request usage visibility in one place matters — this is one of the things SubToAPI adds on top of raw Claude access: usage metadata surfaced per application key, so you can spot runaway conversations before they become a cost problem. See the messages docs for request/response shape details, or the quickstart to get an API key set up in minutes.

Handle multi-turn tool use carefully

If your conversation includes tool calls, the tool_use and tool_result blocks must stay paired in the correct order within history — breaking that pairing when trimming will cause request errors or confused responses. Treat a tool-call round trip as a single atomic unit when deciding what to trim, never split a tool_use message from its corresponding tool_result. See the tools docs for the exact message structure.

Practical checklist

Questions

Does Claude remember past conversations automatically? No. The API is stateless — you must resend relevant history in the messages array on every request, or the model has no awareness of prior turns.

How many previous messages should I include per request? There's no fixed number — budget by token count relative to your model's context window and reserved output tokens, not by message count, since message lengths vary widely.

What's the cheapest way to preserve long conversation context? Summarize older turns into a condensed message once they exceed a size threshold, and keep only recent turns verbatim — this cuts token usage while preserving continuity.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →