Claude API Agent Memory Management Pattern
Building an agent with the Claude API means solving a problem the model itself doesn't solve for you: what to do when a conversation outgrows the context window. Claude has no persistent memory between API calls — every request is stateless, and the "memory" your agent appears to have is just the conversation history you resend each time. This article covers the memory management patterns that actually work in production: when to truncate, when to summarize, when to externalize, and how to combine them.
The short answer to "how do I manage agent memory with the Claude API": you pick a strategy based on how long your sessions run. Short-lived tasks (a few turns) can just resend full history. Long-running agents (support bots, coding assistants, multi-day workflows) need a mix of a sliding window for recent turns, periodic summarization for older context, and an external store for facts that must survive indefinitely.
Why memory management matters for Claude agents
Every message you send counts against the context window and against your token bill. An agent that naively appends every turn to the conversation array will eventually hit two walls:
- Context limit: past a certain size, older messages get truncated or the request fails outright.
- Cost and latency: resending a 50,000-token history on every turn is slow and expensive, even with prompt caching reducing the marginal cost.
Good memory management isn't about remembering everything — it's about deciding, deliberately, what the model needs to see on the next turn to act correctly.
Pattern 1: sliding window with hard cutoff
The simplest pattern: keep the last N messages (or last N tokens) and drop the rest.
function slidingWindow(messages, maxMessages = 20) {
if (messages.length <= maxMessages) return messages;
return messages.slice(messages.length - maxMessages);
}
This works well for chat-style agents where recent context matters far more than early turns (customer support, casual assistants). The failure mode is losing important facts stated early — a user's name, an ID, a constraint set in message 3 that's now gone by message 40.
Mitigation: always keep the system prompt and the first user message pinned, and slide only the middle:
function pinnedWindow(messages, maxMessages = 20) {
const [first, ...rest] = messages;
const tail = rest.slice(-1 * (maxMessages - 1));
return [first, ...tail];
}
Pattern 2: rolling summarization
Instead of dropping old messages, compress them. When the history exceeds a threshold, send the oldest chunk back to Claude with a summarization prompt and replace it with the summary.
async function summarizeChunk(messages, client) {
const response = await client.messages.create({
model: "claude-sonnet-4-20250514",
max_tokens: 500,
messages: [
{
role: "user",
content: `Summarize the key facts, decisions, and open questions from this conversation. Be concise and preserve names, IDs, and numbers exactly.\n\n${JSON.stringify(messages)}`
}
]
});
return response.content[0].text;
}
The summary then becomes a single system-level message prepended to the live window:
[SUMMARY OF EARLIER CONVERSATION]
User is debugging a payment webhook (order #48213). Confirmed retries are failing with 401s. Root cause not yet found.
[RECENT MESSAGES]
... last 10 turns in full ...
This is the pattern most production agents converge on. It costs one extra API call per summarization cycle but keeps token usage roughly flat regardless of session length. Trigger it based on token count, not message count, since message length varies a lot with tool calls and code output.
Pattern 3: structured fact extraction
Summaries are lossy in unpredictable ways. For agents that need reliable recall of specific facts (user preferences, entity IDs, task state), extract structured data instead of prose.
Use tool calling to have Claude emit facts as JSON rather than free text:
{
"name": "record_facts",
"input": {
"facts": [
{ "key": "order_id", "value": "48213" },
{ "key": "user_timezone", "value": "CET" }
]
}
}
Store these key-value pairs in your own database and inject only the relevant subset into the next prompt, rather than replaying the whole conversation. This is the most token-efficient pattern and the one that scales best for agents that run for days or weeks, because retrieval is deterministic instead of relying on the model to notice a fact buried in a long summary. See /docs/tools for how tool calling requests and results are structured when you're building this through SubToAPI.
Pattern 4: external memory store (RAG-adjacent, but simpler)
For agents that need to recall arbitrary past interactions — not just fixed fields — store each turn (or each summary chunk) with an embedding or a searchable index, and retrieve only what's relevant to the current turn. This overlaps with retrieval-augmented generation but is narrower in scope: you're searching the agent's own history, not an external knowledge base.
A minimal version doesn't even need a vector database — a simple keyword or recency-weighted lookup over a table of past facts is often enough for internal tools and support agents. Reach for embeddings only once keyword search starts missing relevant context.
Combining patterns in practice
A production agent typically layers these:
- Pinned system prompt — role, constraints, tools available.
- Structured facts store — IDs, preferences, task state, injected as needed.
- Rolling summary — compressed history older than the active window.
- Sliding window — last 10–20 turns in full, for conversational continuity.
This keeps the prompt sent to Claude bounded in size regardless of how long the session runs, while preserving both recent nuance and long-term facts.
If you're exposing this agent as an internal API rather than calling the Claude SDK directly from every service, SubToAPI sits in front of your Claude access and gives each service its own sub_live_... key, so you can build the memory layer once and let multiple internal tools call it through a single, metered endpoint. Check /docs/quickstart to get a key running, and /docs/streaming if your agent streams partial responses while it assembles the next turn's context.
Practical thresholds to start with
- Trigger summarization around 60–70% of your target context budget, not at the hard limit — you need headroom for the response itself.
- Keep sliding windows at 10–20 turns for chat agents, fewer for agents with large tool outputs (code, logs, search results).
- Extract structured facts eagerly — it's cheap to store a fact and expensive to lose one.
- Re-evaluate what's "important" per session type. A coding agent needs file paths and error messages preserved verbatim; a support agent needs customer identifiers and ticket state.
Questions
Does the Claude API have built-in conversation memory? No. Every request is stateless — you must resend the conversation history (or a compressed version of it) yourself on each call. See /docs/messages for how the messages array is structured.
Should I use summarization or a vector store for agent memory? Use summarization for compressing recent conversational flow, and a structured store or vector search for facts and history that need reliable, targeted recall over long time spans. Most agents use both.
How do I know when to trigger summarization? Track token count of the running conversation and trigger a summarization pass at roughly 60–70% of your context budget, leaving room for the model's response and any tool outputs.