Claude API Chatbot Memory and Context: A Developer Guide
Claude's API doesn't store conversation state between requests — there's no session ID, no server-side memory, no "remembering" anything on Anthropic's end. Every call to the Messages API is stateless: you send the full conversation history as an array of messages, and Claude responds based only on what's in that array. If you want your chatbot to have memory, you build it yourself by managing that array and deciding what context to include on each request.
This is the single most misunderstood part of building chatbots on Claude. Developers expect something like OpenAI's Assistants threads or a built-in memory API, don't find one, and assume they're missing a feature. You're not — you're just responsible for the conversation history, and that responsibility is actually what gives you control over cost, latency, and context quality.
How context actually works in the Messages API
Every request to /v1/messages takes a messages array of alternating user and assistant turns. Claude has no idea what happened in previous requests unless you include it.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-4-20250514",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "My name is Dana and I prefer short answers."},
{"role": "assistant", "content": "Got it, Dana. I will keep answers brief."},
{"role": "user", "content": "What did I just tell you about my preference?"}
]
}'
On your side, you need a store — a database row, a Redis list, a flat file for a prototype — that holds the full exchange for a given user or conversation. On each new user message, you append it, send the whole array (or a trimmed version of it), get the assistant's reply, append that too, and persist the updated array.
Short-term memory: the message history itself
For most chatbots, "memory" just means re-sending prior turns. This works fine until you hit two walls:
- Token limits. Every message you include costs input tokens on every subsequent call, not just once. A 50-turn conversation re-sent in full gets expensive fast and eventually bumps against the model's context window.
- Signal dilution. Long, unfiltered histories bury the information that actually matters under small talk and dead-end tangents, which can degrade response quality.
Common fix: keep a sliding window of the last N turns (say, the last 10–20 messages) and drop older ones. This is simple and works well for task-focused chatbots (support bots, coding assistants) where recent context matters most.
function trimHistory(messages, maxTurns = 20) {
if (messages.length <= maxTurns) return messages;
return messages.slice(-maxTurns);
}
Long-term memory: summarization and extraction
Sliding windows lose information that was important three turns ago but relevant now — a user's stated preference, a decision made earlier, a constraint they mentioned once. For that, you need a second layer:
- Rolling summarization. Periodically (every 10 turns, or when the history exceeds a token budget), ask Claude to compress the older portion of the conversation into a short summary, then replace those raw turns with the summary. Keep the summary as a system-level note or a synthetic early message.
- Fact extraction. Instead of summarizing the whole conversation, extract discrete facts ("user's name is Dana," "user prefers metric units," "user is on the Team plan") into a structured store, and inject only the relevant facts into the system prompt for each new request. This scales better across many sessions and avoids re-feeding irrelevant narrative.
- Retrieval-based memory. For chatbots that need to recall things from weeks or months ago, store past turns in a vector database, embed the new user message, and retrieve the top-k relevant past turns to inject as context. This is effectively RAG applied to conversation history rather than documents.
A practical pattern combines all three: keep a short sliding window of recent raw turns for conversational fluency, a compact rolling summary for mid-term context, and a fact store or retrieval layer for anything that needs to persist indefinitely.
Managing context size in practice
Claude's context window is large, but "fits in the window" and "produces a good response" aren't the same thing. A few habits that help:
- Put stable, slow-changing context (user profile, system instructions, extracted facts) in the system prompt, not buried inside the message array.
- Use the API's usage metadata to track actual token consumption per conversation rather than estimating — this tells you when a conversation is approaching a size where trimming or summarization is overdue.
- Don't summarize every turn; summarize in batches once a threshold is crossed, so you're not spending an extra API call per message.
If you're building this on top of SubToAPI rather than calling Anthropic directly, the mechanics are the same — you still own the messages array and the memory strategy — but you get consistent usage metadata per key and per team seat, which makes it easier to see which conversations are burning the most tokens on history. See /docs/messages for request format and /docs/streaming if your chatbot streams partial responses while you're still assembling long context.
Worked example: summarize-on-threshold
async function maybeSummarize(history, tokenBudget = 6000) {
const estTokens = estimateTokens(history); // your own estimator
if (estTokens < tokenBudget) return history;
const older = history.slice(0, -10);
const recent = history.slice(-10);
const summaryResp = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-opus-4-20250514",
max_tokens: 400,
messages: [
...older,
{ role: "user", content: "Summarize the above conversation in 5 bullet points, keeping any stated facts or preferences." },
],
}),
});
const { content } = await summaryResp.json();
return [
{ role: "user", content: `Conversation summary so far: ${content[0].text}` },
...recent,
];
}
This keeps the live array small while preserving the facts that matter. Check out /docs/quickstart if you're setting up API access for the first time, and /signup for a free trial to test this against real conversations before committing to a plan.
questions
Does Claude remember previous conversations automatically? No. The API is stateless — Claude only sees what you include in the messages array on each request. Any memory has to be implemented in your application.
How much conversation history should I send on each request? Enough to preserve relevant context without exceeding your token budget — typically the last 10–20 turns plus a summary or extracted facts for anything older.
What's the difference between context window size and chatbot memory? The context window is a hard technical limit on how many tokens Claude can process per request. Memory is an application-level design choice about what you store and feed into that window across multiple requests.