Claude API Agent Memory and State Management
The problem: Claude has no memory between calls
Every call to the Claude API is stateless. The model doesn't remember your last request, your user's name, or what tool it called five minutes ago unless you send that information back in the next request. If you're building an agent — something that holds a multi-turn conversation, tracks progress on a task, or calls tools across several steps — you are responsible for memory and state, not the model.
This is the core thing to understand before writing any agent code: "memory" in a Claude-based system is really just "what you include in the messages array of your next request." State management is the engineering problem of deciding what to include, how to store it, and how to keep it small enough to fit in the context window while still being useful.
What "state" actually means for a Claude agent
A Claude agent typically needs to track several distinct kinds of state:
- Conversation history — the back-and-forth turns between user and assistant
- Tool call state — which tools were called, with what arguments, and what they returned
- Task/workflow state — where the agent is in a multi-step process (e.g., "step 3 of 5: waiting for user confirmation")
- Long-term user memory — facts that should persist across sessions, not just within one conversation
- Scratch/working memory — intermediate reasoning or data the agent needs temporarily but doesn't need to show the user
Each of these has a different lifetime and a different storage strategy. Treating them all the same (just appending everything to one growing message list) is the most common mistake.
Pattern 1: Full conversation history (simple, limited)
The simplest approach: store every message in a database row or session object, and replay the full list on each request.
const messages = await db.getMessages(sessionId);
messages.push({ role: "user", content: userInput });
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages
})
});
const data = await response.json();
await db.appendMessage(sessionId, { role: "assistant", content: data.content });
This works fine for short-lived conversations but breaks down as history grows: you'll eventually hit context window limits, pay for tokens you don't need, and risk the model losing track of what matters amid noise. See /docs/messages for the exact request/response shape when building this with SubToAPI.
Pattern 2: Rolling summarization
Once a conversation passes a token threshold, summarize older turns into a compact block and drop the raw messages. A common structure:
const systemPrompt = `
Conversation summary so far:
${storedSummary}
Continue the conversation naturally using this context.
`;
const messages = [
...recentRawMessages, // last N turns, unsummarized
{ role: "user", content: userInput }
];
Every few turns, send the full history to Claude with an instruction like "summarize this conversation in under 200 words, preserving names, decisions, and open questions," store the result, and reset the raw message buffer. This keeps context size bounded regardless of conversation length, at the cost of some fidelity loss.
Pattern 3: External state store for structured data
For anything that isn't conversational prose — a task list, a form being filled out, inventory counts, user preferences — don't rely on the model to "remember" it in natural language. Store it as structured data (JSON in a database, Redis, etc.) and inject only the relevant fields into the system prompt on each call:
{
"task_id": "order-4471",
"status": "awaiting_payment_confirmation",
"collected_fields": {
"shipping_address": "...",
"payment_method": "card"
}
}
const state = await db.getTaskState(taskId);
const systemPrompt = `Current task state: ${JSON.stringify(state)}
Update the task state based on the conversation and respond accordingly.`;
This is far more reliable than letting the model reconstruct structured facts from scrollback. If your agent uses tool calling to update this state, SubToAPI's tool use support (/docs/tools) lets you define the state-mutating functions directly and get structured arguments back instead of parsing prose.
Pattern 4: Long-term memory across sessions
If an agent needs to remember facts about a user across separate sessions (not just within one chat), conversation history isn't the right tool at all — it's a retrieval problem. The common pattern:
- Extract durable facts from conversations (e.g., "user prefers email over SMS") using a Claude call dedicated to extraction
- Store them in a database, keyed by user ID, optionally with embeddings for semantic search
- On each new session, retrieve relevant facts and inject them into the system prompt — not the full history
This keeps long-term memory cheap and precise instead of replaying years of chat logs.
Managing state with streaming agents
Streaming adds a wrinkle: you need to capture the full assistant response as it streams in order to persist it as state, while also forwarding chunks to the client in real time. Buffer the stream server-side, append to your message store once it completes, and only then forward the "done" signal — see /docs/streaming for handling streamed responses correctly so you don't lose partial state on disconnects.
Practical checklist
- Cap raw conversation history by token count, not message count
- Summarize proactively before you hit context limits, not reactively after errors
- Separate structured task state from conversational memory — never mix them in one blob
- Use tool calls to read/write state explicitly rather than hoping the model infers it
- Store long-term memory outside the message array entirely
If you're prototyping this and don't want to manage separate API keys, rate limits, and billing per environment, SubToAPI gives you a single HTTPS endpoint with per-key usage metadata so you can see exactly how much context each session is consuming — useful for tuning your summarization thresholds. Check /docs/quickstart to get a key running in minutes, or /pricing if you're scaling an agent across a team.
Questions
Does Claude remember previous conversations automatically? No. Claude is stateless between API calls. Any memory of prior turns has to be sent explicitly in the messages array of each new request — the model has no built-in persistence.
How do I stop my agent's context window from filling up? Cap raw history by token count, periodically summarize older turns into a compact block, and move structured data (tasks, preferences) out of the chat history into a separate state store you inject only when relevant.
What's the best way to store long-term memory across sessions? Extract durable facts from conversations with a dedicated extraction step, store them in a database (optionally with embeddings for retrieval), and inject only the relevant facts into the system prompt for new sessions — don't replay full chat history.