← Blog

Claude API Agent Memory Persistence Design Guide

2026-09-28 · 5 min read · SubToAPI Team

The core problem: Claude has no memory between requests

Every call to the Claude API is stateless. The model doesn't remember your last conversation, your user's preferences, or what happened five minutes ago unless you send that information back in the request. Memory persistence design is the engineering work of deciding what to store, where to store it, and how to feed it back into the context window without blowing past token limits or degrading response quality.

This matters most for agents that run for hours or days, handle recurring users, or need to recall facts across sessions — support bots, coding agents, personal assistants, research tools. Get the design wrong and you either lose context (agent "forgets" things it should know) or you burn tokens and latency shipping irrelevant history on every call. This guide covers the patterns that actually work in production.

The three layers of agent memory

Most durable Claude agents end up with a similar architecture, split into three layers:

  1. Working memory — the current conversation turn plus recent messages, sent as-is in the messages array.
  2. Session memory — a summarized or compressed version of the current session, used once the raw transcript gets too long.
  3. Long-term memory — facts, preferences, and decisions extracted from past sessions, stored externally and retrieved on demand.

Trying to solve memory with just one layer is where most implementations go wrong. Raw transcript-only memory scales badly. Summarize-everything memory loses detail. You need all three, with clear rules for when each one is used.

Working memory: keep it simple, cap it

For active conversations, just append messages to the array Claude expects:

[
  {"role": "user", "content": "What's the deploy status?"},
  {"role": "assistant", "content": "The last deploy finished 20 minutes ago..."},
  {"role": "user", "content": "Did it include the auth fix?"}
]

Set a hard cap — by message count or token count — beyond which you stop appending and trigger summarization. A common rule of thumb: once the transcript exceeds roughly 60–70% of your target context budget, compress the older portion.

Session memory: summarize, don't truncate

Truncating old messages silently loses information the agent may need later. A better pattern is to periodically ask Claude to summarize the older part of the conversation into a compact block, then replace those messages with the summary:

const summaryResponse = await client.messages.create({
  model: "claude-opus-4",
  max_tokens: 500,
  messages: [
    ...oldMessages,
    { role: "user", content: "Summarize the key facts, decisions, and open questions from this conversation in under 300 words." }
  ]
});

const summary = summaryResponse.content[0].text;
// Replace oldMessages with a single system-style message containing `summary`

Store both the summary and the original transcript (in cold storage, not sent to the model) so you can regenerate a better summary later if your summarization prompt improves.

Long-term memory: structured, retrievable, not just "more context"

Long-term memory is not "paste the entire history in every time." It's a retrieval problem. Store discrete facts — user preferences, past decisions, project state — as structured records with metadata (timestamp, source conversation, confidence), then retrieve only what's relevant to the current turn.

A minimal schema that works well:

{
  "user_id": "usr_123",
  "fact": "Prefers TypeScript over JavaScript for new projects",
  "source_session": "sess_9f2a",
  "created_at": "2024-11-02T10:15:00Z",
  "tags": ["preference", "coding-style"]
}

You don't need a vector database to start. Keyword or tag-based filtering on a Postgres table gets you far. Add embeddings-based retrieval later if recall quality becomes a real problem — most agents never need it.

Deciding what gets promoted to long-term memory

The trickiest design decision is what counts as worth remembering. A practical rule: extract facts only at clear checkpoints — end of session, explicit user statement ("remember that I prefer..."), or a completed task — rather than trying to mine every message. This keeps extraction cheap and avoids polluting long-term storage with noise from casual conversation.

A simple extraction call:

const extraction = await client.messages.create({
  model: "claude-sonnet-4",
  max_tokens: 300,
  messages: [{
    role: "user",
    content: `From this conversation, list only durable facts about the user or project worth remembering long-term. Return JSON array of strings. Ignore anything transient.\n\n${transcript}`
  }]
});

Run this on session close, not on every turn — it's a background job, not part of the live response path.

Assembling context for a new request

At request time, the pattern is: pull relevant long-term facts, prepend the session summary if one exists, append recent raw messages, then send. Keep the assembled context deterministic and testable — log exactly what was sent so you can debug "why did the agent forget X" issues after the fact.

const context = [
  { role: "user", content: `Known facts about this user:\n${relevantFacts.join("\n")}` },
  { role: "assistant", content: "Understood." },
  ...(sessionSummary ? [{ role: "user", content: sessionSummary }] : []),
  ...recentMessages
];

If you're routing this traffic through an API layer, streaming and tool use both need to keep working through the same memory pipeline — check your provider's docs on streaming and tool use before you bolt memory logic on top, since interrupted streams or mid-tool-call summarization can produce inconsistent state.

Where SubToAPI fits

If you're building this on top of an existing Claude subscription rather than a direct Anthropic API contract, SubToAPI turns that access into a standard HTTPS API with sub_live_... keys, so your memory-management code talks to a normal REST endpoint regardless of how the underlying access is provisioned. Usage metadata per request makes it easier to track token cost by session, which matters once you're running summarization jobs alongside live conversations. See the quickstart and messages endpoint docs for the exact request shape, and pricing if you're evaluating it for a team.

Questions

Does Claude have any built-in memory across API calls? No. The API is fully stateless — every request must include the full context you want the model to consider. Any persistence has to be built on your side using external storage and retrieval.

How much conversation history should I keep in raw form before summarizing? There's no universal number, but a common approach is capping raw history at 60–70% of your intended context budget for that request, then summarizing older turns into a compact block rather than truncating them outright.

Do I need a vector database for agent memory? Not initially. Tag- or keyword-based retrieval from a regular database handles most cases. Add embedding-based semantic search only if you find relevant facts are being missed by simpler filtering.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →