← Blog

Claude API Memory Management for Chatbots

2026-10-11 · 5 min read · SubToAPI Team

Claude's API is stateless — every request you send is evaluated independently, with no built-in memory of previous turns. If you're building a chatbot, "memory management" means you are responsible for deciding what conversation history to include in each request, how to compress it as it grows, and where to store it between sessions. Get this wrong and you'll hit context window limits, pay for tokens you don't need, or lose important context mid-conversation.

This guide covers the practical techniques for managing memory in a Claude-powered chatbot: sliding windows, summarization, token budgeting, and persistent storage, plus code patterns you can drop into a Node.js or Python backend today.

Why Claude has no built-in memory

The /v1/messages endpoint accepts a messages array and returns a single response. There's no session ID, no server-side conversation state, no "remember this for later." Every bit of context — system prompt, prior turns, tool results — has to be sent again on every call.

This is a deliberate design choice shared across nearly all LLM APIs: it keeps the API simple, cacheable, and horizontally scalable. The tradeoff is that your application layer owns memory entirely. There are two separate problems to solve:

  1. Short-term memory — what goes into the messages array for the current request.
  2. Long-term memory — what gets persisted across sessions, and how you retrieve it when a user comes back.

Short-term memory: managing the context window

Claude's context windows are large, but "large" isn't "infinite," and large contexts cost more per request and add latency. For a chatbot, the naive approach — append every message forever — breaks down in two ways: eventually you exceed the window, and before that, you're paying to re-send the entire chat history on every single turn.

Sliding window

The simplest fix is a fixed-size sliding window: keep the last N turns, drop the rest.

function buildMessages(history, maxTurns = 20) {
  const trimmed = history.slice(-maxTurns);
  return trimmed.map(({ role, content }) => ({ role, content }));
}

This works well for casual chatbots where old context stops mattering (support widgets, FAQ bots). It fails for anything where earlier facts need to persist — a user mentioning their account ID in turn 2 and asking about it in turn 40.

Summarization + recency buffer

A better pattern combines a running summary with a recent-message buffer:

async function summarizeOlderHistory(oldMessages) {
  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "claude-sonnet-4",
      max_tokens: 300,
      messages: [
        {
          role: "user",
          content: `Summarize the key facts, decisions, and open questions from this conversation in under 150 words:\n\n${JSON.stringify(oldMessages)}`
        }
      ]
    })
  });
  const data = await res.json();
  return data.content[0].text;
}

Store the summary alongside the raw history, and regenerate it every time the buffer grows past your threshold (e.g., every 20 new messages). This keeps token usage roughly constant regardless of conversation length, at the cost of one extra API call per compression cycle.

Token budgeting

Decide a hard ceiling for context tokens per request — say 4,000 tokens for history, leaving headroom for the system prompt and response. Track running token counts (Claude's response includes usage metadata) and trigger summarization proactively rather than reactively failing on an oversized request. If you're using SubToAPI, every response includes usage data in the same shape as the underlying Claude API, so you can log and alert on token growth per conversation without extra tooling — see /docs/messages for the response format.

Long-term memory: persisting facts across sessions

Summarization solves memory within a single long conversation. Long-term memory — remembering a user's preferences, past orders, or support history across separate sessions — is a different problem and shouldn't live in the prompt at all.

Pattern: structured memory store

Instead of stuffing facts into chat history, extract them into a structured store (a Postgres table, Redis, or a vector DB for semantic recall) and inject only what's relevant into the system prompt for each session.

const userFacts = await db.getFacts(userId);
// e.g. ["Prefers email over phone", "Enterprise plan, 50 seats", "Reported bug #4821"]

const systemPrompt = `You are a support assistant. Known facts about this user:
${userFacts.map(f => `- ${f}`).join("\n")}`;

This is cheaper and more reliable than re-summarizing old transcripts, and it lets you edit or delete specific facts without reconstructing an entire conversation history.

Pattern: retrieval-augmented recall

For bots that need to reference large histories (months of support tickets, long document trails), use embeddings to retrieve only the relevant snippets per query instead of loading everything. This keeps your context window focused and your costs predictable, and it scales to users with thousands of past interactions where summarization alone would lose too much detail.

Tool use for memory operations

You can also expose memory as a tool Claude calls explicitly — save_fact(key, value) or recall_facts(query) — so the model decides when something is worth persisting rather than your backend guessing. This is more complex to build but gives much better precision for assistants that need to actively manage what they remember. See /docs/tools for how tool definitions and tool_use responses work.

Putting it together in production

A practical chatbot memory stack usually looks like:

If you're prototyping this, start with sliding window + summarization — it covers the majority of chatbot use cases with minimal infrastructure. Add structured profile storage once you need cross-session continuity, and only reach for retrieval/embeddings when conversation volume per user genuinely outgrows what a summary can hold.

If you're routing requests through SubToAPI, the /v1/messages endpoint mirrors the native Claude request/response shape, so these memory patterns drop in unchanged — you get the same usage fields for token tracking and streaming support for incremental responses. Check /docs/quickstart to get an API key running in a few minutes, and /docs/streaming if your chatbot needs token-by-token output while managing history in parallel.

questions

Does Claude remember previous conversations automatically? No. The API is stateless — it has no memory between requests unless your application resends the relevant history or facts in each call.

How many messages should I keep in the context window? There's no universal number; it depends on your token budget and model. A common starting point is 8–15 recent messages verbatim, with older context summarized or stored separately.

What's the difference between short-term and long-term memory for a chatbot? Short-term memory is what you send in the messages array for the current session. Long-term memory is persisted data (user facts, past sessions) stored in your own database and selectively injected into prompts.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →