← Blog

Claude API Conversation Memory Implementation

2026-10-07 · 5 min read · SubToAPI Team

The Claude API is stateless. Every request you send is evaluated on its own — Claude does not remember what you said five minutes ago, five requests ago, or even in the previous message, unless you include that history yourself. "Conversation memory" is something you build on top of the API, not a feature the API provides automatically.

This means implementing conversation memory is really three separate engineering problems: storing the conversation somewhere durable, reconstructing the right context for each new request, and managing that context so it doesn't blow past token limits or get expensive. Below is a practical breakdown of how to do each one, with working code.

Why Claude doesn't "remember" by default

Each call to the messages endpoint takes a messages array. Claude generates a reply based only on what's in that array for that single request. If you don't send the prior turns, they don't exist as far as the model is concerned. There's no session ID, no server-side chat history, no "continue where we left off" flag.

This is by design — it keeps the API predictable, cacheable, and stateless on Anthropic's side. The tradeoff is that memory is entirely your responsibility.

Step 1: Store the conversation somewhere

The minimum viable implementation is a table (or even a JSON file for a prototype) that stores, per conversation:

A simple schema in Postgres:

create table messages (
  id bigserial primary key,
  conversation_id uuid not null,
  role text not null check (role in ('user', 'assistant')),
  content text not null,
  created_at timestamptz default now()
);

Every time a user sends a message, you insert it. Every time Claude responds, you insert that too. This is your source of truth — the API itself holds nothing.

Step 2: Reconstruct context on each request

When a new user message comes in, pull the relevant prior turns and rebuild the messages array before calling the API:

async function getConversationHistory(conversationId, limit = 20) {
  const rows = await db.query(
    `select role, content from messages
     where conversation_id = $1
     order by created_at asc
     limit $2`,
    [conversationId, limit]
  );
  return rows.map(r => ({ role: r.role, content: r.content }));
}

async function sendMessage(conversationId, userText) {
  const history = await getConversationHistory(conversationId);
  const messages = [...history, { role: 'user', content: userText }];

  const response = await fetch('https://api.subtoapi.app/v1/messages', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      model: 'claude-sonnet-4',
      max_tokens: 1024,
      messages
    })
  });

  const data = await response.json();
  await saveTurn(conversationId, 'user', userText);
  await saveTurn(conversationId, 'assistant', data.content[0].text);
  return data;
}

That's the core loop: fetch history, append the new message, call the API, persist both sides of the exchange. This pattern works identically whether you call Anthropic directly or go through a wrapper like SubToAPI, which exposes the same messages shape over a standard API key.

Step 3: Manage context growth

Unbounded history eventually hits two walls: the model's context window and your per-request cost. Three strategies handle this, usually combined:

1. Sliding window. Only send the last N turns. Simple, cheap, loses old context.

const recent = history.slice(-20);

2. Summarization. Periodically compress older turns into a short summary and prepend it as a system-style message, then drop the raw turns it replaces.

async function summarizeOldTurns(oldMessages) {
  const summaryPrompt = `Summarize this conversation in 3-4 sentences, keeping key facts, decisions, and names:\n\n${oldMessages.map(m => `${m.role}: ${m.content}`).join('\n')}`;

  const response = await fetch('https://api.subtoapi.app/v1/messages', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      model: 'claude-haiku-4',
      max_tokens: 200,
      messages: [{ role: 'user', content: summaryPrompt }]
    })
  });

  const data = await response.json();
  return data.content[0].text;
}

Store that summary and inject it at the start of the messages array on future calls, replacing the turns it covers.

3. Token budgeting. Track approximate token counts per turn and trim from the oldest end until you're under a target (e.g., 70% of the model's context window, leaving room for the response). This is the most robust approach for long-running conversations and is usually combined with summarization for anything that needs to run for days or weeks.

Multi-session and multi-user memory

If your product has many users each with their own conversation threads, key everything by conversation_id and user_id, and index both. For assistants that need to remember facts across sessions (not just within one thread), add a separate facts or memory table — things like user preferences, names, prior decisions — and inject relevant entries as a short system message rather than replaying entire old conversations. This keeps memory accurate without re-sending megabytes of old chat logs.

Where SubToAPI fits

If you're already managing conversation storage yourself, SubToAPI doesn't change that part — you still own the database and the context-assembly logic described above. What it simplifies is everything around the API call itself: you get a standard sub_live_... API key instead of juggling raw Anthropic credentials, streaming and tool use work the same way as documented and documented, and you get per-key usage metadata so you can see token consumption per conversation or per customer without building that tracking yourself. Plans start at €9/month on pricing, with a free trial at signup.

FAQ

Does the Claude API have a built-in chat history or session feature?

No. The API is stateless — there is no server-side session object. Any memory, whether it's the last message or a week-long conversation, has to be stored and resent by your application on every request.

How much conversation history should I send with each request?

Enough to preserve context without approaching the model's context window limit. A common pattern is the last 10–20 turns plus a rolling summary of everything older, which balances continuity against token cost and latency.

Can I store conversation memory in the browser instead of a database?

For short-lived, single-device sessions, yes — localStorage or sessionStorage works fine for prototypes. For anything multi-device, multi-user, or long-running, you need server-side storage so history survives page reloads and can be reconstructed consistently on each API call.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →