Claude API Chatbot Memory: Persistence & Database Design
Why Claude API chatbots need a memory layer
Claude's API is stateless. Every request you send is evaluated independently — the model has no idea what happened in a previous call unless you include that history in the messages array yourself. If you want a chatbot that remembers a user's name, their previous orders, or a decision made three conversations ago, you have to build that memory yourself using a database. The Claude API gives you the conversation reasoning; your application has to give it the conversation history.
This is the core answer to "how do I add memory persistence to a Claude API chatbot": you store conversation turns (and optionally extracted facts) in a database, retrieve the relevant subset before each API call, and re-inject it into the messages array or system prompt. There's no built-in "memory" feature to toggle on — it's a data engineering problem, not a model setting.
The three layers of chatbot memory
Most production chatbots need memory at three different granularities, and conflating them is the most common mistake.
- Short-term / working memory — the last N turns of the current conversation. This is just the raw message history, stored verbatim.
- Session memory — everything from a single conversation session, even if it's too long to fit in context. Requires summarization or truncation.
- Long-term / user memory — durable facts about a user that persist across sessions ("prefers metric units", "is on the Pro plan", "asked about refunds twice"). This is usually a small structured record, not raw transcript.
A well-designed system stores all three separately because they have different retrieval and update patterns.
Database schema for conversation persistence
A minimal relational schema looks like this:
CREATE TABLE conversations (
id UUID PRIMARY KEY,
user_id UUID NOT NULL,
created_at TIMESTAMPTZ DEFAULT now(),
title TEXT
);
CREATE TABLE messages (
id UUID PRIMARY KEY,
conversation_id UUID REFERENCES conversations(id),
role TEXT CHECK (role IN ('user', 'assistant')),
content TEXT NOT NULL,
created_at TIMESTAMPTZ DEFAULT now()
);
CREATE TABLE user_facts (
id UUID PRIMARY KEY,
user_id UUID NOT NULL,
fact TEXT NOT NULL,
source_conversation_id UUID,
created_at TIMESTAMPTZ DEFAULT now()
);
Postgres is a solid default. If you're using a document store like MongoDB or DynamoDB, keep the same conceptual split: raw messages in one collection, distilled facts in another. The important part isn't the database engine — it's separating raw transcript from extracted memory, because you'll query them differently.
Loading memory back into the request
Before calling the API, assemble context from the database rather than replaying an entire history blindly:
async function buildMessages(conversationId, userId) {
const recentMessages = await db.query(
`SELECT role, content FROM messages
WHERE conversation_id = $1
ORDER BY created_at DESC LIMIT 20`,
[conversationId]
);
const facts = await db.query(
`SELECT fact FROM user_facts WHERE user_id = $1`,
[userId]
);
const systemPrompt = facts.rows.length
? `Known facts about this user:\n${facts.rows.map(f => `- ${f.fact}`).join('\n')}`
: undefined;
return {
system: systemPrompt,
messages: recentMessages.rows.reverse()
};
}
This keeps the request small and cheap while still giving Claude the context it needs. Trying to send every message ever exchanged with a user is both expensive and counterproductive — Claude performs better with relevant context than with maximal context.
Summarization for long conversations
Once a conversation exceeds your working window (say 20–30 turns), don't just truncate — summarize the older portion and store the summary as a single message. A simple pattern:
- When a conversation crosses a turn threshold, take the oldest half of the transcript.
- Send it to Claude with a prompt asking for a compact summary of relevant facts and decisions.
- Store that summary as a synthetic system-level message and delete (or archive) the raw turns from the active context window.
This keeps token usage predictable and avoids the "context window creep" that quietly increases your cost per request over time. If you're building this on top of an application API layer, tracking that cost matters — SubToAPI's dashboard shows per-key usage so you can see exactly which conversations are driving token consumption before it becomes a billing surprise. See /pricing for how the Solo, Team, and Scale tiers scale with usage.
Extracting long-term facts
Long-term memory works best when it's structured, not raw text. After a conversation ends (or periodically during it), run a lightweight extraction pass:
const extraction = await fetch('https://api.subtoapi.app/v1/messages', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'claude-sonnet-4-5',
max_tokens: 300,
messages: [{
role: 'user',
content: `Extract durable facts about the user from this transcript as a JSON array of short strings. Transcript:\n${transcript}`
}]
})
});
Store the returned facts in user_facts, deduplicate against existing entries, and prune anything that becomes stale (e.g., "is evaluating the product" should expire once they've been a paying customer for a month). Memory that never expires turns into noise.
Choosing where to run this
None of the database or memory logic depends on which provider serves the model — it's the same pattern whether you're calling Claude directly or through a gateway. If you're already routing requests through an API layer for key management and usage tracking, the memory-building step just becomes another function that runs before your existing request call. Check /docs/messages for the exact request shape, and /docs/quickstart if you're setting up API access for the first time.
Questions
Does Claude have any native memory feature I can enable instead of building this myself? No. The Claude API is stateless per request — there's no session or memory flag. All persistence has to be implemented in your application and database.
Should I store full conversation transcripts forever? Generally no. Keep recent raw turns for a working window, summarize older turns, and extract only durable facts for long-term storage. Storing everything forever increases both cost and privacy risk without improving chatbot quality.
What's the simplest database to start with for a small chatbot project? Postgres or SQLite with two tables — messages and user_facts — is enough for most early-stage products. Move to a dedicated vector store only if you need semantic search across large amounts of historical conversation.