Claude API Multi-Turn Conversation Design Patterns
Designing multi-turn conversations with the Claude API means deciding how you structure the messages array, who owns conversation state, and how you keep context relevant as a conversation grows. Unlike a single request-response call, a multi-turn design has to persist history somewhere, decide what to send on each new turn, and handle the fact that context windows and costs both grow with every message you keep.
This article walks through the core patterns: message array structure, state storage options, context pruning strategies, and error handling for production conversational apps. By the end you'll have a concrete design you can implement regardless of which backend or framework you're using.
How Claude API Conversations Actually Work
Claude's API is stateless. Every request you send contains the full conversation history as a messages array — there is no server-side session that remembers previous turns for you. Each message has a role (user or assistant) and content. To continue a conversation, you append the new user message to the array you already have and send the whole thing again.
{
"model": "claude-opus-4",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "What's the capital of France?" },
{ "role": "assistant", "content": "The capital of France is Paris." },
{ "role": "user", "content": "What's its population?" }
]
}
This is the single most important fact for multi-turn design: you are the state manager. The API doesn't know "who" is talking or what happened five minutes ago unless you put it in the array. That gives you full control, but it also means a bad design (sending too little or too much history) directly hurts response quality or cost.
Core Design Decisions
1. Where conversation state lives
You need a persistence layer independent of the API call. Common options:
- Database row per conversation — a
conversationstable with a JSON column or a relatedmessagestable, keyed by conversation ID and user ID. This is the right default for most apps. - Client-side storage — acceptable for throwaway demos, but fragile for anything multi-device or needing moderation/audit logs.
- Cache layer (Redis) for active sessions — useful if you want fast reads for in-progress conversations, with periodic flush to a database for durability.
Store messages as they're sent and received, not just the final rendered text. You'll want the raw role/content pairs so you can reconstruct exactly what was sent to the model, including tool calls if you use them.
2. Building the messages array per request
On every new user turn, your server should:
- Load the conversation's message history from storage.
- Append the new user message.
- Apply any pruning or summarization (see below).
- Send the array to the API.
- Append the assistant's reply to storage once you receive it.
async function sendTurn(conversationId, userText) {
const history = await loadMessages(conversationId);
const messages = [...history, { role: "user", content: userText }];
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-opus-4",
max_tokens: 1024,
messages
})
});
const data = await res.json();
await saveMessage(conversationId, "user", userText);
await saveMessage(conversationId, "assistant", data.content[0].text);
return data;
}
Routing requests through a layer like this (whether self-built or via a gateway such as SubToAPI) keeps your application code identical regardless of which model or provider sits behind it, and gives you usage metadata per conversation for free. See /docs/messages for the full request/response shape.
3. Pruning and summarizing context
Context windows are large but not infinite, and every token you send costs money and latency. Three practical strategies:
- Sliding window: keep only the last N turns verbatim. Simple, but loses early context (user's name, stated preferences, earlier decisions).
- Rolling summary: periodically ask Claude to summarize the conversation so far into a short system-style message, then replace older turns with that summary plus the last few raw turns. This preserves intent without the token cost of full history.
- Hybrid with system prompt: store durable facts (user profile, preferences, task state) in the
systemparameter rather than the message array, and only use messages for the actual back-and-forth. This separates "what the model needs to know about the user" from "what was literally said."
A reasonable default: keep the last 10–15 turns verbatim, and once you exceed that, collapse everything older into a single summary message injected near the top of the array.
4. Handling roles correctly
The user and assistant roles must alternate correctly — most API clients will reject a request where two user messages appear back to back without an assistant message between them. If your app needs to inject system-level instructions mid-conversation (e.g., "the user just uploaded a file"), put that in the content of the next user message rather than trying to insert a third role.
5. Error handling and retries
Multi-turn conversations amplify the cost of errors because a failed turn can corrupt your stored history if you're not careful. Practical rules:
- Only persist the assistant's message to storage after a successful response — never write a partial or failed turn.
- On rate limits or transient errors, retry the exact same messages array rather than rebuilding it, to avoid duplicating the user's turn.
- If you support streaming, buffer the full response before saving it to storage, even though you render tokens to the user incrementally. See /docs/streaming for details on consuming streamed responses safely.
If you're building this on top of SubToAPI, requests use your application key (sub_live_...) and the same endpoint whether you're doing single-turn or multi-turn calls — multi-turn behavior is entirely a function of what you send in messages, not a special mode you opt into. Check /docs/quickstart for the minimal setup and /pricing if you're deciding between Solo and Team plans for a multi-user conversational product.
A Minimal Reference Design
For most apps, this combination works well without over-engineering:
- One
conversationstable, onemessagestable (role, content, created_at, conversation_id). - System prompt holds durable user context, refreshed on each request.
- Last 12 turns sent verbatim; older turns collapsed into a summary every time the count exceeds 20.
- Assistant replies written to storage only after a 2xx response.
- Conversation ID passed from client on every turn; server owns history, not the client.
This scales from a side project to a production support bot without redesign — you just add caching or sharding under the same storage model as volume grows.
questions
Do I need to send the entire conversation history on every request? Yes — the API is stateless, so every call must include the full context you want the model to consider, whether that's raw messages or a summarized version of them.
How many turns should I keep before summarizing? There's no universal number, but 10–15 raw turns is a practical starting point before collapsing older history into a summary to control token cost and latency.
Should system instructions go in the messages array or the system prompt? Durable instructions and user context belong in the system parameter; the messages array should represent the actual conversational exchange only.