Build an AI Chatbot with Claude API: Full Guide
Building an AI chatbot with Claude API comes down to four parts: sending structured messages to the API, maintaining conversation history, streaming responses back to your UI, and handling edge cases like rate limits and errors. This guide walks through each part with working code so you can go from a blank project to a functioning chatbot.
Unlike older chatbot frameworks that require training data or intent trees, Claude works off prompts and conversation context. You define a system prompt that sets the assistant's behavior, send a list of messages, and Claude returns a response. The hard part isn't the model call — it's managing state, streaming, and cost across many concurrent users, which is where most of this guide focuses.
Core Chatbot Architecture
A Claude-powered chatbot has three moving pieces:
- Frontend — chat UI that sends user input and renders streamed tokens
- Backend — holds your API key, manages conversation history, calls the API
- Storage — persists conversation history per user or session (database, Redis, or even in-memory for prototypes)
Never call the Claude API directly from the browser. Your API key must stay server-side, both for security and because you'll want to add logging, rate limiting, and cost controls between the user and the model.
Setting Up the API Call
The basic request structure is a system prompt plus a messages array:
const response = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-api-key': process.env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01'
},
body: JSON.stringify({
model: 'claude-sonnet-4-5',
max_tokens: 1024,
system: 'You are a helpful support assistant for Acme Inc. Be concise and friendly.',
messages: [
{ role: 'user', content: 'How do I reset my password?' }
]
})
});
const data = await response.json();
console.log(data.content[0].text);
Each new turn in the conversation appends to the messages array — both the user's message and Claude's previous reply. This is how the model "remembers" context; there's no hidden session on Anthropic's side, so your backend owns the history.
Maintaining Conversation State
A simple in-memory store works for prototypes, but production chatbots need persistence:
function buildMessages(history, newUserMessage) {
return [
...history,
{ role: 'user', content: newUserMessage }
];
}
// After getting a response, append it before saving:
history.push({ role: 'user', content: newUserMessage });
history.push({ role: 'assistant', content: assistantReply });
saveToDatabase(sessionId, history);
Watch your context window. Long-running conversations can grow large, and every call resends the full history. Trim older turns, summarize them, or cap history length once you approach your model's context limit.
Streaming Responses for a Real Chat Feel
Users expect tokens to appear incrementally, not a long pause followed by a wall of text. Set stream: true and read the server-sent events:
const response = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-api-key': process.env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01'
},
body: JSON.stringify({
model: 'claude-sonnet-4-5',
max_tokens: 1024,
stream: true,
messages: [{ role: 'user', content: 'Explain recursion simply.' }]
})
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
// parse SSE lines starting with "data: " and append text deltas to your UI
}
On the frontend, append each text delta to the message bubble as it arrives. This single change makes a chatbot feel dramatically faster even if total latency is the same.
Adding Tool Use for Real Actions
A basic chatbot only talks. A useful one can check order status, look up docs, or trigger a workflow. Claude supports tool calling: you describe available functions, Claude decides when to call them, and you execute the actual logic:
{
"tools": [
{
"name": "lookup_order",
"description": "Get order status by order ID",
"input_schema": {
"type": "object",
"properties": {
"order_id": { "type": "string" }
},
"required": ["order_id"]
}
}
],
"messages": [{ "role": "user", "content": "Where's my order #4521?" }]
}
When Claude responds with a tool_use block, your backend runs the actual lookup, then sends the result back as a tool_result message so Claude can phrase the final answer. This pattern turns a chatbot into an agent that can act, not just chat.
Production Concerns: Keys, Costs, and Teams
Once your chatbot works, production introduces new problems: who holds the API key, how do you track usage per customer, and how do multiple developers test without stepping on each other's rate limits.
This is where SubToAPI fits if you're already paying for Claude access and want a faster path to a production-ready setup. Instead of managing raw Anthropic credentials across environments, SubToAPI gives you application-scoped keys (sub_live_...) with streaming, tool use, and usage metadata built in, plus team seats so your whole engineering team works against one shared, trackable dashboard. Plans start at €9/month for solo developers, with Team (€19/seat) and Scale (€49/seat) tiers for growing products. You can try it with a free trial at /signup.
If you're building the chatbot API layer from scratch, the request and response shapes you'll want to reference are covered in /docs/messages, streaming details are in /docs/streaming, and tool calling is documented at /docs/tools. The /docs/quickstart page walks through the first request end to end.
Testing Before You Ship
Before launching, test for:
- Empty or malformed input — users will send blank messages or paste huge walls of text
- Rate limit handling — add retry with backoff on 429 responses
- Context overflow — trim history gracefully instead of erroring out
- Tone drift — long conversations can wander; periodically reinforce the system prompt
A short staging period with real users (even 10–20) surfaces more issues than synthetic testing ever will.
Questions
Do I need to fine-tune Claude to build a chatbot? No. Almost all chatbot behavior — tone, scope, constraints — is controlled through the system prompt and few-shot examples in your messages, not fine-tuning.
How do I keep conversation history across page reloads? Store the message array in a database or session store keyed by user or session ID, and reload it into the messages array on each new request.
What's the fastest way to add a production API layer without managing raw keys myself? Use a service like SubToAPI, which issues scoped application keys, handles streaming and tool use, and gives you usage visibility across your team — see /pricing for plan details.