Claude API Customer Support Bot Example
Building a customer support bot with the Claude API means sending incoming customer messages to Claude with a system prompt that defines your product, tone, and escalation rules, then returning the response through your chat widget or helpdesk. This article walks through a complete, working example: the system prompt, the request structure, handling multi-turn conversations, and deciding when to hand off to a human.
If you just want to see the shape of a real implementation, skip to the code below. If you're evaluating whether to build this yourself or use a wrapper, we'll also cover what changes when you're running this in production with real traffic.
What a Claude support bot actually needs
A support bot is not just "send message, get reply." A usable implementation needs:
- A system prompt that encodes your product knowledge, tone, and boundaries
- Conversation history passed on every request, since Claude is stateless between calls
- Escalation logic so the bot knows when to stop and hand off to a human
- Structured output when you need the bot to trigger actions (open a ticket, tag a conversation, refund a charge)
- Streaming so replies feel instant in a chat widget instead of appearing all at once
Here's a minimal but realistic example using the Messages API.
The system prompt
The system prompt is where most of the actual "bot design" happens. Be specific about what the bot knows, what it should never promise, and what triggers a handoff.
You are a support assistant for Acme Cloud, a file storage product.
Rules:
- Only answer questions about Acme Cloud features, billing, and troubleshooting.
- Never promise refunds, credits, or discounts — escalate those requests.
- If the user is angry, confused after two attempts, or asks for a human, respond with the exact tag [ESCALATE] followed by a one-sentence summary for the agent.
- Keep answers under 4 sentences unless the user asks for detail.
- If you don't know the answer, say so and escalate rather than guessing.
The request
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 400,
system: SUPPORT_SYSTEM_PROMPT,
messages: [
{ role: "user", content: "My file upload keeps failing at 90%." },
{ role: "assistant", content: "That's usually a timeout on large files. Can you tell me the file size and your connection type?" },
{ role: "user", content: "It's a 4GB video file over wifi." }
]
})
});
const data = await response.json();
console.log(data.content[0].text);
Note that the full conversation history is sent on every call — Claude has no memory of previous requests. Your app is responsible for storing and replaying the thread, usually from your database or session store.
Handling the escalation signal
Since the system prompt instructs Claude to emit [ESCALATE], your backend just needs to check for it:
const text = data.content[0].text;
if (text.startsWith("[ESCALATE]")) {
const summary = text.replace("[ESCALATE]", "").trim();
await createSupportTicket({ summary, transcript: messages });
return "I'm connecting you with a support agent now.";
}
return text;
For stricter structured output — like returning a ticket priority or category alongside the reply — use Claude's tool use feature instead of parsing text tags. That's covered in detail in /docs/tools.
Streaming the reply
For chat widgets, streaming tokens as they arrive matters more than raw latency — users perceive a streaming reply as faster even if total time is similar. The Messages API supports server-sent events for this; see /docs/streaming for the full pattern.
Where this breaks down in production
The example above works for a demo. In production you'll hit a few recurring problems:
- Per-user API keys and rate limiting. If multiple products or teams share one Anthropic key, you have no way to attribute usage or cap spend per customer-facing bot.
- Key rotation and exposure. Support bots often run partly client-side or through third-party helpdesk integrations, which increases the risk of a raw API key leaking.
- Usage visibility. When a support bot's token usage spikes, you want to know which conversation or which day caused it, not just a total bill at month-end.
- Multi-seat access. If your support team and engineering team both need API access for testing and tuning prompts, sharing one secret key is a liability.
This is the gap SubToAPI fills. It sits between your app and Claude, turning your existing Claude access into a standard HTTPS API with application-specific keys (sub_live_...), so your support bot, your internal testing tools, and your other integrations each get their own scoped key instead of one shared secret. You get streaming, tool use, and usage metadata per key, plus team seats so support and engineering can each manage their own access without touching production credentials.
Swapping the example above to run through SubToAPI changes almost nothing in the request shape:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 400,
"system": "You are a support assistant for Acme Cloud...",
"messages": [
{ "role": "user", "content": "My file upload keeps failing at 90%." }
]
}'
You get a dashboard showing which key (which bot, which environment) is consuming tokens, which makes debugging a cost spike or a runaway conversation loop much faster than digging through a single account's combined usage. See /docs/messages for the full request reference and /docs/quickstart to get a key running in a few minutes. Plans start at Solo for €9/month, with Team and Scale tiers for per-seat access — details at /pricing.
Putting it together
A production-ready support bot is a system prompt with clear escalation rules, a conversation history your backend owns, a tool-use or tag-based mechanism for structured actions, and streaming for a responsive UI. The Claude API handles the generation; the surrounding infrastructure — key management, usage tracking, access control — is what determines whether it survives contact with real customers.
questions
Can Claude fully replace a human support agent? No — it works best as a first-line responder that handles common questions and escalates anything involving refunds, account security, or unresolved frustration to a human, using the escalation pattern shown above.
How do I stop the bot from answering off-topic questions? Constrain it explicitly in the system prompt ("only answer questions about X") and test with adversarial prompts; Claude follows scoping instructions well but needs the boundary stated, not implied.
Do I need a separate API key for each bot or environment? It's strongly recommended — using a single shared key across staging, production, and multiple bots makes usage attribution and incident response much harder, which is why SubToAPI issues separate scoped keys per application.