← Blog

Building a Customer Support Bot with Claude

2026-09-29 · 5 min read · SubToAPI Team

Building a customer support bot with Claude means combining a well-scoped system prompt, access to your knowledge base and tools, and a reliable API layer that handles streaming, retries, and usage tracking. The model itself is the easy part — Claude is already good at understanding intent and generating helpful, on-brand replies. What actually determines whether your bot ships and stays reliable is the plumbing around it: how you feed it context, how you let it call functions like "look up order status," and how you monitor it in production.

This guide walks through the practical steps: designing the system prompt, grounding answers in your own docs, giving Claude tools to take real actions, handling escalation to a human, and wiring it all up to an API.

Step 1: Define the bot's scope with a system prompt

The single biggest failure mode for support bots is scope creep — the bot tries to answer questions it has no business answering, or invents policy details that don't exist. Start with a tight system prompt that defines role, tone, and boundaries explicitly.

You are a support assistant for Acme Cloud Storage.
Only answer questions about Acme's product, billing, and account settings.
If you don't know the answer or it requires account-specific data,
say so and offer to escalate to a human agent.
Never invent pricing, refund policies, or SLA terms — always check
the provided documentation snippets before answering.
Keep responses under 150 words unless the user asks for detail.

This does three things: it constrains the domain, it discourages hallucination by requiring grounding, and it sets response length so replies feel like chat, not essays.

Step 2: Ground answers in your knowledge base

Claude doesn't know your refund policy or your product's edge cases unless you tell it. The standard pattern is retrieval-augmented generation (RAG): search your docs/FAQ/knowledge base for relevant snippets, then inject them into the prompt alongside the user's question.

async function buildPrompt(userQuestion) {
  const snippets = await searchKnowledgeBase(userQuestion, { topK: 3 });
  const context = snippets.map(s => `- ${s.text}`).join("\n");

  return `Relevant documentation:
${context}

Customer question: ${userQuestion}

Answer using only the documentation above. If it doesn't cover
the question, say you're not sure and offer to escalate.`;
}

You don't need a heavyweight vector database to start — a keyword search over your help center articles is often good enough for a first version. Upgrade to embeddings-based search once you have real conversation logs to test against.

Step 3: Give Claude tools for real actions

A support bot that can only talk is limited. The valuable version can look up order status, check subscription tier, or open a ticket. Claude's tool use lets you define functions the model can call, and you execute them server-side.

{
  "name": "get_order_status",
  "description": "Look up the current status of a customer order by order ID",
  "input_schema": {
    "type": "object",
    "properties": {
      "order_id": { "type": "string" }
    },
    "required": ["order_id"]
  }
}

When a user asks "where's my order #4821," Claude returns a tool call instead of a guess. Your backend runs the actual lookup against your order system and feeds the result back into the conversation. This is the difference between a bot that sounds helpful and one that's actually useful. See /docs/tools for the full tool-calling flow if you're implementing this against SubToAPI.

Step 4: Handle escalation gracefully

Every support bot needs an honest exit path. Bake escalation logic into both the prompt and your application layer:

Don't rely on the model alone to decide when to escalate — add a simple keyword or sentiment check as a backstop.

Step 5: Stream responses for a chat-like feel

Support widgets feel sluggish if the user stares at a blank bubble waiting for a full response. Streaming tokens as they're generated makes even a 3-second response feel instant. If you're consuming Claude through SubToAPI, streaming is a standard SSE connection:

curl -N https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-3.5-sonnet",
    "stream": true,
    "max_tokens": 300,
    "system": "You are a support assistant for Acme Cloud Storage.",
    "messages": [{"role": "user", "content": "How do I reset my password?"}]
  }'
}'

Render each chunk to the chat UI as it arrives instead of waiting for the full completion. See /docs/streaming for event format details.

Step 6: Set up the API layer

Once the logic works in a script, you need a production-grade way to call Claude: application-level API keys per environment, usage metadata per conversation (so you can bill or budget by team), and a dashboard to catch spikes before they become surprise invoices.

This is exactly the layer SubToAPI provides on top of your Claude access. You get sub_live_... keys scoped per app or environment, streaming and tool use out of the box, and per-key usage data so you can see which part of your support flow (RAG lookups, escalations, casual Q&A) consumes the most tokens. Setup takes about the same time as reading /docs/quickstart, and plans start at €9/month on the Solo tier, with Team and Scale tiers for multi-agent or multi-app setups. There's a free trial at /signup if you want to wire this into a prototype before committing.

A minimal request against the Messages endpoint looks like this:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-3.5-sonnet",
    "max_tokens": 500,
    "system": "You are a support assistant. Escalate anything about refunds.",
    "messages": [{"role": "user", "content": "Can I get a refund for last month?"}]
  }'

Full request/response shapes are documented at /docs/messages.

Step 7: Test with real transcripts before launch

Before going live, run your bot against 50–100 real historical support tickets and manually grade the outputs. Look specifically for: made-up policy answers, tool calls with wrong parameters, and missed escalation triggers. This catches most production issues cheaper than a support ticket about a wrong refund promise ever will.

questions

Does Claude need fine-tuning to work as a support bot? No. A well-written system prompt combined with retrieved documentation (RAG) and tool access is sufficient for most support use cases. Fine-tuning is rarely necessary and adds ongoing maintenance cost.

How do I stop the bot from making up policy details? Constrain it explicitly in the system prompt to answer only from provided documentation snippets, and instruct it to say "I'm not sure" rather than guess. Ground every factual claim in retrieved context rather than the model's general knowledge.

Can the bot take actions like issuing refunds automatically? Yes, via tool use — define a function like issue_refund with clear input parameters, but keep a human-approval or hard-limit gate on any action with financial impact rather than letting the model execute it unsupervised.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →