← Blog

How to Build an AI Chatbot for Your SaaS Product

2026-09-28 · 5 min read · SubToAPI Team

Building an AI chatbot for a SaaS product means wiring three things together: a language model that generates responses, a backend that manages conversation state and business logic, and a frontend that streams that response into the UI without freezing the page. There's no magic framework you need to install — the architecture is straightforward once you break it into these pieces, and most of the engineering effort goes into context management, not the chatbot itself.

This guide walks through the practical decisions: which model to use, how to structure conversations, how to give the bot access to your product's data, and how to keep costs predictable as usage grows.

Decide what the chatbot actually needs to do

Before writing code, separate "chatbot" into the three things it usually means for a SaaS product:

Each of these has different requirements. A support bot mostly needs retrieval over your docs. A copilot needs tool use so the model can call your product's functions. A data-aware assistant needs both, plus tight scoping so it never leaks one customer's data into another's session.

Core architecture

A minimal chatbot backend has four responsibilities:

  1. Receive the user's message and the conversation history.
  2. Optionally retrieve relevant context (docs, account data, past tickets).
  3. Send the assembled prompt to the model and stream the response back.
  4. Persist the conversation so the next turn has context.
async function handleChatMessage(conversationId, userMessage) {
  const history = await getConversation(conversationId);
  const context = await retrieveRelevantDocs(userMessage);

  const messages = [
    ...history,
    { role: "user", content: `${userMessage}\n\nRelevant context:\n${context}` }
  ];

  const response = await callModel(messages);
  await saveMessage(conversationId, "assistant", response);
  return response;
}

The callModel function is where you talk to your LLM provider. If you're calling Claude directly, that means managing the Anthropic SDK, your own API key, and rate limits. If you'd rather have a single HTTPS endpoint with a scoped key you can hand to different parts of your app (or different customers on a Team plan), SubToAPI turns your Claude access into an API you call with a sub_live_ key — the request/response shape follows the standard Messages format, so swapping it in doesn't change your chatbot logic. See the quickstart for the exact request format.

Streaming responses

Chatbots feel broken if the user stares at a blank screen while the full response generates. Stream tokens as they arrive instead:

const response = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 1024,
    stream: true,
    messages: [{ role: "user", content: userMessage }]
  })
});

const reader = response.body.getReader();
// forward each chunk to the client as it arrives

On the frontend, render each chunk as it comes in rather than waiting for the full response — this is the single biggest UX improvement you can make to a chatbot. Full streaming setup details are in the streaming docs.

Managing conversation context

Every message you send includes the full conversation history, which means your token usage grows with every turn. Two practical rules keep this under control:

If your chatbot needs to look up live data — a user's subscription status, recent orders, ticket history — that's a job for tool use, not for pre-loading everything into the prompt.

Giving the chatbot access to your product

A copilot-style chatbot becomes genuinely useful once it can call functions in your app, not just answer from static knowledge. Define tools that map to real actions:

{
  "name": "get_account_status",
  "description": "Fetch the current user's subscription plan and billing status",
  "input_schema": {
    "type": "object",
    "properties": {
      "user_id": { "type": "string" }
    },
    "required": ["user_id"]
  }
}

When the model decides it needs account data to answer, it returns a tool-use request instead of a text response; your backend executes the function, returns the result, and the model continues. This is what turns a generic chatbot into one that can say "your trial ends in 3 days" instead of "I don't have access to your account." See tool use for the full request/response cycle.

Keeping costs predictable

Chatbot costs scale with conversation length and usage volume, and it's easy to lose visibility once the feature ships. A few habits help:

If you're issuing separate API access per environment (staging vs. production) or per team, scoped keys make this much easier to audit than one shared credential everywhere. SubToAPI's dashboard gives each application its own sub_live_ key with usage tracked per key — useful once your chatbot has multiple environments or customer-facing integrations. Plans start at €9/month on Solo, with Team and Scale tiers for multiple keys and seats — see pricing.

Shipping it

A working v1 doesn't need retrieval, tools, or fine-tuning — start with a system prompt that describes your product, a message history, and streaming output. Add retrieval once you know what questions users actually ask. Add tool use once you know which actions they want the bot to take on their behalf. Building it incrementally, against real usage, beats trying to design the perfect architecture up front.

FAQs

Do I need to fine-tune a model to build a SaaS chatbot? No. Most SaaS chatbots work well with a strong system prompt, relevant context injected per request, and tool use for live data — fine-tuning is rarely necessary and adds ongoing maintenance cost.

How do I stop the chatbot from answering questions outside my product's scope? Constrain it with a clear system prompt defining its role and boundaries, and only provide context relevant to your product. Avoid giving it broad, unscoped access to search the web or answer arbitrary questions unless that's the intended use case.

Should I build my own API wrapper around Claude or use a service? Either works. Writing your own wrapper is fine for a single app with one team. If you need multiple scoped API keys, per-key usage tracking, or want to avoid managing the SDK and streaming plumbing yourself, a hosted layer like SubToAPI (see the docs) saves setup time.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →