How to Build an AI Chatbot for Your SaaS Product
Building an AI chatbot for a SaaS product means wiring three things together: a language model that generates responses, a backend that manages conversation state and business logic, and a frontend that streams that response into the UI without freezing the page. There's no magic framework you need to install — the architecture is straightforward once you break it into these pieces, and most of the engineering effort goes into context management, not the chatbot itself.
This guide walks through the practical decisions: which model to use, how to structure conversations, how to give the bot access to your product's data, and how to keep costs predictable as usage grows.
Decide what the chatbot actually needs to do
Before writing code, separate "chatbot" into the three things it usually means for a SaaS product:
- Support assistant — answers questions using your docs/knowledge base, escalates to a human when it can't help.
- In-app copilot — helps users complete tasks inside your product (write a query, generate a report, draft an email).
- Data-aware assistant — needs access to the user's account data (their invoices, their tickets, their metrics) to give useful answers.
Each of these has different requirements. A support bot mostly needs retrieval over your docs. A copilot needs tool use so the model can call your product's functions. A data-aware assistant needs both, plus tight scoping so it never leaks one customer's data into another's session.
Core architecture
A minimal chatbot backend has four responsibilities:
- Receive the user's message and the conversation history.
- Optionally retrieve relevant context (docs, account data, past tickets).
- Send the assembled prompt to the model and stream the response back.
- Persist the conversation so the next turn has context.
async function handleChatMessage(conversationId, userMessage) {
const history = await getConversation(conversationId);
const context = await retrieveRelevantDocs(userMessage);
const messages = [
...history,
{ role: "user", content: `${userMessage}\n\nRelevant context:\n${context}` }
];
const response = await callModel(messages);
await saveMessage(conversationId, "assistant", response);
return response;
}
The callModel function is where you talk to your LLM provider. If you're calling Claude directly, that means managing the Anthropic SDK, your own API key, and rate limits. If you'd rather have a single HTTPS endpoint with a scoped key you can hand to different parts of your app (or different customers on a Team plan), SubToAPI turns your Claude access into an API you call with a sub_live_ key — the request/response shape follows the standard Messages format, so swapping it in doesn't change your chatbot logic. See the quickstart for the exact request format.
Streaming responses
Chatbots feel broken if the user stares at a blank screen while the full response generates. Stream tokens as they arrive instead:
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
stream: true,
messages: [{ role: "user", content: userMessage }]
})
});
const reader = response.body.getReader();
// forward each chunk to the client as it arrives
On the frontend, render each chunk as it comes in rather than waiting for the full response — this is the single biggest UX improvement you can make to a chatbot. Full streaming setup details are in the streaming docs.
Managing conversation context
Every message you send includes the full conversation history, which means your token usage grows with every turn. Two practical rules keep this under control:
- Trim history once a conversation gets long — keep the last N turns plus a short summary of what came before, rather than sending every message verbatim.
- Don't stuff unrelated data into every prompt. If your chatbot has access to account data, retrieve only what's relevant to the current question instead of dumping the user's entire account into context on every turn.
If your chatbot needs to look up live data — a user's subscription status, recent orders, ticket history — that's a job for tool use, not for pre-loading everything into the prompt.
Giving the chatbot access to your product
A copilot-style chatbot becomes genuinely useful once it can call functions in your app, not just answer from static knowledge. Define tools that map to real actions:
{
"name": "get_account_status",
"description": "Fetch the current user's subscription plan and billing status",
"input_schema": {
"type": "object",
"properties": {
"user_id": { "type": "string" }
},
"required": ["user_id"]
}
}
When the model decides it needs account data to answer, it returns a tool-use request instead of a text response; your backend executes the function, returns the result, and the model continues. This is what turns a generic chatbot into one that can say "your trial ends in 3 days" instead of "I don't have access to your account." See tool use for the full request/response cycle.
Keeping costs predictable
Chatbot costs scale with conversation length and usage volume, and it's easy to lose visibility once the feature ships. A few habits help:
- Cap
max_tokenson responses so a runaway generation doesn't burn budget. - Set a hard limit on conversation length before forcing a summary/reset.
- Track token usage per customer if you're billing chatbot access as a feature — usage metadata comes back with every response so you can attribute cost without building your own token counter.
If you're issuing separate API access per environment (staging vs. production) or per team, scoped keys make this much easier to audit than one shared credential everywhere. SubToAPI's dashboard gives each application its own sub_live_ key with usage tracked per key — useful once your chatbot has multiple environments or customer-facing integrations. Plans start at €9/month on Solo, with Team and Scale tiers for multiple keys and seats — see pricing.
Shipping it
A working v1 doesn't need retrieval, tools, or fine-tuning — start with a system prompt that describes your product, a message history, and streaming output. Add retrieval once you know what questions users actually ask. Add tool use once you know which actions they want the bot to take on their behalf. Building it incrementally, against real usage, beats trying to design the perfect architecture up front.
FAQs
Do I need to fine-tune a model to build a SaaS chatbot? No. Most SaaS chatbots work well with a strong system prompt, relevant context injected per request, and tool use for live data — fine-tuning is rarely necessary and adds ongoing maintenance cost.
How do I stop the chatbot from answering questions outside my product's scope? Constrain it with a clear system prompt defining its role and boundaries, and only provide context relevant to your product. Avoid giving it broad, unscoped access to search the web or answer arbitrary questions unless that's the intended use case.
Should I build my own API wrapper around Claude or use a service? Either works. Writing your own wrapper is fine for a single app with one team. If you need multiple scoped API keys, per-key usage tracking, or want to avoid managing the SDK and streaming plumbing yourself, a hosted layer like SubToAPI (see the docs) saves setup time.