Claude API Knowledge Base Chatbot: Build Guide
Building a knowledge base chatbot with the Claude API means combining a retrieval step (finding the right documents) with a generation step (Claude writing an answer grounded in those documents). This pattern is called retrieval-augmented generation (RAG), and it's the standard way to make Claude answer questions about your product docs, internal wiki, support tickets, or any private content it wasn't trained on.
The short version: you split your documents into chunks, embed and store them in a vector database, retrieve the most relevant chunks for each user question, inject those chunks into Claude's system prompt or user message, and let Claude generate a grounded answer. Below is a full walkthrough of each piece, plus a working code example.
Why RAG instead of fine-tuning
Fine-tuning a model on your knowledge base is slow, expensive, and goes stale the moment your docs change. RAG keeps the model frozen and swaps out the context at query time, so updating your knowledge base is just re-indexing documents — no retraining required. It also lets you cite sources, which matters for support bots and internal tools where users need to trust the answer.
Step 1: Chunk your documents
Split long documents into smaller pieces (300–800 tokens works well for most knowledge bases). Chunk by semantic boundaries — headings, paragraphs, or sections — rather than fixed character counts, so each chunk stays coherent on its own.
function chunkByHeading(markdown) {
return markdown
.split(/\n##\s/)
.map((section) => section.trim())
.filter((section) => section.length > 50);
}
Step 2: Embed and store chunks
Generate an embedding for each chunk and store it in a vector database (Pinecone, Weaviate, pgvector, or even a flat array for small knowledge bases). At query time, embed the user's question with the same model and run a similarity search to pull back the top 3–5 matching chunks.
Step 3: Build the prompt
This is the part that determines whether your chatbot actually sounds useful or just hallucinates confidently. Put retrieved chunks in the system prompt, instruct Claude to only answer from the provided context, and tell it what to say when the answer isn't there.
const systemPrompt = `You are a support assistant for Acme Docs.
Answer only using the context below. If the answer isn't in the
context, say you don't know and suggest contacting support.
Context:
${retrievedChunks.join("\n\n---\n\n")}`;
Step 4: Call the Claude API
Once you have your system prompt and the user's question, send a request to Claude's messages endpoint. If you're using SubToAPI, the request shape is the same /v1/messages format described in the docs, authenticated with your sub_live_ key instead of juggling raw Anthropic credentials and billing:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"system": "'"$SYSTEM_PROMPT"'",
"messages": [
{"role": "user", "content": "How do I reset my API key?"}
]
}'
For a chatbot UI, you'll almost always want streaming so answers appear token by token instead of waiting for the full response — see streaming for the implementation details.
Step 5: Handle multi-turn conversations
A knowledge base chatbot rarely gets one-shot questions — users ask follow-ups. Keep the conversation history in the messages array, but re-run retrieval on each new user turn rather than reusing the first query's chunks. If a follow-up is "what about on mobile?", retrieval based only on the latest message will miss context, so it's often worth concatenating the last two or three user turns before embedding the query.
const recentTurns = conversation
.filter((m) => m.role === "user")
.slice(-2)
.map((m) => m.content)
.join(" ");
const queryEmbedding = await embed(recentTurns);
Step 6: Consider letting Claude call the retriever directly
Instead of always retrieving before every call, you can give Claude a tool that searches your knowledge base on demand. This is useful when questions vary wildly in scope — some need no context at all, others need multiple searches. Claude decides when to call the tool and with what query, which cuts down on unnecessary retrieval calls for simple questions. The tool use docs cover the request/response format for defining a search_knowledge_base tool and handling the tool_use block Claude returns.
Production considerations
A few things that separate a working prototype from a chatbot people trust:
- Citations: ask Claude to reference which chunk or document it used, and show that to the user as a link.
- Context window budget: don't dump 20 chunks into every prompt. Rank and trim to the top few that actually matched well, or you'll push out room for conversation history.
- Fallback behavior: explicitly instruct the model what to do when retrieval comes back empty — a generic "I don't know" is better than a confident guess.
- Caching: if your knowledge base is mostly static, cache embeddings and even full answers for common questions to cut latency and cost.
- Metering: if you're exposing this chatbot to a team or customers, you'll want per-request usage visibility. SubToAPI tracks usage metadata per API key automatically, which is handy when multiple apps or teams share one Claude setup — see pricing for the Solo, Team, and Scale tiers.
Getting started quickly
If you want to skip the credential and billing setup and just get an API key that works with the code above, sign up for a free trial and follow the quickstart — it walks through generating a sub_live_ key and making your first request in a few minutes.
Questions
Do I need a vector database to build a knowledge base chatbot with Claude? For small knowledge bases (a few hundred chunks), a simple in-memory array with cosine similarity works fine. Once you're past a few thousand chunks, a dedicated vector database like Pinecone or pgvector becomes worth the setup for speed and filtering.
How do I stop Claude from making up answers not in my docs? Be explicit in the system prompt that it should only use the provided context and state clearly what to say when the answer isn't present. Testing with edge-case questions outside your knowledge base is the best way to catch hallucination before launch.
Can I use Claude's tool use feature instead of manual retrieval? Yes — define a search function as a tool and let Claude decide when to call it and with what query. This works well when question complexity varies, since Claude can skip retrieval entirely for simple questions and search multiple times for complex ones.