← Blog

Claude API Knowledge Base Chatbot Setup Guide

2026-10-07 · 5 min read · SubToAPI Team

What "setting up a knowledge base chatbot" actually means

Searching for "claude api knowledge base chatbot setup" usually means one of two things: you want a bot that answers questions from your own docs (not Claude's training data), or you've already built a prototype and need a repeatable process to configure it properly — ingestion, retrieval, prompting, and API access. This guide covers the setup side: the pieces you need to wire together before you write a single line of chatbot logic, and the configuration decisions that determine whether the bot is reliable or hallucinates confidently.

The short version: you need a place to store your knowledge (a vector database or search index), a way to retrieve relevant chunks at query time, a system prompt that tells Claude how to use that retrieved context, and an API setup that handles keys, rate limits, and usage tracking without becoming its own project. We'll walk through each.

Step 1: Decide how you'll store and search your content

Before touching the Claude API, your knowledge base needs to exist somewhere queryable. Three common setups:

For a first setup, pgvector or a managed vector DB with a simple embedding model is usually enough. Don't over-engineer this step — a knowledge base chatbot with 500 FAQ entries doesn't need the same infrastructure as one indexing a million support tickets.

Step 2: Chunk and embed your content correctly

Chunking strategy affects answer quality more than almost anything else in this setup. Guidelines that hold up in practice:

// Simplified ingestion pipeline
for (const doc of documents) {
  const chunks = splitBySection(doc.content, { maxTokens: 400 });
  for (const chunk of chunks) {
    const embedding = await embedText(chunk.text);
    await vectorStore.upsert({
      id: chunk.id,
      vector: embedding,
      metadata: { source: doc.url, title: doc.title, text: chunk.text }
    });
  }
}

Step 3: Build the retrieval step

At query time, embed the user's question, retrieve the top 3–8 matching chunks, and assemble them into context. Retrieve more than you think you need, then let the model's system prompt instruct it to ignore irrelevant chunks — this is more forgiving than retrieving too few.

const queryEmbedding = await embedText(userQuestion);
const matches = await vectorStore.query(queryEmbedding, { topK: 6 });
const context = matches.map(m => `[${m.metadata.title}]\n${m.metadata.text}`).join("\n\n");

Step 4: Write a system prompt that constrains Claude to your knowledge base

This is where most knowledge base chatbots go wrong — they retrieve good context but let Claude answer freely from general knowledge too, which produces mixed answers that look authoritative but aren't grounded. Be explicit:

You are a support assistant that answers only using the provided context.
If the context does not contain the answer, say "I don't have that information
in the knowledge base" instead of guessing. Always cite the source title when
you use information from the context.

Context:
{retrieved_chunks}

Pass the retrieved context as part of the user turn or system prompt, then send the actual question. Details on message formatting are in the docs.

Step 5: Wire up the Claude API connection

Once retrieval and prompting are working, the remaining piece is the API layer itself — authentication, request handling, and tracking what the bot is costing you per conversation. This is where a lot of teams underestimate the work: managing raw API credentials across staging and production, rotating keys when someone leaves the team, and separating usage by project all turn into their own mini-project if you build it from scratch.

SubToAPI handles this layer by converting your existing Claude access into a standard HTTPS API with application-level keys (sub_live_...), so your chatbot code talks to one stable endpoint regardless of how your underlying access changes. A basic request looks like:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 500,
    "system": "You are a support assistant that answers only using the provided context...",
    "messages": [
      { "role": "user", "content": "Context:\n" + context + "\n\nQuestion: " + userQuestion }
    ]
  }'

For chatbots with real-time UI updates, use streaming so answers appear token by token instead of all at once — see the streaming guide. If your knowledge base chatbot also needs to call internal functions (like looking up an order or triggering a search refresh), check tool use for how function calling fits into the same request flow.

Start with the quickstart to get a key and confirm the connection before building retrieval logic on top of it — it's much easier to debug prompting and retrieval issues when you already know the API layer works.

Step 6: Test with questions outside the knowledge base

A properly set up knowledge base chatbot should decline to answer gracefully when the information isn't in the index, rather than falling back to general knowledge. Write a test set of 10–15 questions deliberately outside your knowledge base scope and confirm the bot says so rather than guessing. This single test catches more real-world failures than almost any other validation step.

Keeping the setup maintainable

As your knowledge base grows, revisit chunk sizes and retrieval topK periodically — what worked for 200 documents often needs tuning at 2,000. Log which chunks get retrieved per query so you can spot gaps where users ask about topics with no matching content. And keep an eye on actual API usage per conversation, since retrieval-heavy prompts consume more tokens than simple Q&A; the pricing page and signup flow make it easy to start on a small plan and scale seats as the bot gets used more.

Questions

Do I need a vector database for a small knowledge base chatbot? Not necessarily. If you have under a few hundred documents, full-text search or even keyword matching against a flat file can work fine. Add a vector database once retrieval quality or scale becomes a real problem.

How do I stop Claude from answering outside the knowledge base? Use an explicit system prompt instruction to only use provided context and to say when it doesn't have an answer, and test regularly with out-of-scope questions to confirm the behavior holds.

What's the fastest way to get API access working before building retrieval? Get a key through signup, send a basic test request following the quickstart, and confirm responses work end to end before adding chunking, embeddings, or retrieval logic on top.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →