Build an Internal Knowledge Base with Claude
If you're trying to build an internal knowledge base with Claude, you're almost certainly trying to solve one specific problem: people at your company keep asking the same questions that are already answered somewhere in Notion, Confluence, Slack threads, PDFs, or old tickets — but nobody can find the answer fast enough. Claude is good at synthesizing scattered information into a clear answer, but it only knows what you give it in context. Building a knowledge base means building the pipeline that finds the right documents and feeds them to Claude at query time.
There are two viable approaches, and most real implementations end up using a mix of both. The first is retrieval-augmented generation (RAG): you chunk your documents, embed them, store them in a vector database, and at query time you retrieve the most relevant chunks and pass them to Claude alongside the user's question. The second is long-context stuffing: you rely on Claude's large context window to hold an entire document set (or a curated subset) directly in the prompt, skipping retrieval entirely for smaller corpora. Below is how to decide between them and how to wire up either one.
When RAG makes sense vs. long-context
If your internal docs total under a few hundred thousand tokens — think a single product's documentation, a team's runbooks, or a compact policy handbook — you can often skip a vector database entirely. Just concatenate the relevant files into the system prompt and let Claude search within that. This is simpler to build, easier to debug, and avoids retrieval quality issues where the wrong chunk gets pulled.
Once you're past that size — company-wide wikis, years of support tickets, large codebases — RAG becomes necessary because you can't fit everything in a single request, and you don't want to pay for tokens you don't need on every query.
A reasonable rule of thumb:
- Under ~150K tokens of source material: long-context, no retrieval layer.
- 150K–5M tokens: RAG with a vector store (pgvector, Pinecone, Weaviate, or even a simple SQLite + FAISS setup for smaller teams).
- Multiple sources with different freshness/update cadences: RAG, because re-indexing a changed section is cheaper than re-sending the whole corpus.
The core pipeline
Regardless of scale, the pipeline looks the same:
- Ingest — pull documents from their source (Notion export, Confluence API, Google Drive, markdown files in a repo).
- Chunk — split into 300–800 token pieces, keeping headings and section boundaries intact so chunks are self-contained.
- Embed and store — generate embeddings for each chunk and store them with metadata (source URL, last updated date, team owner).
- Retrieve — on a query, embed the question and fetch the top 5–10 most similar chunks.
- Construct the prompt — inject retrieved chunks into the system prompt with clear delimiters and source attribution.
- Call Claude — send the question with the retrieved context and ask for an answer that cites which document it came from.
Here's a minimal prompt construction step once retrieval has already returned your chunks:
const context = retrievedChunks
.map((c, i) => `[Source ${i + 1}: ${c.title}]\n${c.text}`)
.join('\n\n');
const systemPrompt = `You are an internal knowledge assistant. Answer only using
the sources below. If the answer isn't in the sources, say so explicitly.
Always cite the source number you used.
${context}`;
That system prompt is the single most important part of the whole system. Without an explicit instruction to only use the provided sources and to admit uncertainty, Claude will happily fill gaps with plausible-sounding guesses — fine for a general chatbot, dangerous for an internal source of truth people rely on for compliance or engineering decisions.
Calling the model
Once you have your system prompt and the user's question, the actual API call is straightforward. If you're routing requests through SubToAPI to get a clean HTTPS endpoint with usage metadata and per-team API keys instead of managing raw provider credentials, the call looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"system": "'"$SYSTEM_PROMPT"'",
"messages": [
{"role": "user", "content": "What is our refund policy for enterprise customers?"}
]
}'
This is a good point in the build to think about who inside the company gets to query the knowledge base and how you track usage. If different teams (support, sales, engineering) each need their own access with separate billing visibility, generating a scoped sub_live_... key per team through /signup and checking usage from the dashboard is simpler than building that access control yourself. See /docs/quickstart for the setup and /docs/messages for the full request format.
Keeping it current
A knowledge base that goes stale is worse than no knowledge base, because people trust it and get wrong answers. Two practical habits:
- Re-embed on change, not on schedule. Hook into your source system's webhook (Notion, Confluence, Git) so a document edit triggers re-chunking and re-embedding of just that document, not a full nightly rebuild.
- Attach a "last updated" date to every chunk and instruct Claude to flag when the most relevant source is more than N months old. This catches the common failure mode where an outdated doc scores highest on similarity but is no longer accurate.
Streaming for a better UX
If you're building a chat-style interface on top of this, stream the response so users see the answer forming rather than waiting on a full retrieval-plus-generation round trip. See /docs/streaming for the event format if you're using SubToAPI as the transport layer — it's the same server-sent-events pattern as calling Claude directly, just behind your application key.
Tool use for live lookups
Sometimes the knowledge base needs to answer questions that a static document snapshot can't — current on-call rotation, open ticket counts, live inventory. For those, give Claude a tool definition that calls your internal API directly instead of relying purely on retrieved text. /docs/tools covers the request/response shape for defining and handling tool calls.
questions
Do I need a vector database to build an internal knowledge base with Claude? Not always. If your total document set is small enough to fit in a single context window, you can skip retrieval and pass the documents directly in the prompt. A vector database becomes necessary once your corpus is too large to include in full on every request.
How do I stop Claude from making up answers not in the source documents? Be explicit in the system prompt: instruct it to answer only from the provided sources, cite which source it used, and say "I don't know" when the answer isn't present. This single instruction eliminates most hallucination in RAG setups.
How often should I re-index the knowledge base? Trigger re-indexing on document change via webhooks rather than a fixed schedule. This keeps answers current without wastefully re-embedding unchanged content, and it avoids the lag of a nightly batch job.