Building a Claude API Internal Knowledge Base Assistant
What You're Actually Building
A Claude API internal knowledge base assistant is a tool that lets employees ask natural-language questions about internal documentation — HR policies, engineering runbooks, product specs, support macros — and get accurate answers sourced from your company's own content, not Claude's general training data. The core pattern is retrieval-augmented generation (RAG): you fetch relevant internal documents for a given question, inject them into the prompt, and have Claude answer strictly based on that context.
This is different from a generic chatbot wrapper. The goal is grounded, citable answers that reflect what's actually written in your wiki, Notion, Confluence, or internal docs — and that stop answering when the information isn't there, rather than hallucinating a plausible-sounding guess. Below is a practical architecture you can build this week, plus the decisions that matter most for reliability.
The Core Architecture
A production-grade internal assistant has four pieces:
- Ingestion — pull documents from your sources (Confluence, Notion, Google Drive, S3, Markdown repos) and chunk them into retrievable pieces.
- Embedding + retrieval — convert chunks into vectors, store them in a vector database, and retrieve the top-k relevant chunks per query.
- Prompt assembly — combine the retrieved chunks with the user's question and a system prompt that constrains Claude to the provided context.
- API call + response handling — send the request to Claude, stream the answer back, and optionally return source citations.
1. Chunking and Embedding
Split documents into chunks of roughly 300–800 tokens with some overlap (50–100 tokens) so context isn't cut mid-thought. Store metadata with each chunk: source title, URL, last-updated date, and access permissions if your docs have restricted visibility. Any standard embedding model paired with a vector store (pgvector, Pinecone, Weaviate, or even an in-memory index for a small corpus) works fine here — this part is model-agnostic.
2. Retrieval
At query time, embed the user's question, run a similarity search, and pull the top 3–8 chunks. Re-ranking the results (even a simple keyword overlap boost) noticeably improves answer quality over raw cosine similarity alone, especially for short factual queries like "what's our PTO carryover policy."
3. Prompt Assembly
This is where answer quality is won or lost. A good system prompt for an internal knowledge assistant does three things explicitly: defines the assistant's scope, forces grounding in the provided context, and gives instructions for when information is missing.
You are an internal knowledge assistant for Acme Corp.
Answer the user's question using ONLY the context provided below.
If the answer is not contained in the context, say:
"I couldn't find this in our internal docs — try checking with the team directly."
Always cite the source document title for each claim you make.
Do not use outside knowledge, even if you are confident it's correct.
Keep this prompt stable and version it — small wording changes can shift how aggressively Claude hedges versus answers.
Calling the API
Once you have your context chunks assembled, the actual API call is straightforward. If you're using SubToAPI to turn your existing Claude access into an HTTPS API with a stable key, a basic request looks like this:
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
system: systemPrompt,
messages: [
{
role: "user",
content: `Context:\n${retrievedChunks.join("\n\n")}\n\nQuestion: ${userQuestion}`
}
]
})
});
const data = await response.json();
console.log(data.content);
For an assistant used across a Slack bot or internal web app, streaming matters — users expect a response to start appearing within a second or two, not after a 10-second wait for the full answer. See /docs/streaming for the streaming request format and /docs/messages for the full request schema.
Handling Access Control
Internal knowledge bases often have tiered access — finance docs shouldn't surface to a support intern's query. Handle this at the retrieval layer, not the prompt layer: filter which chunks are eligible for retrieval based on the requesting user's permissions before they ever reach Claude. Never rely on a system prompt instruction like "don't share finance info" as your only safeguard — it's a UX nicety, not an access control mechanism.
Citations and Trust
Employees trust an internal assistant more when they can verify answers. Have Claude return structured output with source references, or post-process the response to match cited document titles back to clickable links. A simple approach: ask Claude to end each answer with a "Sources:" list of document titles from the context you provided, then map those titles to URLs in your application layer.
Keeping It Current
Stale answers are worse than no answers for internal tools — an assistant that confidently cites a deprecated deployment process erodes trust fast. Set up a re-ingestion job (nightly or on-webhook from your docs platform) that re-chunks and re-embeds updated content. Tag each chunk with a last-updated timestamp and consider surfacing it in the response so users can judge freshness themselves.
Where SubToAPI Fits
If your team already has Claude access through a Pro or Team plan, SubToAPI turns that into a proper HTTPS API with application-specific keys (sub_live_...), so your internal assistant doesn't need its own separate Anthropic billing relationship or key management setup. You get streaming, tool use, and usage metadata per key, plus team seats if multiple engineers are building or maintaining the assistant. Check /pricing for plan details or /docs/quickstart to get a key running in a few minutes.
FAQ
Does Claude need to be fine-tuned to answer questions about internal docs? No. Retrieval-augmented generation — injecting relevant document chunks into the prompt at query time — works well for most internal knowledge base use cases and avoids the cost and maintenance burden of fine-tuning.
How do I stop Claude from hallucinating answers not in my docs? Use an explicit system prompt that instructs Claude to answer only from provided context and to say so clearly when the answer isn't present, and test this instruction against edge-case queries before launch.
Can I restrict which employees see which internal content? Yes, but enforce it at the retrieval layer by filtering which document chunks are fetched based on the requesting user's permissions, not by relying on prompt instructions alone.