Claude API Retrieval Augmented Generation Setup Guide
Setting up retrieval augmented generation (RAG) with the Claude API means connecting three pieces: a document store you control, a retrieval step that finds relevant chunks, and a prompt that hands those chunks to Claude alongside the user's question. Claude itself doesn't have a "RAG mode" — you build the pipeline around it, and Claude's job is just to reason well over the context you provide.
This guide walks through a working setup: chunking your documents, generating embeddings, running similarity search, and constructing the final Messages API call so Claude answers grounded in your data instead of hallucinating from training knowledge.
Why RAG instead of a bigger context window
Claude's context window is large enough to paste entire documents directly into a prompt, and for small, static knowledge bases that's a legitimate shortcut — no vector database required. But RAG earns its complexity when:
- Your knowledge base is too large to fit in a single context window (thousands of documents, changing daily)
- You need per-request cost control — retrieving 5 relevant chunks is cheaper than sending 50 pages every time
- Different users should see different subsets of data (multi-tenant knowledge bases)
- You want citations — knowing which chunk an answer came from
If none of that applies, just paste the text in. If it does, here's the setup.
Step 1: Chunk your documents
Split source documents into chunks of 300–800 tokens with some overlap (50–100 tokens) so context isn't cut mid-idea. Chunk by semantic boundaries (headings, paragraphs) rather than fixed character counts when possible.
function chunkText(text, maxTokens = 500, overlap = 80) {
const words = text.split(/\s+/);
const wordsPerChunk = maxTokens * 0.75; // rough token-to-word ratio
const chunks = [];
let i = 0;
while (i < words.length) {
chunks.push(words.slice(i, i + wordsPerChunk).join(" "));
i += wordsPerChunk - overlap;
}
return chunks;
}
Store each chunk with metadata: source document, section title, and a stable ID you can use for citations later.
Step 2: Generate embeddings
Claude doesn't provide an embeddings endpoint, so pair it with a dedicated embedding model (Voyage AI, OpenAI's embedding models, or an open-source model like bge-large running locally). Voyage AI is a common pairing with Claude since it's tuned for retrieval quality.
async function embed(texts) {
const res = await fetch("https://api.voyageai.com/v1/embeddings", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.VOYAGE_API_KEY}`,
},
body: JSON.stringify({ input: texts, model: "voyage-2" }),
});
const data = await res.json();
return data.data.map((d) => d.embedding);
}
Store embeddings in a vector database — Pinecone, Weaviate, pgvector, or even an in-memory array for a prototype with a few thousand chunks.
Step 3: Retrieve relevant chunks at query time
Embed the incoming user question with the same model, then run a similarity search (cosine similarity is standard) against your stored vectors. Pull the top 3–8 chunks depending on how dense your source material is.
function cosineSimilarity(a, b) {
const dot = a.reduce((sum, val, i) => sum + val * b[i], 0);
const magA = Math.sqrt(a.reduce((s, v) => s + v * v, 0));
const magB = Math.sqrt(b.reduce((s, v) => s + v * v, 0));
return dot / (magA * magB);
}
function topK(queryVector, storedChunks, k = 5) {
return storedChunks
.map((c) => ({ ...c, score: cosineSimilarity(queryVector, c.embedding) }))
.sort((a, b) => b.score - a.score)
.slice(0, k);
}
If you're using a managed vector database, this step is a single API call instead of a manual loop, but the logic is identical.
Step 4: Build the prompt for Claude
Structure the retrieved chunks clearly, labeling sources so Claude can cite them and so you can verify grounding. A simple, effective pattern is to inject retrieved context as a system message or as the first block of the user turn, wrapped in XML-style tags Claude parses reliably:
const context = retrievedChunks
.map((c, i) => `<document index="${i}" source="${c.source}">\n${c.text}\n</document>`)
.join("\n\n");
const systemPrompt = `You are a support assistant. Answer only using the documents provided below.
If the answer isn't in the documents, say you don't have that information.
Cite the source document for each claim.
${context}`;
Then send it through the Messages API:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"system": "'"$systemPrompt"'",
"messages": [{"role": "user", "content": "What is our refund policy for annual plans?"}]
}'
The instruction to answer "only using the documents provided" and to say when it doesn't know is what actually prevents hallucination — Claude follows explicit grounding instructions well, but it needs them stated plainly.
Step 5: Handle no-match cases
Set a similarity threshold. If the top retrieved chunk scores below it (say, 0.7 cosine similarity, tune per embedding model), skip the RAG context entirely and tell Claude explicitly that no relevant documents were found, or return a fallback response without calling Claude at all. This avoids Claude confidently answering from irrelevant chunks that happened to rank highest.
Where SubToAPI fits
If you're building this pipeline for an internal tool or a customer-facing app and you're already using Claude through a personal or team subscription rather than a metered Anthropic API key, SubToAPI turns that access into a standard HTTPS API you can call from your RAG backend the same way you'd call any LLM provider — with application-scoped keys (sub_live_...), streaming support, and usage visibility per key. That's useful when multiple services or team members need to hit Claude for retrieval-augmented calls without sharing one raw credential. Check the quickstart or the Messages endpoint docs to see the request format, which mirrors the standard Messages API shown above.
Testing your setup
Before shipping, run a small eval set: 15–20 real questions with known correct answers, run them through the pipeline, and manually check whether retrieval pulled the right chunks and whether Claude's answer stayed grounded. Most RAG quality problems trace back to bad chunking or retrieval, not the model — if answers are wrong, check what was actually retrieved before tuning the prompt.
FAQ
Does Claude have a built-in RAG feature? No. Claude provides the Messages API for generation; you're responsible for chunking, embedding, and retrieval using a separate vector store and embedding model.
Which embedding model works best with Claude? Voyage AI's models are commonly paired with Claude and tuned for retrieval tasks, but any embedding model (OpenAI, Cohere, open-source) works — consistency between the model used for documents and queries matters more than the specific choice.
How many chunks should I retrieve per query? Start with 3–5 chunks of 300–800 tokens each. Retrieve more only if answers are consistently missing context, and watch total prompt size since irrelevant chunks add noise and cost.