Claude API RAG Pipeline Implementation Guide
A RAG (retrieval-augmented generation) pipeline lets Claude answer questions using your own documents instead of relying only on what it learned during training. The core pattern is always the same: split your data into chunks, embed and store those chunks in a vector database, retrieve the most relevant chunks for a given query, and pass them to Claude as context alongside the user's question.
This article walks through a working implementation you can adapt to production: document chunking, embedding generation, vector retrieval, and the prompt structure that gets Claude to answer accurately and cite its sources. It assumes you're calling the Claude API directly or through a proxy like SubToAPI — the retrieval logic is identical either way.
Why RAG instead of fine-tuning or long context
Claude's context window is large, but stuffing entire knowledge bases into every request is slow and expensive. RAG keeps requests small by retrieving only the handful of passages relevant to each query. It also solves the staleness problem: when your underlying documents change, you re-embed and re-index — no retraining, no redeployment of a model.
Fine-tuning changes model behavior; RAG changes what information the model has access to. For most "answer questions about our docs/support tickets/codebase" use cases, RAG is the right tool.
Step 1: Chunk your documents
Chunk size matters more than most people expect. Too large and retrieval returns noisy, unfocused passages. Too small and you lose context needed to answer the question.
A reasonable starting point: 500–800 tokens per chunk with 10–15% overlap between consecutive chunks, split on paragraph or heading boundaries rather than arbitrary character counts.
function chunkText(text, maxTokens = 600, overlapTokens = 80) {
const paragraphs = text.split(/\n\s*\n/);
const chunks = [];
let current = "";
for (const para of paragraphs) {
if ((current + para).length / 4 > maxTokens) {
chunks.push(current.trim());
const overlapChars = overlapTokens * 4;
current = current.slice(-overlapChars) + para;
} else {
current += "\n\n" + para;
}
}
if (current.trim()) chunks.push(current.trim());
return chunks;
}
This is a rough token estimate (4 chars ≈ 1 token), good enough for chunking decisions. For production, use a proper tokenizer if precision matters.
Step 2: Generate embeddings and store them
Claude's API doesn't provide an embeddings endpoint, so you'll pair it with a dedicated embedding model (Voyage AI, OpenAI's embedding models, or an open-source model like bge-large) and a vector store (Pinecone, Qdrant, pgvector, or even an in-memory array for small datasets).
import { VoyageAIClient } from "voyageai";
const voyage = new VoyageAIClient({ apiKey: process.env.VOYAGE_API_KEY });
async function embedChunks(chunks) {
const response = await voyage.embed({
input: chunks,
model: "voyage-3",
});
return response.data.map((d) => d.embedding);
}
Store each embedding alongside the chunk text and any metadata (source document, section title, URL) — you'll need the metadata to build citations later.
Step 3: Retrieve relevant chunks at query time
At query time, embed the user's question with the same model, then run a similarity search (cosine distance is standard) against your vector store and take the top-k results, typically k=4 to k=8.
async function retrieve(query, vectorStore, k = 6) {
const [queryEmbedding] = await embedChunks([query]);
const results = await vectorStore.query({
vector: queryEmbedding,
topK: k,
includeMetadata: true,
});
return results.matches.map((m) => ({
text: m.metadata.text,
source: m.metadata.source,
}));
}
If recall quality is a problem, add a reranking step: retrieve 20–30 candidates with the vector search, then rerank with a cross-encoder or with Claude itself scoring relevance, and keep the top 6.
Step 4: Assemble the prompt and call Claude
Put retrieved chunks in the system prompt or as structured context in the user message, label each chunk with its source, and instruct Claude to only answer from the provided context and to cite sources.
async function askWithContext(query, chunks) {
const context = chunks
.map((c, i) => `[${i + 1}] Source: ${c.source}\n${c.text}`)
.join("\n\n---\n\n");
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
system:
"Answer the user's question using only the context provided below. " +
"Cite sources using [n] notation. If the context doesn't contain " +
"the answer, say so explicitly.\n\n" + context,
messages: [{ role: "user", content: query }],
}),
});
return response.json();
}
Using an explicit "don't guess" instruction materially reduces hallucination on out-of-context questions — Claude follows this reliably when the instruction is unambiguous.
If you're building this through SubToAPI, the request shape matches the standard Messages API — see /docs/messages for the full parameter reference, and /docs/streaming if you want to stream the answer back to users as it's generated, which matters a lot for RAG UIs where the retrieval step already adds latency.
Step 5: Handle grounding and citations
Return the source metadata alongside the response so your UI can render clickable citations. Don't rely on Claude to invent correct citation text — map the [n] markers back to the source list you built during retrieval, server-side.
function renderWithCitations(answerText, chunks) {
return answerText.replace(/\[(\d+)\]/g, (match, n) => {
const chunk = chunks[parseInt(n) - 1];
return chunk ? `[${n}](${chunk.source})` : match;
});
}
Common failure modes
- Chunks too small, losing context: a chunk that says "it increased by 40%" without the surrounding sentence is useless. Favor slightly larger chunks over very small ones.
- No deduplication: overlapping chunks from the same section can dominate the top-k results and crowd out other relevant sources. Deduplicate by source document before sending to Claude.
- Ignoring query rewriting: short or ambiguous user queries embed poorly. Consider having Claude rewrite the query into a more detailed search query before running retrieval — a cheap extra call that improves recall significantly.
- Not monitoring token usage: RAG prompts grow with every chunk you add. If you're routing requests through SubToAPI, usage metadata on each response makes it easy to track per-request cost across the team — useful once you have multiple RAG features sharing one API budget.
Getting started
If you want to prototype quickly without managing separate Anthropic billing and keys, SubToAPI gives you an sub_live_... application key, streaming support, and usage dashboards out of the box — plans start at €9/month with a free trial. Check /docs/quickstart to get a key set up in a few minutes, then plug the askWithContext pattern above into your retrieval layer.
FAQ
Do I need Claude-specific embeddings for RAG? No. Claude's API doesn't generate embeddings — use a dedicated embedding model (Voyage AI, OpenAI, or an open-source alternative) for the retrieval step, and Claude only for generating the final answer from retrieved context.
How many chunks should I retrieve per query? Start with 4–8 chunks (top-k). If answers feel incomplete, increase k or add a reranking step; if answers feel noisy or unfocused, reduce k or shrink chunk size.
Can I stream RAG responses to reduce perceived latency? Yes. Retrieval adds latency before the Claude call starts, but once the request is sent you can stream the completion token by token — see /docs/streaming for implementation details.