Claude API RAG Pipeline From Scratch: Full Guide
A retrieval-augmented generation (RAG) pipeline lets Claude answer questions using your own documents instead of relying only on what it learned during training. You retrieve the most relevant chunks of text from your knowledge base, inject them into the prompt, and let Claude generate an answer grounded in that context. Building one from scratch means you control every step: chunking strategy, embedding model, vector search, and prompt design.
This guide walks through each piece — document chunking, embeddings, vector storage, retrieval, and the final call to the Claude API — with working code you can adapt. By the end you'll have a minimal but functional RAG pipeline you can extend with your own storage backend, reranking, or UI.
Why build RAG instead of relying on a long context window
Claude's context window is large, but RAG still matters for three practical reasons:
- Cost: sending your entire knowledge base on every request wastes tokens. Retrieval sends only what's relevant.
- Freshness: RAG lets you update your knowledge base without retraining or re-uploading a giant prompt each time.
- Accuracy: narrowing the context to relevant chunks reduces the chance Claude gets distracted by irrelevant text.
Step 1: Chunk your documents
Split documents into chunks small enough to embed meaningfully but large enough to preserve context — 300–800 tokens is a common range. Overlap chunks slightly so you don't cut sentences mid-idea.
function chunkText(text, maxChars = 1500, overlap = 200) {
const chunks = [];
let start = 0;
while (start < text.length) {
const end = Math.min(start + maxChars, text.length);
chunks.push(text.slice(start, end));
start += maxChars - overlap;
}
return chunks;
}
For structured documents (Markdown, HTML), chunk by heading or paragraph boundaries instead of a fixed character count — it produces cleaner retrieval results.
Step 2: Generate embeddings
You need a vector representation for each chunk so you can compare it against a query at search time. Use any embedding model you have access to — the pipeline logic is the same regardless of provider.
async function embed(texts, embeddingApiUrl, apiKey) {
const res = await fetch(embeddingApiUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ input: texts }),
});
const data = await res.json();
return data.embeddings;
}
Step 3: Store vectors
For a from-scratch pipeline, you don't need a managed vector database to get started. An in-memory store with cosine similarity works fine for a few thousand chunks and lets you understand the mechanics before adopting something like pgvector, Pinecone, or Qdrant.
function cosineSimilarity(a, b) {
let dot = 0, normA = 0, normB = 0;
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i];
normA += a[i] * a[i];
normB += b[i] * b[i];
}
return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}
function topK(queryVector, store, k = 5) {
return store
.map(item => ({ ...item, score: cosineSimilarity(queryVector, item.vector) }))
.sort((a, b) => b.score - a.score)
.slice(0, k);
}
store here is just an array of { text, vector, metadata } objects built in step 2. At real scale, swap this for a proper vector database — the retrieval logic (embed query, find top matches) stays identical.
Step 4: Retrieve and build the prompt
When a user asks a question, embed the query, retrieve the top-k chunks, and assemble a prompt that clearly separates retrieved context from the question.
function buildPrompt(question, retrievedChunks) {
const context = retrievedChunks
.map((c, i) => `[${i + 1}] ${c.text}`)
.join("\n\n");
return `Answer the question using only the context below. If the context doesn't contain the answer, say you don't know.
Context:
${context}
Question: ${question}`;
}
Keeping the instruction explicit ("use only the context below") reduces hallucination and makes it easier to debug bad answers — you can check whether the retrieved chunks actually contained the information.
Step 5: Call Claude with the retrieved context
Once you have the assembled prompt, send it to Claude. If you're using SubToAPI as your Claude access layer, the call looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [
{ "role": "user", "content": "<your assembled RAG prompt here>" }
]
}'
For longer answers or a chat-style UI, stream the response instead of waiting for the full completion — see /docs/streaming for the event format. Full request and response fields are documented at /docs/messages, and /docs/quickstart walks through getting your first key.
Step 6: Cite sources and handle gaps
Two practical additions make the pipeline usable in production:
- Citations: ask Claude to reference chunk numbers (
[1],[2]) in its answer, then map those back to source documents in your UI. - No-answer handling: if retrieval scores are all low (below a similarity threshold you tune empirically), skip the Claude call entirely and tell the user no relevant information was found. This saves tokens and avoids confident-sounding wrong answers.
const MIN_SCORE = 0.72;
function hasRelevantContext(retrieved) {
return retrieved.some(chunk => chunk.score >= MIN_SCORE);
}
Where this pipeline tends to break
- Chunking too coarse or too fine: test a few chunk sizes against real queries before locking one in.
- No reranking: raw cosine similarity on embeddings is a good first pass but imprecise. A lightweight reranking step (even just asking Claude to score relevance of the top 10 candidates) improves quality a lot.
- Stale indexes: if your source documents change, re-embed and re-index — don't assume the pipeline updates itself.
- Prompt bloat: retrieving too many chunks (k=15+) often hurts more than it helps. Start with k=3–5 and increase only if answers are consistently missing context.
If you're already managing Claude access across a team, routing these RAG calls through a single API key with usage tracking makes it easier to see which queries are expensive and which chunking strategy is actually reducing token spend — check /pricing for plan details or /signup to get a key.
questions
Do I need a vector database to build a RAG pipeline? No. An in-memory array with cosine similarity is enough for prototypes and small knowledge bases. Move to a dedicated vector database once you're indexing tens of thousands of chunks or need persistence across restarts.
How many chunks should I retrieve per query? Start with 3–5. Retrieving too many dilutes relevance and increases token cost; too few risks missing necessary context. Tune based on how your documents are structured.
Can I build RAG without an embedding model? Not effectively — embeddings are what let you find semantically relevant chunks instead of just keyword matches. Keyword search (like BM25) can supplement embeddings but rarely replaces them for natural-language questions.