← Blog

Claude API RAG Pipeline From Scratch: Full Guide

2026-10-04 · 5 min read · SubToAPI Team

A retrieval-augmented generation (RAG) pipeline lets Claude answer questions using your own documents instead of relying only on what it learned during training. You retrieve the most relevant chunks of text from your knowledge base, inject them into the prompt, and let Claude generate an answer grounded in that context. Building one from scratch means you control every step: chunking strategy, embedding model, vector search, and prompt design.

This guide walks through each piece — document chunking, embeddings, vector storage, retrieval, and the final call to the Claude API — with working code you can adapt. By the end you'll have a minimal but functional RAG pipeline you can extend with your own storage backend, reranking, or UI.

Why build RAG instead of relying on a long context window

Claude's context window is large, but RAG still matters for three practical reasons:

Step 1: Chunk your documents

Split documents into chunks small enough to embed meaningfully but large enough to preserve context — 300–800 tokens is a common range. Overlap chunks slightly so you don't cut sentences mid-idea.

function chunkText(text, maxChars = 1500, overlap = 200) {
  const chunks = [];
  let start = 0;
  while (start < text.length) {
    const end = Math.min(start + maxChars, text.length);
    chunks.push(text.slice(start, end));
    start += maxChars - overlap;
  }
  return chunks;
}

For structured documents (Markdown, HTML), chunk by heading or paragraph boundaries instead of a fixed character count — it produces cleaner retrieval results.

Step 2: Generate embeddings

You need a vector representation for each chunk so you can compare it against a query at search time. Use any embedding model you have access to — the pipeline logic is the same regardless of provider.

async function embed(texts, embeddingApiUrl, apiKey) {
  const res = await fetch(embeddingApiUrl, {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ input: texts }),
  });
  const data = await res.json();
  return data.embeddings;
}

Step 3: Store vectors

For a from-scratch pipeline, you don't need a managed vector database to get started. An in-memory store with cosine similarity works fine for a few thousand chunks and lets you understand the mechanics before adopting something like pgvector, Pinecone, or Qdrant.

function cosineSimilarity(a, b) {
  let dot = 0, normA = 0, normB = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    normA += a[i] * a[i];
    normB += b[i] * b[i];
  }
  return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}

function topK(queryVector, store, k = 5) {
  return store
    .map(item => ({ ...item, score: cosineSimilarity(queryVector, item.vector) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, k);
}

store here is just an array of { text, vector, metadata } objects built in step 2. At real scale, swap this for a proper vector database — the retrieval logic (embed query, find top matches) stays identical.

Step 4: Retrieve and build the prompt

When a user asks a question, embed the query, retrieve the top-k chunks, and assemble a prompt that clearly separates retrieved context from the question.

function buildPrompt(question, retrievedChunks) {
  const context = retrievedChunks
    .map((c, i) => `[${i + 1}] ${c.text}`)
    .join("\n\n");

  return `Answer the question using only the context below. If the context doesn't contain the answer, say you don't know.

Context:
${context}

Question: ${question}`;
}

Keeping the instruction explicit ("use only the context below") reduces hallucination and makes it easier to debug bad answers — you can check whether the retrieved chunks actually contained the information.

Step 5: Call Claude with the retrieved context

Once you have the assembled prompt, send it to Claude. If you're using SubToAPI as your Claude access layer, the call looks like this:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [
      { "role": "user", "content": "<your assembled RAG prompt here>" }
    ]
  }'

For longer answers or a chat-style UI, stream the response instead of waiting for the full completion — see /docs/streaming for the event format. Full request and response fields are documented at /docs/messages, and /docs/quickstart walks through getting your first key.

Step 6: Cite sources and handle gaps

Two practical additions make the pipeline usable in production:

const MIN_SCORE = 0.72;

function hasRelevantContext(retrieved) {
  return retrieved.some(chunk => chunk.score >= MIN_SCORE);
}

Where this pipeline tends to break

If you're already managing Claude access across a team, routing these RAG calls through a single API key with usage tracking makes it easier to see which queries are expensive and which chunking strategy is actually reducing token spend — check /pricing for plan details or /signup to get a key.

questions

Do I need a vector database to build a RAG pipeline? No. An in-memory array with cosine similarity is enough for prototypes and small knowledge bases. Move to a dedicated vector database once you're indexing tens of thousands of chunks or need persistence across restarts.

How many chunks should I retrieve per query? Start with 3–5. Retrieving too many dilutes relevance and increases token cost; too few risks missing necessary context. Tune based on how your documents are structured.

Can I build RAG without an embedding model? Not effectively — embeddings are what let you find semantically relevant chunks instead of just keyword matches. Keyword search (like BM25) can supplement embeddings but rarely replaces them for natural-language questions.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →