← Blog

Claude API RAG Pipeline Implementation Guide

2026-10-07 · 5 min read · SubToAPI Team

A RAG (retrieval-augmented generation) pipeline lets Claude answer questions using your own documents instead of relying only on what it learned during training. The core pattern is always the same: split your data into chunks, embed and store those chunks in a vector database, retrieve the most relevant chunks for a given query, and pass them to Claude as context alongside the user's question.

This article walks through a working implementation you can adapt to production: document chunking, embedding generation, vector retrieval, and the prompt structure that gets Claude to answer accurately and cite its sources. It assumes you're calling the Claude API directly or through a proxy like SubToAPI — the retrieval logic is identical either way.

Why RAG instead of fine-tuning or long context

Claude's context window is large, but stuffing entire knowledge bases into every request is slow and expensive. RAG keeps requests small by retrieving only the handful of passages relevant to each query. It also solves the staleness problem: when your underlying documents change, you re-embed and re-index — no retraining, no redeployment of a model.

Fine-tuning changes model behavior; RAG changes what information the model has access to. For most "answer questions about our docs/support tickets/codebase" use cases, RAG is the right tool.

Step 1: Chunk your documents

Chunk size matters more than most people expect. Too large and retrieval returns noisy, unfocused passages. Too small and you lose context needed to answer the question.

A reasonable starting point: 500–800 tokens per chunk with 10–15% overlap between consecutive chunks, split on paragraph or heading boundaries rather than arbitrary character counts.

function chunkText(text, maxTokens = 600, overlapTokens = 80) {
  const paragraphs = text.split(/\n\s*\n/);
  const chunks = [];
  let current = "";

  for (const para of paragraphs) {
    if ((current + para).length / 4 > maxTokens) {
      chunks.push(current.trim());
      const overlapChars = overlapTokens * 4;
      current = current.slice(-overlapChars) + para;
    } else {
      current += "\n\n" + para;
    }
  }
  if (current.trim()) chunks.push(current.trim());
  return chunks;
}

This is a rough token estimate (4 chars ≈ 1 token), good enough for chunking decisions. For production, use a proper tokenizer if precision matters.

Step 2: Generate embeddings and store them

Claude's API doesn't provide an embeddings endpoint, so you'll pair it with a dedicated embedding model (Voyage AI, OpenAI's embedding models, or an open-source model like bge-large) and a vector store (Pinecone, Qdrant, pgvector, or even an in-memory array for small datasets).

import { VoyageAIClient } from "voyageai";

const voyage = new VoyageAIClient({ apiKey: process.env.VOYAGE_API_KEY });

async function embedChunks(chunks) {
  const response = await voyage.embed({
    input: chunks,
    model: "voyage-3",
  });
  return response.data.map((d) => d.embedding);
}

Store each embedding alongside the chunk text and any metadata (source document, section title, URL) — you'll need the metadata to build citations later.

Step 3: Retrieve relevant chunks at query time

At query time, embed the user's question with the same model, then run a similarity search (cosine distance is standard) against your vector store and take the top-k results, typically k=4 to k=8.

async function retrieve(query, vectorStore, k = 6) {
  const [queryEmbedding] = await embedChunks([query]);
  const results = await vectorStore.query({
    vector: queryEmbedding,
    topK: k,
    includeMetadata: true,
  });
  return results.matches.map((m) => ({
    text: m.metadata.text,
    source: m.metadata.source,
  }));
}

If recall quality is a problem, add a reranking step: retrieve 20–30 candidates with the vector search, then rerank with a cross-encoder or with Claude itself scoring relevance, and keep the top 6.

Step 4: Assemble the prompt and call Claude

Put retrieved chunks in the system prompt or as structured context in the user message, label each chunk with its source, and instruct Claude to only answer from the provided context and to cite sources.

async function askWithContext(query, chunks) {
  const context = chunks
    .map((c, i) => `[${i + 1}] Source: ${c.source}\n${c.text}`)
    .join("\n\n---\n\n");

  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      system:
        "Answer the user's question using only the context provided below. " +
        "Cite sources using [n] notation. If the context doesn't contain " +
        "the answer, say so explicitly.\n\n" + context,
      messages: [{ role: "user", content: query }],
    }),
  });

  return response.json();
}

Using an explicit "don't guess" instruction materially reduces hallucination on out-of-context questions — Claude follows this reliably when the instruction is unambiguous.

If you're building this through SubToAPI, the request shape matches the standard Messages API — see /docs/messages for the full parameter reference, and /docs/streaming if you want to stream the answer back to users as it's generated, which matters a lot for RAG UIs where the retrieval step already adds latency.

Step 5: Handle grounding and citations

Return the source metadata alongside the response so your UI can render clickable citations. Don't rely on Claude to invent correct citation text — map the [n] markers back to the source list you built during retrieval, server-side.

function renderWithCitations(answerText, chunks) {
  return answerText.replace(/\[(\d+)\]/g, (match, n) => {
    const chunk = chunks[parseInt(n) - 1];
    return chunk ? `[${n}](${chunk.source})` : match;
  });
}

Common failure modes

Getting started

If you want to prototype quickly without managing separate Anthropic billing and keys, SubToAPI gives you an sub_live_... application key, streaming support, and usage dashboards out of the box — plans start at €9/month with a free trial. Check /docs/quickstart to get a key set up in a few minutes, then plug the askWithContext pattern above into your retrieval layer.

FAQ

Do I need Claude-specific embeddings for RAG? No. Claude's API doesn't generate embeddings — use a dedicated embedding model (Voyage AI, OpenAI, or an open-source alternative) for the retrieval step, and Claude only for generating the final answer from retrieved context.

How many chunks should I retrieve per query? Start with 4–8 chunks (top-k). If answers feel incomplete, increase k or add a reranking step; if answers feel noisy or unfocused, reduce k or shrink chunk size.

Can I stream RAG responses to reduce perceived latency? Yes. Retrieval adds latency before the Claude call starts, but once the request is sent you can stream the completion token by token — see /docs/streaming for implementation details.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →