← Blog

Claude API Vector Database Integration Tutorial

2026-10-01 · 5 min read · SubToAPI Team

If you're searching for how to connect the Claude API to a vector database, you're almost certainly building retrieval-augmented generation (RAG): a system that fetches relevant chunks of your own data and feeds them to Claude as context before it answers a question. This tutorial walks through the full pipeline — embedding your documents, storing and querying vectors, and constructing a Claude prompt that produces grounded answers instead of hallucinated ones.

The pattern is the same regardless of which vector database you pick (Pinecone, Weaviate, Qdrant, pgvector, Chroma). The steps are: chunk your text, embed the chunks, store the vectors with metadata, embed the user's query at runtime, retrieve the top matches, and pass them to Claude inside the system or user message. We'll use pgvector in the examples since it's free to self-host and works with plain SQL, but the logic transfers directly to any other vector store.

Step 1: Chunk and embed your documents

Vector databases don't understand raw text — they store numeric embeddings. You need an embedding model (OpenAI's text-embedding-3-small, Cohere, or a local model like all-MiniLM-L6-v2) to convert text chunks into vectors. Claude itself doesn't provide an embeddings endpoint, so this step always uses a separate provider.

import OpenAI from "openai";
const openai = new OpenAI();

async function embed(text) {
  const res = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: text,
  });
  return res.data[0].embedding;
}

function chunkText(text, size = 800, overlap = 100) {
  const chunks = [];
  for (let i = 0; i < text.length; i += size - overlap) {
    chunks.push(text.slice(i, i + size));
  }
  return chunks;
}

Keep chunks between 500–1000 characters with some overlap so you don't cut relevant context in half. Store the original text alongside each vector — you'll need it later as the content you hand to Claude.

Step 2: Store vectors in your database

With pgvector, create a table with a vector column:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE documents (
  id SERIAL PRIMARY KEY,
  content TEXT,
  embedding VECTOR(1536),
  metadata JSONB
);

CREATE INDEX ON documents USING ivfflat (embedding vector_cosine_ops);

Insert each chunk with its embedding and any metadata (source URL, page number, document title) you'll want to show as citations later.

async function storeChunk(content, embedding, metadata) {
  await db.query(
    "INSERT INTO documents (content, embedding, metadata) VALUES ($1, $2, $3)",
    [content, JSON.stringify(embedding), metadata]
  );
}

Step 3: Retrieve relevant chunks at query time

When a user asks a question, embed the query with the same model you used for indexing, then run a similarity search:

SELECT content, metadata, 1 - (embedding <=> $1) AS similarity
FROM documents
ORDER BY embedding <=> $1
LIMIT 5;

The <=> operator computes cosine distance in pgvector. Pull back 3–5 chunks — more than that tends to dilute Claude's context with noise rather than help it.

async function retrieve(query, k = 5) {
  const queryEmbedding = await embed(query);
  const { rows } = await db.query(
    `SELECT content, metadata, 1 - (embedding <=> $1) AS similarity
     FROM documents ORDER BY embedding <=> $1 LIMIT $2`,
    [JSON.stringify(queryEmbedding), k]
  );
  return rows;
}

Step 4: Build the Claude prompt

This is where retrieval meets generation. Put the retrieved chunks in the system prompt or as a prefixed block in the user message, clearly separated from the actual question, and instruct Claude to only answer from the provided context.

function buildPrompt(chunks, question) {
  const context = chunks
    .map((c, i) => `[${i + 1}] ${c.content}`)
    .join("\n\n");

  return {
    system: `Answer the user's question using only the context below. Cite sources using [number]. If the context doesn't contain the answer, say so.\n\nContext:\n${context}`,
    user: question,
  };
}

Now send this to Claude. If you're calling Claude directly through Anthropic, that's the Messages API. If you're already proxying Claude through SubToAPI to get a stable HTTPS key and usage tracking, the request shape is nearly identical — see the Messages API docs for the full schema:

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-3-5-sonnet-20241022",
    max_tokens: 1024,
    system: prompt.system,
    messages: [{ role: "user", content: prompt.user }],
  }),
});

const data = await res.json();
console.log(data.content[0].text);

For chat-style RAG apps where the answer should start appearing immediately, use streaming instead of waiting for the full response — see /docs/streaming for the server-sent events format. This matters a lot when your retrieved context is large, since the model has more to process before the first token comes back.

Handling citations and grounding

Ask Claude explicitly to cite the bracketed source numbers you included in the context, then map those numbers back to the metadata field you stored per chunk (URL, document title, page). This gives users a way to verify the answer instead of trusting it blindly — important for any RAG system dealing with internal docs, support tickets, or compliance content.

If your app also needs to call functions after retrieval — for example, pulling a live record from your database before finalizing an answer — combine this pattern with tool use, covered in /docs/tools.

Common mistakes to avoid

If you're prototyping and want to avoid managing a separate Anthropic billing account and key rotation while you iterate on this pipeline, SubToAPI turns your existing Claude access into an API key in a few minutes — start with the quickstart guide or check pricing before committing to a plan.

FAQ

Does Claude have a built-in vector database? No. Claude is a language model accessed through the Messages API; it has no native storage or retrieval layer. You need a separate vector database (pgvector, Pinecone, Weaviate, etc.) and a separate embedding model to build RAG.

Which vector database works best with Claude? Any of them — the integration point is just a text string you insert into the prompt. Choose based on your infrastructure: pgvector if you already run Postgres, Pinecone or Weaviate if you want a managed service with less ops overhead.

How many retrieved chunks should I send to Claude? Start with 3–5 chunks of 500–800 characters each. Test with your real queries and adjust — too few chunks miss context, too many add noise and unnecessary token cost without improving accuracy.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →