Claude API Vector Database Connection Guide
Connecting the Claude API to a vector database means building a retrieval-augmented generation (RAG) pipeline: you store embedded chunks of your documents in a vector store, retrieve the most relevant ones for a given query, and inject them into the prompt you send to Claude. Claude itself has no built-in vector database — Anthropic doesn't ship one, and there's no special "vector mode" in the API. The connection happens entirely on your side: embed, store, retrieve, then call the Messages API with the retrieved text as context.
This article walks through the actual mechanics of that connection — which embedding model to use, how to structure the retrieval step, how to format retrieved chunks for Claude, and where this gets tricky in production (stale indexes, chunk size, citation, cost).
Why Claude Doesn't Need Its Own Vector Store
Some API providers bundle vector search directly into their platform. Claude's API is intentionally unopinionated: it's a text-in, text-out (or tool-call-out) interface. This is actually an advantage for RAG — you're not locked into one vendor's retrieval quality, and you can swap vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma) without touching how you call Claude.
The pipeline always has the same shape:
- Embed your documents with an embedding model (Claude doesn't provide one — use Voyage AI, OpenAI's embedding models, or an open-source model like
bge-large). - Store the vectors plus metadata in your vector database.
- Query the vector database with an embedded version of the user's question.
- Inject the top-k results into the Claude prompt as context.
- Call the Messages API and let Claude reason over that context.
Step 1: Embedding and Storing Documents
Pick an embedding model independent of Claude. Anthropic has previously recommended Voyage AI embeddings for Claude-based RAG because they're tuned for retrieval quality, but any embedding model that fits your budget and latency requirements works.
import { VoyageAIClient } from "voyageai";
const voyage = new VoyageAIClient({ apiKey: process.env.VOYAGE_API_KEY });
async function embedChunks(chunks) {
const result = await voyage.embed({
input: chunks,
model: "voyage-2",
});
return result.data.map(d => d.embedding);
}
Store each embedding alongside the original text and a reference ID in your vector database. Chunk size matters more than people expect — 300–800 tokens per chunk with some overlap tends to balance recall against noise. Smaller chunks retrieve more precisely but lose surrounding context; larger chunks carry context but dilute relevance scoring.
Step 2: Querying the Vector Database
At query time, embed the user's question with the same model and run a similarity search:
const queryEmbedding = await embedChunks([userQuestion]);
const results = await vectorDb.query({
vector: queryEmbedding[0],
topK: 5,
includeMetadata: true,
});
Keep topK small (3–8) initially. Dumping 20 chunks into the prompt increases cost and can actually hurt answer quality by burying the relevant passage among irrelevant ones.
Step 3: Building the Claude Prompt
This is the part that's actually "the connection" — formatting retrieved text so Claude can use it well. A clear structure with XML-style tags works reliably with Claude models:
const context = results.matches
.map((m, i) => `<document index="${i}">\n${m.metadata.text}\n</document>`)
.join("\n\n");
const systemPrompt = `You are a helpful assistant. Answer the user's question using only the information in the documents below. If the answer isn't in the documents, say you don't know.
${context}`;
Then send this as context with the user's actual question. If you're calling Claude directly through Anthropic's API, this is a standard Messages API call. If you're routing through SubToAPI, the request shape is the same — you're just hitting your own sub_live_... key instead of managing Anthropic credentials and billing directly:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"system": "'"$systemPrompt"'",
"messages": [
{"role": "user", "content": "'"$userQuestion"'"}
]
}'
SubToAPI doesn't change anything about how RAG works — it's still your embedding pipeline and your vector database. What it handles is turning your Claude access into a stable HTTPS endpoint with its own API keys, so the retrieval service and the generation service in your architecture both authenticate against one dashboard instead of juggling raw Anthropic credentials per environment. See /docs/messages for the full request format.
Handling Citations and Grounding
For production RAG, users usually want to know where an answer came from. Pass metadata (source URL, document title, page number) alongside each chunk and ask Claude to cite it explicitly:
const context = results.matches
.map((m, i) => `<document index="${i}" source="${m.metadata.source}">\n${m.metadata.text}\n</document>`)
.join("\n\n");
Then instruct Claude: "Cite the source attribute for any claim you make." Claude follows structured citation instructions well when the source format is consistent and labeled clearly.
Keeping the Index Fresh
A common failure mode in Claude + vector database setups isn't the API connection — it's a stale index. If your documents change and you don't re-embed, Claude will confidently answer with outdated context. Set up a re-indexing job (cron, webhook-triggered, or event-driven on document updates) rather than treating the vector store as a one-time load.
Streaming the Final Answer
Once retrieval is done, the generation call behaves like any other Claude request — streaming works the same way. See /docs/streaming if you want token-by-token output for a RAG chat interface rather than waiting for the full response.
Common Mistakes
- Embedding model mismatch: using one model to embed documents and a different one to embed queries breaks similarity scoring entirely.
- No chunk overlap: hard cuts between chunks can split a sentence that contains the answer.
- Too much context: stuffing 15+ chunks into the prompt increases token cost and can reduce answer precision.
- No fallback: always instruct Claude to say "I don't know" when retrieval returns nothing relevant, instead of letting it guess.
FAQ
Does Claude have a native vector database integration? No. Anthropic's API doesn't include a vector store. You connect Claude to a vector database yourself: embed documents externally, retrieve relevant chunks, and pass them as context in your Messages API call.
Which embedding model should I use with Claude? Any model independent of Claude works — Voyage AI, OpenAI embeddings, or open-source models like bge-large are all common choices. Just make sure you use the same model for both documents and queries.
Can SubToAPI replace my vector database? No. SubToAPI turns your Claude access into an API with keys, streaming, and usage metadata — see /docs/quickstart. Your vector database and retrieval logic stay in your own stack; SubToAPI just handles the generation call.