Claude API Vector Database Integration Guide
If you're building anything that needs Claude to answer questions about your own data — docs, support tickets, product catalogs, internal wikis — you need a vector database in the loop. Claude itself doesn't store or search your data; it only sees what you put in the prompt. A vector database is what finds the relevant slice of your data to put there. This guide covers the actual integration pattern: how embeddings, chunking, retrieval, and the Claude API call fit together, with working code.
The short version: you embed your documents once, store the vectors, embed the user's query at request time, retrieve the top matches, and stuff those matches into the Claude prompt as context. This is retrieval-augmented generation (RAG), and it works with any vector database — Qdrant, Weaviate, Chroma, pgvector, Milvus. The database choice matters less than getting chunking, embedding, and prompt construction right, which is what most "integration guides" skip.
The architecture
A RAG pipeline with Claude has four stages:
- Ingestion — split documents into chunks, embed each chunk, store the vector + metadata (source, chunk text) in the vector DB.
- Retrieval — embed the incoming query, run a similarity search, pull back the top-k chunks.
- Context assembly — format the retrieved chunks into the prompt, usually in the system prompt or as a prefixed block in the user message.
- Generation — send the assembled prompt to Claude and return the answer.
Claude only touches steps 3 and 4. Steps 1 and 2 are entirely your vector database and embedding model's job — Claude's API has no native vector storage, so this separation is unavoidable and, honestly, a good thing: it lets you swap embedding models or databases without touching your Claude integration code.
Step 1: chunk and embed your documents
Chunk size matters more than most people think. Too large and retrieval gets noisy; too small and you lose context. 300–800 tokens per chunk with ~10-15% overlap is a reasonable default for most prose documents.
import { OpenAI } from "openai"; // any embedding provider works
const openai = new OpenAI();
async function embedChunks(chunks) {
const res = await openai.embeddings.create({
model: "text-embedding-3-small",
input: chunks,
});
return res.data.map(d => d.embedding);
}
Claude's API doesn't generate embeddings itself, so you'll pair it with an embedding model from any provider, then store the vectors in your chosen database:
await qdrantClient.upsert("docs", {
points: chunks.map((text, i) => ({
id: i,
vector: embeddings[i],
payload: { text, source: filename },
})),
});
Step 2: retrieve at query time
When a user asks a question, embed the query with the same model used for ingestion, then search:
const [queryVector] = await embedChunks([userQuestion]);
const results = await qdrantClient.search("docs", {
vector: queryVector,
limit: 5,
});
const context = results.map(r => r.payload.text).join("\n\n---\n\n");
Five chunks is a decent starting point. Fewer and you risk missing the answer; more and you pad the prompt with irrelevant text that can confuse the model or push you past context budgets.
Step 3: call Claude with the retrieved context
This is where SubToAPI comes in if you're already using Claude through it — you get a straightforward HTTPS endpoint with an application API key instead of managing the raw Anthropic SDK and credentials yourself. The request shape is a standard messages call:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-3-7-sonnet-20250219",
"max_tokens": 1024,
"system": "Answer only using the provided context. If the answer isn'"'"'t in the context, say so.",
"messages": [
{ "role": "user", "content": "Context:\n'"$context"'\n\nQuestion: '"$userQuestion"'" }
]
}'
The system prompt instruction to stick to the provided context is doing real work here — without it, Claude will happily fall back on general knowledge, which defeats the point of RAG when you need grounded, source-specific answers. See /docs/messages for the full request/response reference.
Keeping the pipeline maintainable
A few practices that save headaches as this scales:
- Re-embed on content changes, not on a schedule. Hook ingestion into your content pipeline (CMS webhook, file watcher) rather than cron jobs that re-embed everything nightly.
- Store source metadata with every vector. You'll want to cite sources in the final answer, and debugging bad retrievals is much easier when you can see where a chunk came from.
- Log retrieval scores separately from generation. If answers go wrong, you need to know whether retrieval failed (wrong chunks returned) or generation failed (right chunks, bad answer). Mixing these into one log makes debugging painful.
- Cache embeddings for repeated queries. FAQs and common questions don't need re-embedding every time.
- Stream the final answer. Once you have context in hand, retrieval is usually the slow part; streaming the Claude response back to the user keeps perceived latency low. SubToAPI supports streaming — see /docs/streaming.
Picking a vector database
For most teams starting out, the decision tree is simple:
- Already on Postgres? Use pgvector. One less service to run.
- Need a managed, zero-ops option? A hosted vector DB (Qdrant Cloud, Weaviate Cloud) removes the infra burden.
- Running fully self-hosted / air-gapped? Qdrant or Milvus are solid, battle-tested open-source choices.
- Prototyping fast? Chroma runs in-process with almost no setup.
None of these choices change how you talk to Claude — the integration code in steps 2-3 above is identical regardless of which database sits behind qdrantClient.search(...).
If you're evaluating whether to build this on raw Anthropic API access or through a managed layer, SubToAPI gives you the same Claude models behind a standard API key with usage metadata per key, which is useful once you have multiple services (ingestion jobs, chat endpoints, internal tools) all calling Claude and you want visibility into which one is burning tokens. Check /pricing or start with a free trial at /signup.
FAQs
Does Claude have a built-in vector database or embeddings endpoint? No. Claude's API handles chat/completion, tool use, and vision, but embeddings and vector storage are separate concerns you handle with an embedding model (OpenAI, Voyage, Cohere, or open-source) and a vector database of your choice.
How many retrieved chunks should I send to Claude? Start with 3-5 chunks of 300-800 tokens each. Tune based on answer quality — too few chunks miss context, too many dilute relevance and increase cost and latency without improving accuracy.
Can I use SubToAPI instead of calling the Anthropic API directly for RAG? Yes. The messages endpoint at /docs/messages accepts the same context-plus-question pattern shown above; you just swap your endpoint and API key. Streaming and tool use work the same way, documented at /docs/streaming and /docs/tools.