Claude API and Pinecone for RAG: A Build Guide
What RAG with Claude and Pinecone actually solves
Retrieval augmented generation (RAG) pairs a vector database like Pinecone with a language model like Claude so the model can answer questions using your own documents instead of relying only on what it learned during training. Pinecone stores embeddings of your content and finds the most relevant chunks for a given query; Claude then reads those chunks and generates an answer grounded in them. This is the standard pattern for building internal knowledge-base bots, customer support assistants, and documentation search tools that need accurate, up-to-date answers.
The short version of how it works: you embed your documents once, store the vectors in Pinecone, and at query time you embed the user's question, retrieve the closest matches, and stuff those matches into Claude's context window before asking it to answer. The rest of this article walks through the architecture and a working implementation.
The RAG architecture, step by step
A Claude + Pinecone RAG pipeline has four stages:
- Chunk and embed your source documents. Split PDFs, wiki pages, or support tickets into chunks of a few hundred tokens each, then run them through an embedding model to get vectors.
- Upsert vectors into Pinecone. Each vector is stored with metadata (document title, source URL, chunk text) so you can retrieve the original text later.
- Query time retrieval. Embed the incoming user question, query Pinecone for the top-k nearest vectors, and pull the associated text back out.
- Generate with Claude. Pass the retrieved chunks as context in the prompt, along with the user's question, and let Claude produce a grounded answer — ideally citing which chunk it used.
The quality of the whole system depends more on steps 1–3 (chunking strategy, embedding model, metadata filtering) than on the generation step itself. Claude is good at synthesizing an answer from provided context — it's the retrieval side that usually needs tuning.
Setting up the Pinecone index
Create an index sized to your embedding model's dimensions (1536 for many OpenAI embedding models, 1024 for some open-source alternatives):
curl -X POST "https://api.pinecone.io/indexes" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "docs-index",
"dimension": 1536,
"metric": "cosine",
"spec": { "serverless": { "cloud": "aws", "region": "us-east-1" } }
}'
Upsert embeddings with the source text stored in metadata so you don't need a separate lookup step:
await index.upsert([
{
id: "doc-42-chunk-3",
values: embeddingVector,
metadata: {
text: "Refunds are processed within 5 business days...",
source: "refund-policy.md"
}
}
]);
Querying Pinecone and calling Claude
At query time, embed the user's question, retrieve the top matches, and build a prompt that includes the retrieved text as context:
const queryEmbedding = await embed(userQuestion);
const results = await index.query({
vector: queryEmbedding,
topK: 5,
includeMetadata: true,
});
const context = results.matches
.map((m) => `Source: ${m.metadata.source}\n${m.metadata.text}`)
.join("\n\n---\n\n");
const prompt = `Answer the question using only the context below.
If the context doesn't contain the answer, say so.
Context:
${context}
Question: ${userQuestion}`;
You can send that prompt to Claude through SubToAPI's /v1/messages endpoint, which mirrors the standard Messages API shape:
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [{ role: "user", content: prompt }],
}),
});
const data = await response.json();
console.log(data.content[0].text);
This is useful if you're building the RAG pipeline as part of a product with multiple team members or environments: SubToAPI turns your Claude access into a standard HTTPS API with sub_live_... application keys, so each service (retrieval worker, chat UI, admin tool) gets its own key and usage is tracked separately in one dashboard instead of sharing a single raw Anthropic key. See the quickstart and Messages API reference for the full request/response shape.
Streaming and tool use in a RAG app
Two things improve a production RAG assistant beyond the basic request/response loop:
- Streaming the response so users see tokens as they're generated, which matters a lot for longer answers built from retrieved context. SubToAPI supports streaming the same way the standard API does — see streaming docs.
- Tool use for cases where the model needs to decide whether to retrieve at all, or needs to call a second lookup (e.g., fetch a specific document by ID after an initial search). Claude's tool-calling lets you expose a
search_documentsfunction that the model invokes only when it determines retrieval is necessary, rather than always forcing a retrieval step. Details are in the tool use guide.
Common mistakes in Claude + Pinecone RAG pipelines
- Chunking too large or too small. Chunks that are too big dilute relevance scores; chunks that are too small lose context. 200–500 tokens per chunk with slight overlap is a reasonable starting point.
- Not filtering by metadata. If your index holds multiple document types or tenants, always filter Pinecone queries by namespace or metadata — otherwise you'll leak irrelevant or cross-tenant content into the prompt.
- Skipping citation instructions. Tell Claude explicitly to cite which source chunk it used; this makes it much easier to debug bad answers and builds user trust.
- Ignoring retrieval failures. If Pinecone returns low-similarity matches, instruct Claude to say it doesn't know rather than guessing from weak context.
Getting started
If you already have Claude access and want to wire it into a RAG pipeline without managing raw API credentials across every service, sign up for a free trial and generate a scoped key for your retrieval worker. Plans start at €9/month for solo projects, with team pricing at €19/seat and scale pricing at €49/seat — see pricing for details.
Questions
Does Claude have built-in retrieval, or do I need Pinecone? Claude doesn't ship with a built-in vector database. You need an external retrieval layer — Pinecone, pgvector, or similar — to find relevant chunks before passing them to Claude in the prompt.
Which Claude model works best for RAG answers? Claude 3.5 Sonnet is a strong default for RAG: it's fast enough for interactive use and reliable at synthesizing answers strictly from provided context without over-relying on prior knowledge.
How many chunks should I retrieve per query? Start with 3–5 chunks (top-k) and adjust based on answer quality. Too few chunks miss context; too many add noise and increase token cost without improving accuracy.