Claude API Embeddings Generation Tutorial
Does the Claude API generate embeddings?
No — and this trips up a lot of developers searching for a "Claude embeddings endpoint." Anthropic's API (and SubToAPI, which proxies it) is built for chat completions: sending messages and getting text, tool calls, or streamed tokens back. It does not expose a /embeddings route the way OpenAI's API does.
What you actually want, in almost every case, is a two-model pipeline: a dedicated embedding model turns your text into vectors for search/retrieval, and Claude handles the generation step — reading the retrieved context and writing the answer. This is the standard RAG (retrieval-augmented generation) architecture, and it's what Anthropic itself recommends. This tutorial shows you how to build it correctly, including which embedding model to use and how to wire the generation half through the Claude API.
Why Claude doesn't do embeddings
Embedding models and chat/completion models are trained for different jobs. Embedding models are optimized to produce fixed-length vectors where semantic similarity maps to vector distance — good for search, clustering, and deduplication. Claude is optimized for reasoning, instruction-following, and long-context generation. Anthropic keeps these separate on purpose and partners with Voyage AI for embeddings, since Voyage's models are tuned for retrieval quality and tested against Claude's context handling.
If you're building a RAG app, knowledge base search, or semantic recommendation system, you need both pieces working together, not one model trying to do both jobs.
Step 1: Generate embeddings with Voyage AI
Sign up for a Voyage AI API key, then embed your documents:
curl https://api.voyageai.com/v1/embeddings \
-H "Authorization: Bearer $VOYAGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": ["Claude is a family of large language models made by Anthropic."],
"model": "voyage-3"
}'
The response gives you a float array per input string:
{
"data": [
{ "embedding": [0.0123, -0.0456, ...] }
],
"model": "voyage-3"
}
Batch this across every chunk of your document set. A practical chunking rule: 300–800 tokens per chunk, with slight overlap, so each chunk is self-contained enough to be useful on its own.
Step 2: Store vectors and run similarity search
Store the embeddings in a vector database — Pinecone, Weaviate, Qdrant, or even Postgres with pgvector for smaller projects. Each record needs: the vector, the original text chunk, and metadata (source document, section, URL).
await vectorDB.upsert([
{
id: "doc1-chunk3",
vector: embedding,
metadata: { text: chunkText, source: "handbook.pdf" }
}
]);
At query time, embed the user's question with the same model, then run a nearest-neighbor search:
const queryEmbedding = await embed(userQuestion);
const matches = await vectorDB.query({
vector: queryEmbedding,
topK: 5
});
You now have the 5 most relevant chunks. This is the retrieval half of RAG — done entirely without Claude.
Step 3: Generate the answer with Claude
Now pass the retrieved chunks as context into a Claude API call. This is where SubToAPI fits in: it gives you a plain HTTPS endpoint (sub_live_... key) so you don't have to manage separate Anthropic billing or OAuth just for this generation step.
const context = matches.map(m => m.metadata.text).join("\n\n");
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
system: "Answer only using the provided context. If the answer isn't in the context, say so.",
messages: [
{
role: "user",
content: `Context:\n${context}\n\nQuestion: ${userQuestion}`
}
]
})
});
const result = await response.json();
console.log(result.content[0].text);
That's the full pipeline: embed → retrieve → generate. Embeddings never touch Claude directly; they just decide what text gets put in front of it.
Keeping the two systems in sync
A few practical notes that save debugging time later:
- Re-embed on content changes. If source documents update, re-run embeddings for the affected chunks only — don't re-embed the whole corpus on every edit.
- Match embedding model versions. Mixing vectors from
voyage-2andvoyage-3in the same index gives inconsistent similarity scores. Re-embed fully if you switch model versions. - Keep context under Claude's limit. Even with large context windows, stuffing in 20 retrieved chunks instead of 5 degrades answer quality — it's not just a token-count problem, it's a signal-to-noise problem.
- Stream the generation step. For chat-style RAG apps, streaming the Claude response back to the user (see /docs/streaming) makes the latency of the retrieval step feel invisible.
Where SubToAPI fits
SubToAPI doesn't replace your embedding model — you still need Voyage AI or an equivalent for the vector side. What it simplifies is the generation half: instead of managing raw Anthropic API credentials and tracking usage manually, you get a sub_live_... key, usage metadata per request, and team seats if more than one person on your team is building against the same Claude backend. Check the docs for the full request/response reference, or the quickstart if you're setting this up for the first time. Plans start at €9/month for solo use, with team and scale tiers at /pricing, and every plan starts with a free trial at /signup.
Questions
Does Anthropic have its own embeddings model? No. Anthropic does not ship a native embeddings endpoint. It officially recommends Voyage AI's embedding models for use alongside Claude in retrieval and RAG applications.
Can I use OpenAI embeddings with Claude for generation? Yes. Embeddings and generation models don't need to be from the same provider — plenty of production RAG systems use OpenAI or Cohere embeddings paired with Claude for the final answer. Just be consistent about which embedding model you use across your whole index.
Do I need a vector database, or can I use plain search? For small datasets (a few hundred documents), keyword search or even a simple cosine-similarity loop in memory can work fine. Once you're past a few thousand chunks, a dedicated vector database (Pinecone, Qdrant, pgvector) becomes worth the setup for query speed and filtering.