Claude API Embeddings Generation: What Actually Works
Does the Claude API generate embeddings?
No — not directly. If you've been searching for a "Claude API embeddings generation example" and hitting a wall, that's because Anthropic's Messages API is built for text generation, tool use, and conversation, not for returning vector embeddings. There is no /v1/embeddings endpoint on Claude the way there is on some other providers.
What you actually want in almost every case is a two-part pipeline: an embedding model (Anthropic officially recommends Voyage AI for this) to turn text into vectors for search and retrieval, and Claude itself to generate the final answer once you've retrieved relevant context. This article shows you exactly how to set that up, with working code, so you're not stuck trying to force Claude into a job it isn't designed for.
Why Claude doesn't do embeddings
Claude is a generative, instruction-following model. Embedding models are a different architecture entirely — typically smaller, trained with contrastive objectives, and optimized to output fixed-length vectors that capture semantic similarity rather than produce text. Anthropic made a deliberate choice not to ship its own embeddings endpoint and instead partnered with Voyage AI, whose models (voyage-2, voyage-large-2, domain-specific variants like voyage-code-2) are tuned for retrieval tasks and benchmark well against general-purpose alternatives.
This separation is actually a feature, not a gap. It means you can:
- Pick the embedding model that fits your domain (code, legal text, multilingual content) independently of which generation model you use.
- Swap embedding providers without touching your generation layer.
- Keep embedding costs (usually cents per million tokens) separate from generation costs.
Step 1: generate embeddings with Voyage AI
Here's a minimal example generating embeddings for a batch of documents:
curl https://api.voyageai.com/v1/embeddings \
-H "Authorization: Bearer $VOYAGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": [
"Claude is a family of large language models by Anthropic.",
"Embeddings represent text as numerical vectors for similarity search."
],
"model": "voyage-2"
}'
The response contains a data array with one vector per input string. You store those vectors in a vector database (Pinecone, Weaviate, pgvector, Qdrant — any of them work fine) alongside the original text and metadata.
In JavaScript:
const response = await fetch("https://api.voyageai.com/v1/embeddings", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.VOYAGE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
input: documents.map(d => d.text),
model: "voyage-2",
}),
});
const { data } = await response.json();
// data[i].embedding is a float array — persist it with documents[i]
Step 2: retrieve relevant chunks at query time
When a user asks a question, embed their query the same way, then run a nearest-neighbor search against your vector store:
const queryEmbedding = await embedText(userQuestion); // same Voyage call as above
const matches = await vectorDb.query({
vector: queryEmbedding,
topK: 5,
});
const context = matches.map(m => m.metadata.text).join("\n\n");
This gives you the chunks of text most semantically similar to the question — the "retrieval" half of RAG (retrieval-augmented generation).
Step 3: let Claude generate the answer
Now feed the retrieved context into Claude as part of the prompt. This is where most teams run into the second half of the problem: managing API keys, rate limits, and usage tracking for the generation calls. If you're already using SubToAPI to turn your Claude access into a standard HTTPS API, this step looks like a normal Messages call:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Context:\n'"$context"'\n\nQuestion: What is Claude trained on?"
}
]
}'
Or in JavaScript, combining both halves of the pipeline:
const answer = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [{
role: "user",
content: `Context:\n${context}\n\nQuestion: ${userQuestion}`,
}],
}),
});
const result = await answer.json();
console.log(result.content[0].text);
Because SubToAPI issues application-scoped sub_live_ keys, you can wire this generation step into a product or internal tool without exposing your underlying Claude access, and you get per-key usage metadata so you can see exactly how much the generation half of your RAG pipeline costs versus the embedding half. See the quickstart and the Messages API reference for full request/response shapes, including streaming responses documented at /docs/streaming.
Putting it together: a simple RAG architecture
- Ingest: chunk your documents, embed each chunk with Voyage AI, store vectors + text in a vector database.
- Query: embed the incoming user question with the same model.
- Retrieve: fetch the top-k most similar chunks.
- Generate: pass the question plus retrieved chunks to Claude via the Messages API and return the answer.
This pattern scales from a weekend project to a production support bot. The only parts that change as you grow are chunking strategy, vector database choice, and how many documents you retrieve per query — the embedding-then-generate shape stays the same.
If you're building this into a customer-facing product, check pricing for the generation layer and start with a free trial at signup before committing to a plan.
questions
Can I use Claude to generate embeddings directly? No. The Claude Messages API only returns generated text (or tool calls), not vectors. For embeddings, use a dedicated model — Anthropic recommends Voyage AI, though OpenAI's or open-source embedding models work too if you're not already in the Anthropic ecosystem.
Which embedding model should I pick for a RAG pipeline with Claude? voyage-2 or voyage-large-2 are solid general-purpose choices and are the ones Anthropic specifically recommends alongside Claude. If your content is code-heavy, voyage-code-2 performs better on code similarity tasks.
Do I need embeddings if I'm just using Claude's large context window? Not always. If your documents fit comfortably in context, you can skip retrieval entirely and paste the full text into the prompt. Embeddings and retrieval become necessary once your knowledge base exceeds what you can afford to send on every request, both for cost and latency reasons.