Claude API Semantic Search: Implementation Guide
Semantic search means finding content by meaning rather than exact keyword matches. If someone searches "how to cancel a subscription" your system should also surface a document titled "ending your plan" even though the words don't overlap. The Claude API doesn't do this alone — it's not a search engine — but it's the component that makes semantic search actually useful: it turns raw retrieved chunks into accurate, well-cited answers, and it can rerank or filter results that a pure vector similarity search gets wrong.
This guide walks through the full pipeline: generating embeddings, storing and querying them, retrieving candidates, and using Claude to rerank and synthesize a final answer. It's written for developers who already have a dataset (docs, support tickets, product catalog, internal wiki) and want to let users ask natural-language questions against it.
The Three-Stage Architecture
A production semantic search system built around Claude has three distinct stages, and conflating them is the most common mistake:
- Retrieval — an embedding model finds the top N candidate chunks by vector similarity. This is fast and cheap but imprecise.
- Reranking — Claude (or a cross-encoder) re-scores those N candidates against the actual query to pick the best K.
- Synthesis — Claude generates a natural-language answer grounded in the top K chunks, with citations back to source documents.
Skipping stage 2 is fine for simple use cases. Skipping stage 3 turns your product into "search results" instead of "an answer" — which is usually what users actually want.
Step 1: Chunk and Embed Your Content
Before anything touches Claude, split your source documents into chunks of roughly 200–500 tokens with some overlap, and generate embeddings for each chunk using a dedicated embedding model (not Claude — Claude's API doesn't expose embeddings, so you'll need a separate embedding provider or an open-source model). Store the vectors alongside the original text and metadata (source URL, title, timestamp) in whatever vector store you're using.
// Pseudocode — embedding step uses your embedding provider, not Claude
async function indexDocument(doc) {
const chunks = splitIntoChunks(doc.text, { size: 400, overlap: 50 });
for (const chunk of chunks) {
const embedding = await embedText(chunk.text);
await vectorStore.upsert({
id: chunk.id,
vector: embedding,
metadata: { text: chunk.text, source: doc.url, title: doc.title },
});
}
}
Step 2: Retrieve Candidates
At query time, embed the user's question with the same embedding model and pull the top 10–20 nearest neighbors from your vector store. This gets you a rough shortlist — fast, but noisy, especially for ambiguous queries or questions that mix multiple concepts.
const queryVector = await embedText(userQuery);
const candidates = await vectorStore.query(queryVector, { topK: 15 });
Step 3: Rerank with Claude
Raw cosine similarity often ranks a tangentially related chunk above a directly relevant one, especially when the query is phrased informally. Claude is good at this kind of judgment call because it can read the actual text, not just compare vectors. Send the candidates and the query, and ask for a ranked subset:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 500,
"messages": [{
"role": "user",
"content": "Query: \"how do I cancel my plan\"\n\nCandidates:\n[1] \"Ending your subscription...\"\n[2] \"Billing cycle overview...\"\n[3] \"Downgrade vs cancel...\"\n\nReturn the IDs of the 3 most relevant candidates, ranked, as a JSON array."
}]
}'
This is a cheap call relative to full generation because you're only sending chunk snippets, not entire documents, and the response is a short JSON array. If your candidate list is large, truncate each chunk to its first 100–150 tokens for reranking — you only need enough text for Claude to judge relevance, not the full passage.
Step 4: Synthesize the Answer
Once you have the top K reranked chunks, build a final prompt that includes them as context and asks Claude to answer using only that material, with citations:
const response = await fetch('https://api.subtoapi.app/v1/messages', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
'content-type': 'application/json',
},
body: JSON.stringify({
model: 'claude-sonnet-4-5',
max_tokens: 1024,
system:
'Answer using only the provided context. Cite the source title for each claim. If the context does not contain the answer, say so explicitly.',
messages: [
{
role: 'user',
content: `Context:\n${topChunks.map(c => `[${c.title}] ${c.text}`).join('\n\n')}\n\nQuestion: ${userQuery}`,
},
],
}),
});
The explicit instruction to say "I don't know" when the context is insufficient matters a lot — without it, Claude will sometimes extrapolate confidently from a loosely related chunk, which defeats the point of grounding the answer in your actual data.
For answers that stream to the UI as they're generated — which matters a lot for perceived latency in a search product — route this same request through streaming instead of waiting for the full response; see /docs/streaming for the implementation details.
Handling Multi-Turn Search Refinement
Users rarely get their query right the first time. A good semantic search UX lets them follow up — "no, I meant the enterprise plan" — without re-explaining context. Keep the conversation history in the messages array across turns, and re-run retrieval only when the follow-up introduces new terms that aren't covered by the existing context. Claude can also help decide this: ask it whether the follow-up requires new retrieval or can be answered from the existing context alone, as a cheap classification call before you hit the vector store again.
Cost and Latency Tradeoffs
Reranking and synthesis both consume tokens on every query, which adds up at scale. A few practical levers:
- Cache reranking results for frequently repeated queries.
- Use a smaller/faster Claude model for reranking and a stronger one only for final synthesis.
- Keep context chunks tight — send only what's needed, not entire source documents.
- Track usage per endpoint so you know which stage is actually driving cost.
If you're running this through SubToAPI, usage metadata is attached to every response, so you can see token counts per call and separate reranking cost from synthesis cost without building your own instrumentation. Getting a working key takes a couple of minutes — see /docs/quickstart — and the request/response format is documented at /docs/messages if you want the full field reference while wiring this up.
questions
Does Claude generate embeddings for semantic search? No. The Claude API handles reasoning, reranking, and answer generation, but not vector embeddings — you'll need a separate embedding model or provider for the retrieval step.
Do I need a reranking step, or can I just use vector similarity? Vector similarity alone works for simple, well-phrased queries. Reranking with Claude noticeably improves precision on ambiguous or multi-concept queries, at the cost of one extra API call per search.
What's the biggest mistake teams make implementing this? Skipping the "say so explicitly" instruction when context doesn't contain the answer. Without it, Claude will often generate a plausible-sounding but ungrounded answer instead of admitting the retrieved content is insufficient.