← Blog

Claude API LlamaIndex Integration Guide

2026-10-05 · 4 min read · SubToAPI Team

Claude API LlamaIndex Integration Guide

If you're building a retrieval-augmented generation (RAG) pipeline and want Claude as your LLM instead of GPT, LlamaIndex supports this directly through its llama-index-llms-anthropic package. The integration is straightforward: you configure Claude as the LLM object, point LlamaIndex's query engine at it, and your existing index, retriever, and embedding setup stay unchanged.

This guide walks through the actual setup — installing the right packages, authenticating, configuring Claude as the model for both indexing-adjacent tasks and query-time generation, and the two things people usually get wrong: embedding model choice (Claude doesn't do embeddings) and context window management when stuffing retrieved chunks into prompts.

Why use Claude with LlamaIndex

LlamaIndex is a data framework — it handles chunking, embedding, storage, and retrieval. The LLM itself is just the final step: turning retrieved context into an answer. Claude is a strong fit here because of its large context window (useful when retrieval returns several large chunks) and its lower hallucination rate on grounded Q&A tasks compared to some alternatives. None of that changes how LlamaIndex works — you're swapping one LLM provider for another.

Step 1: Install dependencies

pip install llama-index
pip install llama-index-llms-anthropic
pip install llama-index-embeddings-openai  # or another embedding provider

Claude does not offer an embeddings API, so you need a separate embedding model. OpenAI's text-embedding-3-small, Cohere, or a local model like BAAI/bge-small-en all work fine. This is the most common point of confusion — people expect Claude to handle embeddings because it handles generation, but it's a generation-only model.

Step 2: Configure Claude as the LLM

from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings

llm = Anthropic(
    model="claude-sonnet-4-20250514",
    api_key="YOUR_ANTHROPIC_API_KEY",
    max_tokens=1024,
)

Settings.llm = llm

Setting Settings.llm globally means every query engine, chat engine, and agent you build afterward uses Claude by default. If you only want Claude for a specific query engine, pass llm=llm directly to that constructor instead.

Step 3: Build the index and query engine

from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.embeddings.openai import OpenAIEmbedding

Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")

documents = SimpleDirectoryReader("./data").load_data()
index = VectorStoreIndex.from_documents(documents)

query_engine = index.as_query_engine()
response = query_engine.query("What does the Q3 report say about churn?")
print(response)

At query time, LlamaIndex retrieves the top-k relevant chunks from your vector store, builds a prompt with those chunks as context, and sends it to Claude. The response object includes the generated answer plus source nodes you can inspect for citation purposes.

Step 4: Tune retrieval and context window usage

Claude's larger context models can handle more retrieved chunks per query than smaller-context LLMs, but more context isn't automatically better — irrelevant chunks dilute the signal and increase cost. Start with similarity_top_k=3 to 5 and adjust based on answer quality:

query_engine = index.as_query_engine(similarity_top_k=5)

If you're working with long documents (contracts, research papers, transcripts), consider a SentenceWindowNodeParser or hierarchical node parser instead of fixed-size chunking — Claude handles longer, more coherent chunks well, so you don't need to chop text as aggressively as you might for a smaller-context model.

Step 5: Streaming and chat engines

For conversational RAG, use a chat engine instead of a query engine:

chat_engine = index.as_chat_engine(chat_mode="context", llm=llm)
response = chat_engine.chat("Summarize the key risks mentioned in these documents.")

Streaming works the same way LlamaIndex streams any LLM response:

response = query_engine.query("Explain the methodology section.")
for token in response.response_gen:
    print(token, end="")

Handling rate limits and production concerns

Two issues come up once you move past prototyping: rate limits on the Anthropic API during high query volume, and the need to track usage per application or per team when multiple services share the same Claude access. Anthropic's native API doesn't give you per-key usage breakdowns or team-level dashboards out of the box — if you're running several LlamaIndex-based apps off one Claude subscription, that visibility gap becomes a real operational problem.

This is where routing requests through SubToAPI helps. Instead of pointing the Anthropic LLM client at Anthropic directly, you point it at SubToAPI's endpoint and use an application-scoped key (sub_live_...). You get the same streaming and tool-use behavior LlamaIndex expects, but with per-key usage metadata so you can see exactly which app or team is burning tokens. Swapping the base URL and key is a config change, not a code rewrite — check the quickstart for the exact client setup, and messages and streaming docs for request format details that map directly onto what LlamaIndex sends under the hood.

Common pitfalls

questions

Does LlamaIndex support Claude's tool use / function calling? Yes, through the FunctionAgent or ReActAgent classes combined with the Anthropic LLM integration. Tool definitions get passed through to Claude's native tool-use API, so you get structured outputs rather than parsed free text.

Can I use Claude for both retrieval and generation in LlamaIndex? No — Claude doesn't provide an embeddings endpoint, so retrieval (embedding and similarity search) always relies on a separate embedding provider. Claude only handles the final generation step.

What's the easiest way to monitor token usage across multiple LlamaIndex apps using Claude? Anthropic's raw API doesn't break down usage per application. Routing calls through a layer like SubToAPI with separate application keys gives you that breakdown without changing your LlamaIndex code structure — see pricing for plan details.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →