← Blog

Claude API LlamaIndex Integration Tutorial

2026-10-02 · 5 min read · SubToAPI Team

LlamaIndex is one of the most popular frameworks for building retrieval-augmented generation (RAG) pipelines, and Claude is a strong choice for the generation step because of its large context window and reliable instruction-following. This tutorial walks through connecting Claude to LlamaIndex, indexing documents, and running queries against your own data — from a basic setup to a production-ready configuration.

By the end you'll have a working query engine that retrieves relevant chunks from your documents and sends them to Claude for a grounded answer, plus notes on handling streaming, cost control, and API key management when you move past prototyping.

Prerequisites

LlamaIndex ships an official Anthropic integration package, so you don't need to write a custom LLM wrapper for the standard case.

Step 1: Install dependencies

pip install llama-index llama-index-llms-anthropic llama-index-embeddings-huggingface

You'll also need an embedding model. LlamaIndex defaults to OpenAI embeddings unless you configure something else, so this tutorial uses a free local embedding model to keep the stack independent of any single provider — you can swap this for whatever embeddings you already use.

Step 2: Configure Claude as the LLM

from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings

Settings.llm = Anthropic(
    model="claude-sonnet-4-20250514",
    api_key="your-anthropic-key",
    max_tokens=1024,
)

If you're routing requests through SubToAPI instead of calling Anthropic directly, point the client at SubToAPI's base URL and use your sub_live_... key:

from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings

Settings.llm = Anthropic(
    model="claude-sonnet-4-20250514",
    api_key="sub_live_your_key",
    base_url="https://api.subtoapi.app/v1",
    max_tokens=1024,
)

This is useful when multiple services or team members share one underlying Claude subscription but each needs its own scoped key, usage visibility, and rate limits — see the docs for how application keys work.

Step 3: Set the embedding model

from llama_index.embeddings.huggingface import HuggingFaceEmbedding

Settings.embed_model = HuggingFaceEmbedding(
    model_name="BAAI/bge-small-en-v1.5"
)

Embeddings are only used for retrieval, not generation, so they don't need to come from the same provider as your LLM.

Step 4: Load and index your documents

from llama_index.core import SimpleDirectoryReader, VectorStoreIndex

documents = SimpleDirectoryReader("./data").load_data()
index = VectorStoreIndex.from_documents(documents)

SimpleDirectoryReader handles PDFs, Markdown, and plain text out of the box. For larger document sets, persist the index to disk so you don't rebuild embeddings on every run:

index.storage_context.persist(persist_dir="./storage")

And reload it later:

from llama_index.core import StorageContext, load_index_from_storage

storage_context = StorageContext.from_defaults(persist_dir="./storage")
index = load_index_from_storage(storage_context)

Step 5: Query with Claude as the generator

query_engine = index.as_query_engine(similarity_top_k=4)
response = query_engine.query(
    "What are the termination clauses in this contract?"
)
print(response)

LlamaIndex handles the retrieval, builds a prompt with the relevant chunks, and sends it to Claude through the LLM you configured in step 2. Because Claude handles long context well, you can raise similarity_top_k to pull in more chunks without worrying as much about truncation compared to smaller-context models — just watch your token costs as you increase it.

Step 6: Stream responses

For chat-style interfaces, streaming matters more than a single blocking response. LlamaIndex supports streaming query engines:

query_engine = index.as_query_engine(streaming=True)
response = query_engine.query("Summarize section 3")
response.print_response_stream()

If you're calling Claude through SubToAPI, streaming works the same way over SSE — see the streaming docs for the raw event format if you need to build a custom consumer outside LlamaIndex.

Step 7: Use a chat engine for multi-turn RAG

A plain query engine answers one question at a time with no memory. For conversational retrieval, use a chat engine:

chat_engine = index.as_chat_engine(
    chat_mode="context",
    system_prompt=(
        "You are a support assistant. Answer only using the "
        "provided context. If the answer isn't in the context, say so."
    ),
)

response = chat_engine.chat("How do I reset a user's password?")
print(response)

response = chat_engine.chat("What if they don't have access to email?")
print(response)

The context chat mode retrieves relevant chunks on every turn and keeps conversation history, which works well for Claude since it handles system prompts and multi-turn context reliably.

Adding tool use

If your RAG pipeline needs to call external functions — looking up order status, hitting an internal API — LlamaIndex's agent abstractions work with Claude's native tool use. Define a FunctionTool, pass it to a ReActAgent or FunctionCallingAgentWorker configured with the Anthropic LLM, and Claude will decide when to invoke it based on the conversation. For the raw tool-calling format Claude expects, the tools docs are a useful reference if you ever need to debug what's actually being sent over the wire.

Running this in production

A few things that matter once you move past a local script:

If you're setting up Claude access for the first time, the quickstart covers getting an API key and making your first request before you plug it into LlamaIndex, and pricing has the plan breakdown if you're evaluating options for a team.

Questions

Does LlamaIndex have official support for Claude? Yes. The llama-index-llms-anthropic package is maintained as part of the LlamaIndex integrations ecosystem and supports chat, streaming, and tool calling.

Can I use Claude for generation and a different provider for embeddings? Yes, and it's common. Embeddings and the LLM are configured independently in LlamaIndex's Settings object, so you can mix providers freely.

Do I need Anthropic's API directly, or can I route through a gateway? Either works. LlamaIndex's Anthropic client accepts a custom base_url, so you can point it at a gateway like SubToAPI if you want scoped application keys and shared usage tracking instead of a raw Anthropic key.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →