Claude API FastAPI Backend Setup Guide
Setting up a FastAPI backend to call the Claude API means wiring three things correctly: an async HTTP client that talks to the Messages endpoint, environment-based key management, and a response layer that either returns JSON or streams tokens back to your frontend. This guide walks through a minimal but production-ready setup you can extend.
If you just want a working endpoint in the next ten minutes, skip to the "Basic setup" section below. If you're trying to decide between calling Claude directly or going through a proxy layer, the later section on production considerations covers that tradeoff.
Project structure
A clean FastAPI project for an LLM backend typically looks like this:
app/
main.py
config.py
clients/
claude.py
routers/
chat.py
requirements.txt
.env
Keep the Claude client isolated in its own module. This matters more than it sounds — once you add retries, timeouts, or switch providers, you want one place to change.
Install dependencies:
pip install fastapi uvicorn httpx python-dotenv pydantic
httpx is the right choice here because it supports async requests and streaming natively, which requests does not.
Environment and config
Store your API key outside the codebase:
# .env
CLAUDE_API_KEY=sk-ant-xxxxxxxx
CLAUDE_API_URL=https://api.anthropic.com/v1/messages
# app/config.py
import os
from dotenv import load_dotenv
load_dotenv()
CLAUDE_API_KEY = os.environ["CLAUDE_API_KEY"]
CLAUDE_API_URL = os.getenv("CLAUDE_API_URL", "https://api.anthropic.com/v1/messages")
Fail fast if the key is missing — using os.environ[...] instead of .get() will raise a KeyError at startup rather than a confusing 401 on first request.
Basic setup: a single chat endpoint
Here's a minimal non-streaming endpoint:
# app/clients/claude.py
import httpx
from app.config import CLAUDE_API_KEY, CLAUDE_API_URL
HEADERS = {
"x-api-key": CLAUDE_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
}
async def send_message(messages: list, model: str = "claude-sonnet-4-5", max_tokens: int = 1024):
payload = {
"model": model,
"max_tokens": max_tokens,
"messages": messages,
}
async with httpx.AsyncClient(timeout=60.0) as client:
resp = await client.post(CLAUDE_API_URL, headers=HEADERS, json=payload)
resp.raise_for_status()
return resp.json()
# app/routers/chat.py
from fastapi import APIRouter, HTTPException
from pydantic import BaseModel
from app.clients.claude import send_message
router = APIRouter()
class ChatRequest(BaseModel):
prompt: str
@router.post("/chat")
async def chat(req: ChatRequest):
try:
result = await send_message([{"role": "user", "content": req.prompt}])
except httpx.HTTPStatusError as e:
raise HTTPException(status_code=502, detail=str(e))
return {"reply": result["content"][0]["text"]}
# app/main.py
from fastapi import FastAPI
from app.routers import chat
app = FastAPI()
app.include_router(chat.router)
Run it with uvicorn app.main:app --reload and test:
curl -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"prompt": "Explain async/await in one sentence"}'
Streaming responses from FastAPI
Chat UIs feel broken without streaming. FastAPI's StreamingResponse combined with httpx's async streaming makes this straightforward:
from fastapi.responses import StreamingResponse
import json
async def stream_claude(messages: list):
payload = {
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": messages,
"stream": True,
}
async with httpx.AsyncClient(timeout=None) as client:
async with client.stream("POST", CLAUDE_API_URL, headers=HEADERS, json=payload) as resp:
async for line in resp.aiter_lines():
if line.startswith("data:"):
yield line[5:].strip() + "\n"
@router.post("/chat/stream")
async def chat_stream(req: ChatRequest):
messages = [{"role": "user", "content": req.prompt}]
return StreamingResponse(stream_claude(messages), media_type="text/event-stream")
On the frontend, consume this with EventSource or a fetch reader. The key detail is timeout=None on the client — a default timeout will kill long-running streams mid-response.
Error handling and retries
Claude's API returns standard HTTP status codes: 429 for rate limits, 529 for overload, 401 for bad keys. A production backend should retry transient errors with backoff:
import asyncio
async def send_message_with_retry(messages: list, retries: int = 3):
for attempt in range(retries):
try:
return await send_message(messages)
except httpx.HTTPStatusError as e:
if e.response.status_code in (429, 529) and attempt < retries - 1:
await asyncio.sleep(2 ** attempt)
continue
raise
Don't swallow errors silently — log the status code and request ID from Anthropic's response headers so you can trace issues later.
Multiple apps, one Claude account
If you're building more than one service on top of your Claude access — an internal tool plus a customer-facing app, for example — managing a single raw API key across all of them gets messy fast. There's no per-app usage breakdown, no way to revoke one integration without breaking the others, and no built-in way to add teammates without sharing the actual key.
This is where SubToAPI fits into a FastAPI setup: it sits between your backend and Claude, issuing scoped sub_live_... keys per application while you keep using the same underlying access. Your FastAPI client code barely changes — swap the base URL and header:
HEADERS = {
"Authorization": f"Bearer {SUBTOAPI_KEY}",
"content-type": "application/json",
}
CLAUDE_API_URL = "https://api.subtoapi.app/v1/messages"
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "claude-sonnet-4-5", "max_tokens": 1024, "messages": [{"role":"user","content":"hello"}]}'
You get streaming, tool use, and usage metadata per key out of the box — check the quickstart and streaming docs for the full request format. Pricing starts at €9/month on the Solo plan; see /pricing for Team and Scale tiers if you need shared seats.
Deployment checklist
Before shipping your FastAPI + Claude backend:
- Set request timeouts explicitly on every
httpxclient - Never log full request/response bodies containing user data
- Rate-limit your own
/chatendpoint to avoid surprise Claude bills - Use
async defconsistently — mixing sync blocking calls into async routes will stall your event loop - Validate
max_tokensserver-side; don't trust client input directly
questions
Do I need a special library to call Claude from FastAPI? No. httpx with async support is enough — Anthropic's API is plain REST over HTTPS, so no SDK is required, though the official Python SDK works fine too if you prefer it.
How do I stream Claude responses through FastAPI to a browser? Use StreamingResponse with an async generator that reads server-sent events from Claude's streaming endpoint via httpx.AsyncClient.stream(), then forward each chunk as it arrives.
Can I run Claude API calls inside FastAPI background tasks? Yes, for non-interactive work like batch summarization. For chat UIs, prefer the streaming approach above since background tasks don't return partial results to the original request.