Claude API FastAPI Backend Integration Guide
Integrating the Claude API into a FastAPI backend means building an async Python service that forwards requests to Claude, handles streaming responses, manages errors, and exposes a clean endpoint for your frontend or other services to consume. The core pattern is straightforward: FastAPI handles HTTP routing and validation, an async HTTP client talks to the Claude API, and you wrap the result in a response model your own clients can rely on.
This article walks through a working setup — project structure, async client calls, streaming with Server-Sent Events, error handling, and the operational details (rate limits, retries, API key management) that matter once this is running in production rather than a notebook.
Why FastAPI Fits Well Here
FastAPI's async-first design matches how LLM APIs behave: requests take seconds, not milliseconds, so you want non-blocking I/O while waiting for a response. FastAPI also gives you automatic request/response validation via Pydantic, which is useful because Claude's API has a specific message format you'll want to validate before sending (and before returning to your own clients).
The typical architecture:
Client → FastAPI endpoint → Claude API (or a compatible gateway) → FastAPI → Client
If you're building this for an internal tool or a product, you'll also want a layer that handles API key distribution, usage tracking per user, and rate limiting — things the raw Claude API doesn't give you out of the box. That's a reasonable use case for a gateway like SubToAPI sitting between FastAPI and the model, since it adds per-application API keys and usage metadata without you building that infrastructure yourself. More on that later; first, the integration itself.
Basic Setup
Install dependencies:
pip install fastapi uvicorn httpx python-dotenv
Project structure:
app/
main.py
claude_client.py
models.py
.env
Defining Request/Response Models
# models.py
from pydantic import BaseModel
from typing import List, Literal
class Message(BaseModel):
role: Literal["user", "assistant"]
content: str
class ChatRequest(BaseModel):
messages: List[Message]
model: str = "claude-sonnet-4"
max_tokens: int = 1024
class ChatResponse(BaseModel):
content: str
model: str
usage: dict
The Async Claude Client
Use httpx.AsyncClient so calls don't block the event loop:
# claude_client.py
import httpx
import os
API_URL = "https://api.subtoapi.app/v1/messages"
API_KEY = os.environ["SUBTOAPI_KEY"]
async def call_claude(messages: list, model: str, max_tokens: int) -> dict:
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"model": model,
"max_tokens": max_tokens,
"messages": messages,
}
async with httpx.AsyncClient(timeout=60.0) as client:
response = await client.post(API_URL, headers=headers, json=payload)
response.raise_for_status()
return response.json()
Keep the timeout generous — longer completions with larger max_tokens values take time, and a default 5-second httpx timeout will fail requests that would otherwise succeed.
The Endpoint
# main.py
from fastapi import FastAPI, HTTPException
from models import ChatRequest, ChatResponse
from claude_client import call_claude
import httpx
app = FastAPI()
@app.post("/chat", response_model=ChatResponse)
async def chat(request: ChatRequest):
try:
result = await call_claude(
messages=[m.dict() for m in request.messages],
model=request.model,
max_tokens=request.max_tokens,
)
except httpx.HTTPStatusError as e:
raise HTTPException(status_code=e.response.status_code, detail=e.response.text)
except httpx.TimeoutException:
raise HTTPException(status_code=504, detail="Upstream request timed out")
return ChatResponse(
content=result["content"][0]["text"],
model=result["model"],
usage=result.get("usage", {}),
)
This gives you a self-contained /chat endpoint that validates input, forwards it, and returns a typed response. Run it with uvicorn main:app --reload and test with curl or your frontend.
Streaming Responses from FastAPI
For chat UIs, streaming matters — users shouldn't wait for the full completion before seeing text. FastAPI supports this via StreamingResponse, and you proxy the upstream stream chunk by chunk:
from fastapi.responses import StreamingResponse
import json
@app.post("/chat/stream")
async def chat_stream(request: ChatRequest):
async def event_generator():
async with httpx.AsyncClient(timeout=None) as client:
async with client.stream(
"POST",
"https://api.subtoapi.app/v1/messages",
headers={"Authorization": f"Bearer {os.environ['SUBTOAPI_KEY']}"},
json={
"model": request.model,
"max_tokens": request.max_tokens,
"messages": [m.dict() for m in request.messages],
"stream": True,
},
) as response:
async for line in response.aiter_lines():
if line:
yield f"{line}\n\n"
return StreamingResponse(event_generator(), media_type="text/event-stream")
Set timeout=None for the client used in streaming — a fixed timeout will cut off long-running generations mid-response. See the SubToAPI streaming docs for the exact event format if you're parsing chunks client-side rather than just relaying them.
Error Handling and Retries
Three failure modes you'll hit in production:
- Rate limits (429) — back off and retry with jitter, don't retry instantly
- Timeouts — common with long completions behind slow networks; increase client timeouts rather than retrying immediately
- Malformed input (400) — validate message roles and content length before sending, since Pydantic validation won't catch everything Claude's API rejects
A minimal retry wrapper using tenacity:
from tenacity import retry, wait_exponential, stop_after_attempt
@retry(wait=wait_exponential(min=1, max=10), stop=stop_after_attempt(3))
async def call_claude_with_retry(*args, **kwargs):
return await call_claude(*args, **kwargs)
Only retry on 429s and 5xx responses — retrying a 400 just wastes calls.
Managing API Keys and Multiple Clients
If this FastAPI service backs a product with multiple users or applications, don't share one Claude API key across everything. Instead, issue per-application keys so you can track usage, revoke access individually, and set team seats without redeploying code. SubToAPI handles this with sub_live_... keys you generate per app from a dashboard, which pairs well with the FastAPI pattern above — your backend just swaps the key per tenant rather than building key management from scratch. Check pricing if you're scoping this for a team.
Store keys in environment variables or a secrets manager, never in source control, and rotate them if a key is exposed in logs or client-side code by mistake.
Deployment Notes
- Run with multiple Uvicorn workers (
--workers 4) behind Gunicorn or directly with Uvicorn's worker flag for production traffic - Set explicit
httpxconnection pool limits if you're handling high concurrency — the default pool size can bottleneck under load - Log request IDs and latency per call so you can debug slow completions separately from network issues
- If streaming, make sure your reverse proxy (nginx, etc.) doesn't buffer the response — buffering defeats the purpose of SSE
For getting a working key and testing the full request/response cycle before wiring up FastAPI, start with the quickstart and the messages endpoint docs.
questions
Do I need to handle streaming manually, or can FastAPI do it automatically? FastAPI doesn't stream LLM output automatically — you need StreamingResponse with an async generator that relays chunks from the upstream API as they arrive, as shown above.
What's the biggest mistake people make integrating Claude into FastAPI? Using a short default timeout on the HTTP client. Claude completions can take well over 5–10 seconds depending on max_tokens, and the default timeouts in most HTTP libraries will cut requests off early.
Should I call the Claude API directly or through a gateway? Calling directly works fine for a single internal tool. For products with multiple users, teams, or apps, a gateway that issues scoped API keys and tracks usage per key — like SubToAPI — saves you from building that layer yourself.