← Blog

Claude API FastAPI Backend Integration Guide

2026-10-07 · 5 min read · SubToAPI Team

Integrating the Claude API into a FastAPI backend means building an async Python service that forwards requests to Claude, handles streaming responses, manages errors, and exposes a clean endpoint for your frontend or other services to consume. The core pattern is straightforward: FastAPI handles HTTP routing and validation, an async HTTP client talks to the Claude API, and you wrap the result in a response model your own clients can rely on.

This article walks through a working setup — project structure, async client calls, streaming with Server-Sent Events, error handling, and the operational details (rate limits, retries, API key management) that matter once this is running in production rather than a notebook.

Why FastAPI Fits Well Here

FastAPI's async-first design matches how LLM APIs behave: requests take seconds, not milliseconds, so you want non-blocking I/O while waiting for a response. FastAPI also gives you automatic request/response validation via Pydantic, which is useful because Claude's API has a specific message format you'll want to validate before sending (and before returning to your own clients).

The typical architecture:

Client → FastAPI endpoint → Claude API (or a compatible gateway) → FastAPI → Client

If you're building this for an internal tool or a product, you'll also want a layer that handles API key distribution, usage tracking per user, and rate limiting — things the raw Claude API doesn't give you out of the box. That's a reasonable use case for a gateway like SubToAPI sitting between FastAPI and the model, since it adds per-application API keys and usage metadata without you building that infrastructure yourself. More on that later; first, the integration itself.

Basic Setup

Install dependencies:

pip install fastapi uvicorn httpx python-dotenv

Project structure:

app/
  main.py
  claude_client.py
  models.py
  .env

Defining Request/Response Models

# models.py
from pydantic import BaseModel
from typing import List, Literal

class Message(BaseModel):
    role: Literal["user", "assistant"]
    content: str

class ChatRequest(BaseModel):
    messages: List[Message]
    model: str = "claude-sonnet-4"
    max_tokens: int = 1024

class ChatResponse(BaseModel):
    content: str
    model: str
    usage: dict

The Async Claude Client

Use httpx.AsyncClient so calls don't block the event loop:

# claude_client.py
import httpx
import os

API_URL = "https://api.subtoapi.app/v1/messages"
API_KEY = os.environ["SUBTOAPI_KEY"]

async def call_claude(messages: list, model: str, max_tokens: int) -> dict:
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
    }
    payload = {
        "model": model,
        "max_tokens": max_tokens,
        "messages": messages,
    }
    async with httpx.AsyncClient(timeout=60.0) as client:
        response = await client.post(API_URL, headers=headers, json=payload)
        response.raise_for_status()
        return response.json()

Keep the timeout generous — longer completions with larger max_tokens values take time, and a default 5-second httpx timeout will fail requests that would otherwise succeed.

The Endpoint

# main.py
from fastapi import FastAPI, HTTPException
from models import ChatRequest, ChatResponse
from claude_client import call_claude
import httpx

app = FastAPI()

@app.post("/chat", response_model=ChatResponse)
async def chat(request: ChatRequest):
    try:
        result = await call_claude(
            messages=[m.dict() for m in request.messages],
            model=request.model,
            max_tokens=request.max_tokens,
        )
    except httpx.HTTPStatusError as e:
        raise HTTPException(status_code=e.response.status_code, detail=e.response.text)
    except httpx.TimeoutException:
        raise HTTPException(status_code=504, detail="Upstream request timed out")

    return ChatResponse(
        content=result["content"][0]["text"],
        model=result["model"],
        usage=result.get("usage", {}),
    )

This gives you a self-contained /chat endpoint that validates input, forwards it, and returns a typed response. Run it with uvicorn main:app --reload and test with curl or your frontend.

Streaming Responses from FastAPI

For chat UIs, streaming matters — users shouldn't wait for the full completion before seeing text. FastAPI supports this via StreamingResponse, and you proxy the upstream stream chunk by chunk:

from fastapi.responses import StreamingResponse
import json

@app.post("/chat/stream")
async def chat_stream(request: ChatRequest):
    async def event_generator():
        async with httpx.AsyncClient(timeout=None) as client:
            async with client.stream(
                "POST",
                "https://api.subtoapi.app/v1/messages",
                headers={"Authorization": f"Bearer {os.environ['SUBTOAPI_KEY']}"},
                json={
                    "model": request.model,
                    "max_tokens": request.max_tokens,
                    "messages": [m.dict() for m in request.messages],
                    "stream": True,
                },
            ) as response:
                async for line in response.aiter_lines():
                    if line:
                        yield f"{line}\n\n"

    return StreamingResponse(event_generator(), media_type="text/event-stream")

Set timeout=None for the client used in streaming — a fixed timeout will cut off long-running generations mid-response. See the SubToAPI streaming docs for the exact event format if you're parsing chunks client-side rather than just relaying them.

Error Handling and Retries

Three failure modes you'll hit in production:

A minimal retry wrapper using tenacity:

from tenacity import retry, wait_exponential, stop_after_attempt

@retry(wait=wait_exponential(min=1, max=10), stop=stop_after_attempt(3))
async def call_claude_with_retry(*args, **kwargs):
    return await call_claude(*args, **kwargs)

Only retry on 429s and 5xx responses — retrying a 400 just wastes calls.

Managing API Keys and Multiple Clients

If this FastAPI service backs a product with multiple users or applications, don't share one Claude API key across everything. Instead, issue per-application keys so you can track usage, revoke access individually, and set team seats without redeploying code. SubToAPI handles this with sub_live_... keys you generate per app from a dashboard, which pairs well with the FastAPI pattern above — your backend just swaps the key per tenant rather than building key management from scratch. Check pricing if you're scoping this for a team.

Store keys in environment variables or a secrets manager, never in source control, and rotate them if a key is exposed in logs or client-side code by mistake.

Deployment Notes

For getting a working key and testing the full request/response cycle before wiring up FastAPI, start with the quickstart and the messages endpoint docs.

questions

Do I need to handle streaming manually, or can FastAPI do it automatically? FastAPI doesn't stream LLM output automatically — you need StreamingResponse with an async generator that relays chunks from the upstream API as they arrive, as shown above.

What's the biggest mistake people make integrating Claude into FastAPI? Using a short default timeout on the HTTP client. Claude completions can take well over 5–10 seconds depending on max_tokens, and the default timeouts in most HTTP libraries will cut requests off early.

Should I call the Claude API directly or through a gateway? Calling directly works fine for a single internal tool. For products with multiple users, teams, or apps, a gateway that issues scoped API keys and tracks usage per key — like SubToAPI — saves you from building that layer yourself.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →