Claude API Python Flask Backend Example
If you're searching for a Claude API Python Flask backend example, you probably want a working server that accepts requests from a frontend, calls Claude, and returns a response — without hardcoding keys in client-side code. This article walks through exactly that: a minimal Flask app, a streaming version, error handling, and the production concerns that show up once real users hit the endpoint.
The short version: install flask and anthropic, create a route that forwards the request body to client.messages.create(), and return the result as JSON. Below is the full pattern, plus what changes when you move from a prototype to something you'd actually deploy.
Minimal Flask Backend with Claude
Install the dependencies:
pip install flask anthropic python-dotenv
Store your key in a .env file (never commit it):
ANTHROPIC_API_KEY=sk-ant-...
Then a basic app.py:
from flask import Flask, request, jsonify
from dotenv import load_dotenv
import anthropic
import os
load_dotenv()
app = Flask(__name__)
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
@app.route("/api/chat", methods=["POST"])
def chat():
data = request.get_json(force=True)
prompt = data.get("prompt", "")
if not prompt:
return jsonify({"error": "prompt is required"}), 400
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
)
return jsonify({"response": message.content[0].text})
if __name__ == "__main__":
app.run(port=5000, debug=True)
This is enough for a local prototype: your frontend sends { "prompt": "..." } to /api/chat, Flask calls Claude, and you return the text. The key reason to put Flask in the middle rather than calling Claude from the browser is simple — API keys in client-side JavaScript are visible to anyone who opens dev tools.
Adding Streaming Responses
For chat-style UIs, users expect tokens to appear as they're generated rather than waiting for the full response. Flask supports this with a generator and Response with text/event-stream:
from flask import Response, stream_with_context
@app.route("/api/chat/stream", methods=["POST"])
def chat_stream():
data = request.get_json(force=True)
prompt = data.get("prompt", "")
def generate():
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}]
) as stream:
for text in stream.text_stream:
yield f"data: {text}\n\n"
yield "event: done\ndata: end\n\n"
return Response(stream_with_context(generate()), mimetype="text/event-stream")
On the frontend, consume this with EventSource or a fetch call that reads the response body as a stream. This is the pattern most production chat apps use instead of polling or waiting for a full response.
Handling Multi-Turn Conversations
A real backend needs to track conversation history, not just a single prompt. The simplest approach keeps the message array on the client and passes it through on each request:
@app.route("/api/chat", methods=["POST"])
def chat():
data = request.get_json(force=True)
messages = data.get("messages", [])
if not messages:
return jsonify({"error": "messages array is required"}), 400
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=messages
)
return jsonify({
"response": message.content[0].text,
"usage": {
"input_tokens": message.usage.input_tokens,
"output_tokens": message.usage.output_tokens
}
})
Returning usage data alongside the response matters once you need to track cost per user or per feature — it's easy to skip early on and painful to retrofit later.
Error Handling You'll Actually Need
Claude's API can return rate limit errors, overloaded errors, or malformed request errors. A production Flask route should catch these explicitly rather than letting a generic 500 leak to the client:
import anthropic
@app.route("/api/chat", methods=["POST"])
def chat():
data = request.get_json(force=True)
try:
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=data.get("messages", [])
)
return jsonify({"response": message.content[0].text})
except anthropic.RateLimitError:
return jsonify({"error": "rate limited, retry shortly"}), 429
except anthropic.APIStatusError as e:
return jsonify({"error": str(e)}), 502
except Exception as e:
return jsonify({"error": "unexpected error"}), 500
When a Direct SDK Integration Isn't Enough
The pattern above works fine for a single app with a single key. It gets harder once you have multiple environments, several developers, or a product that needs usage tracking per customer. At that point you're building key rotation, per-team quotas, and logging on top of your Flask app — infrastructure that has nothing to do with your actual product.
This is the gap SubToAPI fills. Instead of managing a raw Anthropic key inside your Flask app, you generate a scoped application key (sub_live_...) from the SubToAPI dashboard and point your existing code at it. The request shape stays nearly identical — you're still calling a messages endpoint with a model and a messages array — but you get streaming, tool use, usage metadata, and team seats without writing that layer yourself. Swapping it into the Flask example above is a matter of changing the base URL and the Authorization header; see /docs/quickstart and /docs/messages for the exact request format, and /docs/streaming if you're adapting the SSE route.
If you're already comfortable with the Flask + Anthropic SDK pattern, nothing here breaks that — SubToAPI sits behind an HTTPS API that looks the same shape, so your route handlers barely change. Plans start at Solo for €9/month, with Team at €19/seat and Scale at €49/seat for larger setups; details at /pricing. A free trial is available at /signup if you want to compare it against your current direct-key setup.
Deploying the Flask App
A few things to check before shipping:
- Set
debug=Falseand run behindgunicornoruwsgi, not the Flask dev server. - Store the API key in environment variables on the host, not in the repo.
- Add request timeouts and a reasonable
max_tokenscap to avoid runaway costs from malformed input. - Log token usage per request if you need to bill or budget by customer.
- Add CORS headers (
flask-cors) if your frontend is on a different origin.
Questions
Do I need Flask specifically, or does this work with FastAPI too? The pattern is the same — a route handler that calls the Claude client and returns JSON or a stream. FastAPI's async support makes streaming slightly cleaner, but Flask with stream_with_context works fine for most traffic levels.
Should I call Claude directly from my frontend instead of using a backend? No. Any API key embedded in frontend JavaScript is exposed to users. A backend like the Flask example here — or a gateway like SubToAPI — keeps keys server-side and lets you add rate limiting, logging, and auth in front of the model call.
How do I handle conversation history across requests? Store the message array either client-side (pass the full history each request) or server-side in a database keyed by session/user ID. For most chat UIs, passing the array back on each request is simpler and avoids needing session state in Flask.