← Blog

Claude API Python Flask Backend Example

2026-10-08 · 5 min read · SubToAPI Team

If you're searching for a Claude API Python Flask backend example, you probably want a working server that accepts requests from a frontend, calls Claude, and returns a response — without hardcoding keys in client-side code. This article walks through exactly that: a minimal Flask app, a streaming version, error handling, and the production concerns that show up once real users hit the endpoint.

The short version: install flask and anthropic, create a route that forwards the request body to client.messages.create(), and return the result as JSON. Below is the full pattern, plus what changes when you move from a prototype to something you'd actually deploy.

Minimal Flask Backend with Claude

Install the dependencies:

pip install flask anthropic python-dotenv

Store your key in a .env file (never commit it):

ANTHROPIC_API_KEY=sk-ant-...

Then a basic app.py:

from flask import Flask, request, jsonify
from dotenv import load_dotenv
import anthropic
import os

load_dotenv()

app = Flask(__name__)
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

@app.route("/api/chat", methods=["POST"])
def chat():
    data = request.get_json(force=True)
    prompt = data.get("prompt", "")

    if not prompt:
        return jsonify({"error": "prompt is required"}), 400

    message = client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}]
    )

    return jsonify({"response": message.content[0].text})

if __name__ == "__main__":
    app.run(port=5000, debug=True)

This is enough for a local prototype: your frontend sends { "prompt": "..." } to /api/chat, Flask calls Claude, and you return the text. The key reason to put Flask in the middle rather than calling Claude from the browser is simple — API keys in client-side JavaScript are visible to anyone who opens dev tools.

Adding Streaming Responses

For chat-style UIs, users expect tokens to appear as they're generated rather than waiting for the full response. Flask supports this with a generator and Response with text/event-stream:

from flask import Response, stream_with_context

@app.route("/api/chat/stream", methods=["POST"])
def chat_stream():
    data = request.get_json(force=True)
    prompt = data.get("prompt", "")

    def generate():
        with client.messages.stream(
            model="claude-sonnet-4-20250514",
            max_tokens=1024,
            messages=[{"role": "user", "content": prompt}]
        ) as stream:
            for text in stream.text_stream:
                yield f"data: {text}\n\n"
        yield "event: done\ndata: end\n\n"

    return Response(stream_with_context(generate()), mimetype="text/event-stream")

On the frontend, consume this with EventSource or a fetch call that reads the response body as a stream. This is the pattern most production chat apps use instead of polling or waiting for a full response.

Handling Multi-Turn Conversations

A real backend needs to track conversation history, not just a single prompt. The simplest approach keeps the message array on the client and passes it through on each request:

@app.route("/api/chat", methods=["POST"])
def chat():
    data = request.get_json(force=True)
    messages = data.get("messages", [])

    if not messages:
        return jsonify({"error": "messages array is required"}), 400

    message = client.messages.create(
        model="claude-sonnet-4-20250514",
        max_tokens=1024,
        messages=messages
    )

    return jsonify({
        "response": message.content[0].text,
        "usage": {
            "input_tokens": message.usage.input_tokens,
            "output_tokens": message.usage.output_tokens
        }
    })

Returning usage data alongside the response matters once you need to track cost per user or per feature — it's easy to skip early on and painful to retrofit later.

Error Handling You'll Actually Need

Claude's API can return rate limit errors, overloaded errors, or malformed request errors. A production Flask route should catch these explicitly rather than letting a generic 500 leak to the client:

import anthropic

@app.route("/api/chat", methods=["POST"])
def chat():
    data = request.get_json(force=True)
    try:
        message = client.messages.create(
            model="claude-sonnet-4-20250514",
            max_tokens=1024,
            messages=data.get("messages", [])
        )
        return jsonify({"response": message.content[0].text})
    except anthropic.RateLimitError:
        return jsonify({"error": "rate limited, retry shortly"}), 429
    except anthropic.APIStatusError as e:
        return jsonify({"error": str(e)}), 502
    except Exception as e:
        return jsonify({"error": "unexpected error"}), 500

When a Direct SDK Integration Isn't Enough

The pattern above works fine for a single app with a single key. It gets harder once you have multiple environments, several developers, or a product that needs usage tracking per customer. At that point you're building key rotation, per-team quotas, and logging on top of your Flask app — infrastructure that has nothing to do with your actual product.

This is the gap SubToAPI fills. Instead of managing a raw Anthropic key inside your Flask app, you generate a scoped application key (sub_live_...) from the SubToAPI dashboard and point your existing code at it. The request shape stays nearly identical — you're still calling a messages endpoint with a model and a messages array — but you get streaming, tool use, usage metadata, and team seats without writing that layer yourself. Swapping it into the Flask example above is a matter of changing the base URL and the Authorization header; see /docs/quickstart and /docs/messages for the exact request format, and /docs/streaming if you're adapting the SSE route.

If you're already comfortable with the Flask + Anthropic SDK pattern, nothing here breaks that — SubToAPI sits behind an HTTPS API that looks the same shape, so your route handlers barely change. Plans start at Solo for €9/month, with Team at €19/seat and Scale at €49/seat for larger setups; details at /pricing. A free trial is available at /signup if you want to compare it against your current direct-key setup.

Deploying the Flask App

A few things to check before shipping:

Questions

Do I need Flask specifically, or does this work with FastAPI too? The pattern is the same — a route handler that calls the Claude client and returns JSON or a stream. FastAPI's async support makes streaming slightly cleaner, but Flask with stream_with_context works fine for most traffic levels.

Should I call Claude directly from my frontend instead of using a backend? No. Any API key embedded in frontend JavaScript is exposed to users. A backend like the Flask example here — or a gateway like SubToAPI — keeps keys server-side and lets you add rate limiting, logging, and auth in front of the model call.

How do I handle conversation history across requests? Store the message array either client-side (pass the full history each request) or server-side in a database keyed by session/user ID. For most chat UIs, passing the array back on each request is simpler and avoids needing session state in Flask.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →