← Blog

Claude API Python SDK Quickstart Example

2026-10-10 · 4 min read · SubToAPI Team

If you want to send your first request to Claude from Python in the next five minutes, here's the short version: install anthropic, set your API key as an environment variable, create a client, and call messages.create() with a model name, a max token limit, and a list of messages. That's the entire quickstart. The rest of this guide walks through each step with working code, plus streaming, error handling basics, and a few things that trip people up the first time.

This is the same pattern Anthropic's official SDK uses, and it's also the pattern you'd use against a Claude-compatible HTTPS endpoint like SubToAPI's, since the request shape mirrors the standard Messages API. If you're building something that needs to go into production with team seats, usage tracking, or a key that isn't tied to a personal Claude account, keep that in mind as you read — we'll touch on it near the end.

Install the SDK

pip install anthropic

This installs the official Python client, which wraps HTTP calls to the Claude API in a typed, synchronous (and async) interface.

Set your API key

The SDK looks for ANTHROPIC_API_KEY in your environment by default. On Linux/macOS:

export ANTHROPIC_API_KEY="your-key-here"

On Windows (PowerShell):

$env:ANTHROPIC_API_KEY="your-key-here"

Avoid hardcoding the key directly in your script. If you're committing code to a repo, use a .env file with python-dotenv and add .env to .gitignore.

Your first request

import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-opus-4-20250514",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Explain what a quicksort algorithm does in two sentences."}
    ]
)

print(response.content[0].text)

A few notes on what's happening here:

Adding a system prompt

response = client.messages.create(
    model="claude-opus-4-20250514",
    max_tokens=1024,
    system="You are a terse technical writer. Answer in bullet points only.",
    messages=[
        {"role": "user", "content": "What are the tradeoffs of using a message queue?"}
    ]
)

print(response.content[0].text)

The system field shapes behavior across the whole conversation and is kept separate from the conversational turns, which keeps your message history clean if you're logging or replaying it later.

Multi-turn conversations

Claude doesn't retain memory between calls — your Python code owns the conversation history. Each request needs the full list of prior turns:

messages = [
    {"role": "user", "content": "I'm building a REST API in Flask. Any naming conventions I should follow?"}
]

response = client.messages.create(
    model="claude-opus-4-20250514",
    max_tokens=1024,
    messages=messages
)

messages.append({"role": "assistant", "content": response.content[0].text})
messages.append({"role": "user", "content": "Now show me an example route following those conventions."})

response = client.messages.create(
    model="claude-opus-4-20250514",
    max_tokens=1024,
    messages=messages
)

print(response.content[0].text)

This is a common source of bugs: forgetting to append the assistant's reply before sending the next turn, which causes Claude to lose context on the second call.

Streaming responses

For chat UIs or long completions, streaming gives you tokens as they're generated instead of waiting for the full response:

with client.messages.stream(
    model="claude-opus-4-20250514",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Write a short poem about debugging."}]
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

The context manager handles connection cleanup automatically, even if you break out of the loop early.

Basic error handling

The SDK raises typed exceptions you can catch directly:

from anthropic import APIStatusError, APIConnectionError

try:
    response = client.messages.create(
        model="claude-opus-4-20250514",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello"}]
    )
except APIConnectionError:
    print("Network issue — check connectivity and retry")
except APIStatusError as e:
    print(f"API returned an error: {e.status_code} - {e.message}")

For anything beyond a quick script, you'll want retry logic with backoff on rate limits and transient 5xx errors — but that's a separate, deeper topic worth its own guide.

Async usage

If your app is already async (FastAPI, aiohttp), use AsyncAnthropic instead:

import asyncio
from anthropic import AsyncAnthropic

client = AsyncAnthropic()

async def main():
    response = await client.messages.create(
        model="claude-opus-4-20250514",
        max_tokens=1024,
        messages=[{"role": "user", "content": "Summarize REST vs GraphQL in one paragraph."}]
    )
    print(response.content[0].text)

asyncio.run(main())

Going from a script to a shared service

This quickstart covers a single developer running a script with one API key. It breaks down quickly once a team gets involved: shared keys get copy-pasted into Slack, there's no per-person usage visibility, and billing is one lump charge with no breakdown.

If you've outgrown the single-key setup and need application-scoped keys, usage metadata, and team seats without re-architecting your Python code, SubToAPI exposes the same Messages-style endpoint your SDK already expects. You generate sub_live_... keys per app or environment, point your existing client at https://api.subtoapi.app/v1/messages, and keep the same request/response shape — streaming and tool use included. Check /docs/quickstart for the setup, or /pricing if you're comparing Solo, Team, and Scale seat-based plans. There's a free trial at /signup if you want to try it against a real project before committing.

Questions

Do I need the anthropic package, or can I just use requests? You can call the Claude API with plain HTTP and requests, but the official SDK handles retries, typed responses, and streaming parsing for you. Use requests only if you need to avoid the dependency for a minimal environment.

Why does my request fail with a "max_tokens required" error? max_tokens is a mandatory parameter in the Messages API — it defines the maximum length of Claude's reply, not the input. Set it explicitly on every call.

Can I use this same Python code against SubToAPI instead of Anthropic directly? Yes — point the base_url at https://api.subtoapi.app/v1 and use a sub_live_... key as your bearer token. The request and response format follow the same Messages structure, so existing code needs minimal changes. See /docs/messages for field details.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →