Claude API Python SDK Quickstart Guide
If you're searching for a Claude API Python SDK quickstart, you want the shortest path from an empty terminal to a working script that sends a prompt and gets a response back. This guide covers exactly that: installing the anthropic package, authenticating, making your first call, streaming output, and handling the errors you'll actually run into in production.
By the end you'll have a working Python script, understand the shape of the Messages API, and know where it gets complicated — rate limits, retries, multi-user key management — and how to avoid building that infrastructure yourself.
Installing the SDK
The official Python SDK is published as anthropic on PyPI. Install it with pip:
pip install anthropic
It requires Python 3.8+. If you're using a virtual environment (recommended), activate it first:
python -m venv venv
source venv/bin/activate
pip install anthropic
Setting your API key
The SDK reads your key from the ANTHROPIC_API_KEY environment variable by default:
export ANTHROPIC_API_KEY="sk-ant-..."
On Windows (PowerShell):
$env:ANTHROPIC_API_KEY="sk-ant-..."
You can also pass the key explicitly when constructing the client, which is useful in testing or when you're managing multiple keys:
from anthropic import Anthropic
client = Anthropic(api_key="sk-ant-...")
Never hardcode the key in source control. Use environment variables, a .env file with python-dotenv, or a secrets manager.
Your first request
Here's the minimal script to send a message and print the reply:
from anthropic import Anthropic
client = Anthropic()
message = client.messages.create(
model="claude-opus-4-20250514",
max_tokens=1024,
messages=[
{"role": "user", "content": "Explain what a quicksort algorithm does in two sentences."}
],
)
print(message.content[0].text)
Run it:
python quickstart.py
A few things worth noting about the request shape:
modelis required — check Anthropic's docs for the current model identifiers, since names change over time.max_tokenscaps the length of the response and is also required.messagesis a list of role/content turns, alternatinguserandassistant. There's no separate "system prompt" field in the messages list itself — use the top-levelsystemparameter for that.
Adding a system prompt
message = client.messages.create(
model="claude-opus-4-20250514",
max_tokens=1024,
system="You are a terse, senior backend engineer. Answer in bullet points.",
messages=[
{"role": "user", "content": "What are the tradeoffs of using UUIDs as primary keys?"}
],
)
Multi-turn conversations
Since the API is stateless, you maintain conversation history yourself by appending to the messages list:
conversation = [
{"role": "user", "content": "What's the capital of Portugal?"}
]
reply = client.messages.create(
model="claude-opus-4-20250514",
max_tokens=512,
messages=conversation,
)
conversation.append({"role": "assistant", "content": reply.content[0].text})
conversation.append({"role": "user", "content": "What's its population?"})
reply2 = client.messages.create(
model="claude-opus-4-20250514",
max_tokens=512,
messages=conversation,
)
Streaming responses
For chat UIs or anything latency-sensitive, stream tokens instead of waiting for the full response:
with client.messages.stream(
model="claude-opus-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Write a haiku about distributed systems."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
The stream() context manager handles the server-sent events for you and exposes text_stream as a simple iterator.
Handling errors
The SDK raises typed exceptions you should catch explicitly:
from anthropic import Anthropic, APIStatusError, APIConnectionError, RateLimitError
client = Anthropic()
try:
message = client.messages.create(
model="claude-opus-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
except RateLimitError:
print("Rate limited — back off and retry")
except APIConnectionError:
print("Network issue reaching the API")
except APIStatusError as e:
print(f"API returned an error: {e.status_code} {e.message}")
In production you'll want retry logic with exponential backoff around RateLimitError and transient APIConnectionErrors, since rate limits are the most common failure mode once you have real traffic.
Where this gets harder
The quickstart above works fine for a single script or prototype. It gets more complicated once you need to:
- Issue separate API keys per customer or environment, without sharing your root Anthropic key
- Track token usage and cost per application or team, not just globally
- Add multiple team members with their own keys and usage visibility
- Put a stable HTTPS endpoint in front of your Claude usage that doesn't change when you rotate credentials
This is the gap SubToAPI fills. It sits on top of your existing Claude access and gives you application-scoped keys (sub_live_...) that work with the same Messages-style requests shown above, plus a dashboard for usage metadata, streaming, tool use, and team seats. If you're building something that other people or services will call — not just running personal scripts — it saves you from rebuilding key management and billing visibility from scratch.
To try it with the same request shape as the official SDK, a basic call looks like:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-20250514",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello"}]
}'
See the quickstart and messages docs for the full request reference, streaming for SSE details, and tools for function-calling support. Plans start at Solo €9/month, with Team and Scale tiers for multi-seat usage — see pricing — and you can start with a free trial at signup.
FAQs
Do I need the official SDK to use Claude from Python, or can I just use requests? The SDK isn't required — you can call the HTTP API directly with requests or httpx. The SDK just saves you from writing serialization, retry, and streaming boilerplate yourself.
What's the difference between max_tokens and the model's context window? max_tokens limits only the length of the generated response. The context window is the total budget for input plus output combined, and varies by model.
Can I use the Python SDK for streaming and tool use together? Yes. You can pass a tools parameter alongside stream(), and the SDK will emit tool-use content blocks as part of the streamed events, which you handle the same way as a non-streamed tool call.