Claude API Python SDK Quickstart Guide
If you want to send your first request to Claude from Python in the next five minutes, here's the short version: install the anthropic package, set your API key as an environment variable, create a client, and call messages.create() with a model, a max token limit, and a list of messages. That's the entire quickstart.
This guide walks through each step in detail — installation, authentication, making your first call, streaming responses, handling errors, and a couple of patterns you'll actually use in production code. It also covers an alternative setup if you want a single HTTPS API key that works across tools without juggling SDK versions per language.
Prerequisites
Before you start, you need:
- Python 3.8 or newer
- An API key (either directly from Anthropic, or from a service that proxies Claude access)
pipinstalled and working
Check your Python version first:
python3 --version
Step 1: Install the SDK
Install the official Python SDK with pip:
pip install anthropic
If you're working inside a virtual environment (recommended), activate it first:
python3 -m venv venv
source venv/bin/activate
pip install anthropic
Step 2: Set Your API Key
Never hardcode API keys in your source files. Set the key as an environment variable instead:
export ANTHROPIC_API_KEY="your-key-here"
The SDK automatically reads this variable, so you don't need to pass it explicitly in code. For local development, a .env file with python-dotenv works well too:
pip install python-dotenv
from dotenv import load_dotenv
load_dotenv()
Step 3: Your First Request
Here's the minimal code to send a message and get a response:
import anthropic
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
messages=[
{"role": "user", "content": "Explain what a quicksort algorithm does in two sentences."}
]
)
print(message.content[0].text)
A few things worth noting:
max_tokensis required and caps the length of the response, not the input.messagesis a list of role/content dictionaries —userandassistantroles alternate for multi-turn conversations.- The response object contains a
contentlist (usually one text block), plususagedata showing input and output token counts.
Step 4: Add a System Prompt
System prompts set behavior and context separately from the conversation itself:
message = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
system="You are a concise technical writer. Answer in bullet points only.",
messages=[
{"role": "user", "content": "What are the benefits of connection pooling?"}
]
)
Keep system prompts stable across requests — they're a good place for persona, output format rules, and constraints that shouldn't change per-message.
Step 5: Stream Responses
For chat interfaces or anything latency-sensitive, streaming token-by-token output feels far more responsive than waiting for the full response:
with client.messages.stream(
model="claude-sonnet-4-5",
max_tokens=1024,
messages=[
{"role": "user", "content": "Write a short poem about debugging."}
]
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
The stream() context manager handles the SSE parsing for you — you just consume text_stream chunk by chunk.
Step 6: Handle Errors Properly
Production code needs to handle rate limits, timeouts, and invalid requests gracefully:
import anthropic
client = anthropic.Anthropic()
try:
message = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}]
)
except anthropic.RateLimitError:
print("Rate limited — back off and retry")
except anthropic.APIConnectionError:
print("Network issue — retry with backoff")
except anthropic.APIStatusError as e:
print(f"API error: {e.status_code} - {e.message}")
The SDK raises typed exceptions, which makes it easy to branch logic (retry vs. fail vs. log) without parsing raw HTTP status codes manually.
Step 7: Multi-Turn Conversations
To maintain context, append each exchange to the messages list and resend the whole history:
conversation = [
{"role": "user", "content": "What's the capital of France?"}
]
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=512,
messages=conversation
)
conversation.append({"role": "assistant", "content": response.content[0].text})
conversation.append({"role": "user", "content": "What's its population?"})
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=512,
messages=conversation
)
Claude has no memory between calls — the full conversation history has to be sent each time.
An Alternative: One API Key Across Tools
The official SDK is the right choice if you're building directly against Anthropic's infrastructure and want full control over SDK versions, model access, and rate limit tiers. But if you're already paying for Claude through a subscription and want to call it from scripts, internal tools, or multiple projects without separate billing accounts, SubToAPI turns that access into a standard HTTPS API with a sub_live_... key.
The request shape is nearly identical to what you've just written, just pointed at a different base URL:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Claude"}]
}'
It supports streaming, tool use, and usage metadata, with team seats if you need shared access across a project. Check the quickstart and messages docs for details, or see pricing — plans start at Solo (€9), Team (€19/seat), and Scale (€49/seat), with a free trial at signup.
Next Steps
Once the basics are working, look into:
- Streaming for real-time output in chat UIs
- Tool use for letting Claude call functions and return structured results
- Setting per-request
temperatureandtop_pvalues to tune output variability - Logging
usage.input_tokensandusage.output_tokensfrom every response for cost tracking
questions
Do I need the official Anthropic SDK, or can I just use requests? The SDK isn't strictly required — Claude's API is plain HTTPS/JSON, so requests or httpx works fine. The SDK just saves you from writing retry logic, streaming parsers, and typed error handling yourself.
What's the difference between max_tokens and the model's context window? max_tokens limits only the length of the response Claude generates. The context window is the total space available for your input plus that response, and it's much larger — exceeding it returns an error, not a truncated reply.
Can I use the Python SDK with an API key from a service other than Anthropic directly? Yes, as long as the service exposes the same request/response format at its own base URL. You typically just change the base_url parameter and the key — the rest of your code, including message formatting and streaming, stays the same.