Claude API SDK Python Quickstart Guide
If you're searching for a Claude API SDK Python quickstart guide, you want two things: a working script you can copy, and a clear understanding of what each line does so you're not stuck debugging an opaque error on your first attempt. This guide covers both — installing the official anthropic Python package, authenticating, sending your first message, streaming responses, and handling the errors you'll actually hit in production.
By the end, you'll have a script that sends a prompt to Claude and prints the response, plus enough context to extend it into a real feature: chat history, streaming output, tool calls, and usage tracking.
Prerequisites
You need three things before writing any code:
- Python 3.8 or newer — check with
python3 --version - An API key — either directly from Anthropic's console, or a
sub_live_...key from SubToAPI if you want usage dashboards, team seats, and a unified billing view without managing Anthropic console access per developer - pip for installing the SDK
Installing the SDK
pip install anthropic
That's the entire install step. The package is lightweight and has no heavy dependencies — it wraps HTTP calls, request retries, and response typing so you don't have to hand-roll JSON parsing.
Setting your API key
Never hardcode your key in source files. Set it as an environment variable:
export ANTHROPIC_API_KEY="sk-ant-..."
On Windows (PowerShell):
$env:ANTHROPIC_API_KEY="sk-ant-..."
The SDK automatically reads this variable, so you don't need to pass the key explicitly in code — though you can if you're managing multiple keys or environments.
Your first request
Here's a minimal script that sends a message and prints the reply:
from anthropic import Anthropic
client = Anthropic() # reads ANTHROPIC_API_KEY from environment
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[
{"role": "user", "content": "Explain what a quickstart guide should include."}
]
)
print(message.content[0].text)
Run it with python3 quickstart.py. If your key is valid, you'll get a text response printed to your terminal within a couple of seconds.
A few things worth noting:
max_tokensis required — it caps the length of the response, not the inputmessagesis a list, so multi-turn conversations just append more{"role": ..., "content": ...}entriesmessage.contentis a list of content blocks, since responses can include text, tool calls, or other block types —[0].textgrabs the first text block
Adding a system prompt
System prompts set behavior and context separately from the conversation itself:
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
system="You are a concise technical writer. Answer in bullet points.",
messages=[
{"role": "user", "content": "List three benefits of using an SDK over raw HTTP calls."}
]
)
Keep system prompts stable across requests — this is also what makes prompt caching effective if you're sending the same instructions repeatedly.
Streaming responses
For chat interfaces or anything latency-sensitive, stream tokens as they're generated instead of waiting for the full response:
with client.messages.stream(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Write a short poem about debugging."}]
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
This prints each chunk as it arrives, which is what you want for anything user-facing where perceived speed matters.
Handling errors properly
Production code needs to handle rate limits, timeouts, and invalid requests without crashing. The SDK raises typed exceptions you can catch:
from anthropic import Anthropic, APIStatusError, APIConnectionError, RateLimitError
client = Anthropic()
try:
message = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello Claude"}]
)
print(message.content[0].text)
except RateLimitError:
print("Rate limited — back off and retry")
except APIConnectionError:
print("Network issue — check connectivity")
except APIStatusError as e:
print(f"API returned an error: {e.status_code} {e.message}")
Catching these three categories separately matters because the right response differs: rate limits need backoff, connection errors need retries, and status errors (like a malformed request) need you to fix the code, not retry blindly.
Where SubToAPI fits in
The script above talks directly to Anthropic's API. If you're building a product on top of Claude rather than a personal script, you'll eventually need things the raw SDK doesn't give you: per-application API keys instead of one shared secret, usage breakdowns by key, and seat-based access for a team without sharing a single console login.
SubToAPI sits in front of the same Claude models and gives you sub_live_... keys that work with the same request shape — messages, streaming, and tool use all behave the same way, just routed through https://api.subtoapi.app/v1/messages with your SubToAPI key in the Authorization header. If you started with the quickstart above and now need to issue separate keys per app or teammate, check the quickstart docs and messages API reference — the request and response format will already look familiar.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello from SubToAPI"}]
}'
Plans start at €9/month for solo use, with team and scale tiers for multi-seat setups — see pricing or start with a free trial at signup.
Next steps
Once the basic request/response loop works, the natural extensions are:
- Streaming for chat UIs — see /docs/streaming
- Tool use for structured function calling — see /docs/tools
- Multi-turn conversations by appending assistant and user turns to the
messageslist - Retry logic with exponential backoff for
RateLimitErrorin production
questions
Do I need a paid Anthropic account to use the Python SDK? Yes, the SDK itself is free and open source, but you need a valid API key tied to a billed account to make actual requests. There's no SDK-only free tier.
What's the difference between client.messages.create and client.messages.stream? create waits for the full response and returns it at once. stream returns content incrementally as it's generated, which is better for interactive applications where you want to show output as it arrives rather than after a delay.
Can I use the same Python code with SubToAPI instead of Anthropic directly? The request and response structure is compatible — you point requests at https://api.subtoapi.app/v1/messages with a sub_live_... key instead of an Anthropic key. Check /docs/messages for the exact endpoint details before switching.