Claude API Async Requests with Python asyncio
Claude API Async Requests with Python asyncio
If you're calling the Claude API from Python and your script processes requests one at a time, you're leaving a lot of speed on the table. The fix is to use Python's asyncio together with an async HTTP client so multiple requests run concurrently instead of waiting on each other. This article shows exactly how to structure that code, what libraries to use, and the mistakes that cause people to hit rate limits or silently lose responses.
The short answer: use httpx.AsyncClient (or Anthropic's official async client) inside an asyncio.gather() or asyncio.Semaphore-bounded loop, await each request, and collect results as they complete. Below is a full working pattern, plus the concurrency limits and error handling you need for production use.
Why synchronous Claude calls are slow
A typical synchronous Python loop calling an LLM API looks like this:
import requests
results = []
for prompt in prompts:
r = requests.post(url, json={"prompt": prompt}, headers=headers)
results.append(r.json())
Each request blocks the entire program until the response arrives. If a single Claude call takes 2–4 seconds and you have 100 prompts, that's 200–400 seconds of pure waiting, even though the CPU is idle the whole time. Since LLM API calls are I/O-bound — you're waiting on a network response, not doing computation — asyncio is the right tool. It lets your program start many requests, then resume each one as its response arrives, without spinning up threads or processes.
Setting up async calls with httpx
The official Anthropic SDK supports async, but the underlying pattern is the same regardless of provider, so this works whether you're calling Claude directly or through a proxy like SubToAPI. Here's a self-contained example using httpx:
import asyncio
import httpx
API_URL = "https://api.anthropic.com/v1/messages"
HEADERS = {
"x-api-key": "YOUR_KEY",
"anthropic-version": "2023-06-01",
"content-type": "application/json",
}
async def call_claude(client, prompt):
payload = {
"model": "claude-sonnet-4",
"max_tokens": 512,
"messages": [{"role": "user", "content": prompt}],
}
resp = await client.post(API_URL, json=payload, headers=HEADERS, timeout=60)
resp.raise_for_status()
return resp.json()
async def main(prompts):
async with httpx.AsyncClient() as client:
tasks = [call_claude(client, p) for p in prompts]
results = await asyncio.gather(*tasks, return_exceptions=True)
return results
prompts = ["Summarize photosynthesis", "Explain TCP handshakes", "Write a haiku about Rust"]
results = asyncio.run(main(prompts))
Key points:
- One
AsyncClientfor all requests. Creating a new client per call adds connection overhead. Reuse it viaasync with. asyncio.gatherruns tasks concurrently, not sequentially. All three prompts above go out roughly at the same time.return_exceptions=Truestops one failed request from crashing the entire batch — you get the exception object back instead, which you can inspect and retry.
Controlling concurrency with a semaphore
Firing off 500 requests simultaneously will almost certainly trigger rate limit errors. Bound concurrency with asyncio.Semaphore:
import asyncio
import httpx
sem = asyncio.Semaphore(10) # max 10 concurrent requests
async def call_claude_limited(client, prompt):
async with sem:
return await call_claude(client, prompt)
async def main(prompts):
async with httpx.AsyncClient() as client:
tasks = [call_claude_limited(client, p) for p in prompts]
return await asyncio.gather(*tasks, return_exceptions=True)
This caps how many requests are in flight at once. Start conservative (5–10) and increase based on your account's rate limits and observed error rates. This matters for both raw Anthropic API keys and application keys issued through a service like SubToAPI — every provider enforces concurrency and rate limits, so async code without a cap will eventually get throttled.
Adding retries for transient failures
Async doesn't remove the need for retry logic — it just means you need to retry inside each coroutine rather than in a blocking loop:
async def call_with_retry(client, prompt, max_retries=3):
for attempt in range(max_retries):
try:
return await call_claude(client, prompt)
except httpx.HTTPStatusError as e:
if e.response.status_code == 429 and attempt < max_retries - 1:
await asyncio.sleep(2 ** attempt)
continue
raise
Use asyncio.sleep(), never time.sleep(), inside async functions — time.sleep() blocks the entire event loop, defeating the purpose of going async in the first place.
Streaming responses asynchronously
If you also need token-by-token streaming inside an async context, httpx.AsyncClient supports it with client.stream():
async def stream_claude(client, prompt):
async with client.stream("POST", API_URL, json={...}, headers=HEADERS) as resp:
async for line in resp.aiter_lines():
if line.startswith("data:"):
print(line)
This lets you run several streaming conversations concurrently, each yielding tokens independently — useful for chat backends serving multiple users at once.
Where SubToAPI fits
If you're building this against a hosted Claude proxy rather than raw Anthropic credentials, SubToAPI issues application keys (sub_live_...) that work with the same async patterns above — just swap the URL for https://api.subtoapi.app/v1/messages and the header for Authorization: Bearer $SUBTOAPI_KEY. It supports streaming and tool use over the same async client code, and gives each team member their own key with usage tracked centrally. See the quickstart and streaming docs for exact payload shapes, or check pricing if you need per-seat billing for a team.
FAQ
Does asyncio actually make Claude API calls faster?
Yes, for I/O-bound workloads like HTTP calls. asyncio doesn't speed up any single request, but it lets many requests wait on network responses at the same time instead of one after another, cutting total wall-clock time for batches significantly.
Should I use requests with threads instead of asyncio?
You can — ThreadPoolExecutor with requests also achieves concurrency and is simpler to reason about for small scripts. asyncio with httpx scales better to hundreds of concurrent calls with lower memory overhead, which matters for larger batch jobs.
How many concurrent Claude requests can I safely send?
It depends on your account's rate limits, which vary by plan and model. Start with a semaphore limit around 5–10, monitor for 429 responses, and increase gradually. Always implement exponential backoff regardless of your concurrency cap.