← Blog

Claude API Async Requests with Python asyncio

2026-09-29 · 4 min read · SubToAPI Team

Claude API Async Requests with Python asyncio

If you're calling the Claude API from Python and your script processes requests one at a time, you're leaving a lot of speed on the table. The fix is to use Python's asyncio together with an async HTTP client so multiple requests run concurrently instead of waiting on each other. This article shows exactly how to structure that code, what libraries to use, and the mistakes that cause people to hit rate limits or silently lose responses.

The short answer: use httpx.AsyncClient (or Anthropic's official async client) inside an asyncio.gather() or asyncio.Semaphore-bounded loop, await each request, and collect results as they complete. Below is a full working pattern, plus the concurrency limits and error handling you need for production use.

Why synchronous Claude calls are slow

A typical synchronous Python loop calling an LLM API looks like this:

import requests

results = []
for prompt in prompts:
    r = requests.post(url, json={"prompt": prompt}, headers=headers)
    results.append(r.json())

Each request blocks the entire program until the response arrives. If a single Claude call takes 2–4 seconds and you have 100 prompts, that's 200–400 seconds of pure waiting, even though the CPU is idle the whole time. Since LLM API calls are I/O-bound — you're waiting on a network response, not doing computation — asyncio is the right tool. It lets your program start many requests, then resume each one as its response arrives, without spinning up threads or processes.

Setting up async calls with httpx

The official Anthropic SDK supports async, but the underlying pattern is the same regardless of provider, so this works whether you're calling Claude directly or through a proxy like SubToAPI. Here's a self-contained example using httpx:

import asyncio
import httpx

API_URL = "https://api.anthropic.com/v1/messages"
HEADERS = {
    "x-api-key": "YOUR_KEY",
    "anthropic-version": "2023-06-01",
    "content-type": "application/json",
}

async def call_claude(client, prompt):
    payload = {
        "model": "claude-sonnet-4",
        "max_tokens": 512,
        "messages": [{"role": "user", "content": prompt}],
    }
    resp = await client.post(API_URL, json=payload, headers=HEADERS, timeout=60)
    resp.raise_for_status()
    return resp.json()

async def main(prompts):
    async with httpx.AsyncClient() as client:
        tasks = [call_claude(client, p) for p in prompts]
        results = await asyncio.gather(*tasks, return_exceptions=True)
    return results

prompts = ["Summarize photosynthesis", "Explain TCP handshakes", "Write a haiku about Rust"]
results = asyncio.run(main(prompts))

Key points:

Controlling concurrency with a semaphore

Firing off 500 requests simultaneously will almost certainly trigger rate limit errors. Bound concurrency with asyncio.Semaphore:

import asyncio
import httpx

sem = asyncio.Semaphore(10)  # max 10 concurrent requests

async def call_claude_limited(client, prompt):
    async with sem:
        return await call_claude(client, prompt)

async def main(prompts):
    async with httpx.AsyncClient() as client:
        tasks = [call_claude_limited(client, p) for p in prompts]
        return await asyncio.gather(*tasks, return_exceptions=True)

This caps how many requests are in flight at once. Start conservative (5–10) and increase based on your account's rate limits and observed error rates. This matters for both raw Anthropic API keys and application keys issued through a service like SubToAPI — every provider enforces concurrency and rate limits, so async code without a cap will eventually get throttled.

Adding retries for transient failures

Async doesn't remove the need for retry logic — it just means you need to retry inside each coroutine rather than in a blocking loop:

async def call_with_retry(client, prompt, max_retries=3):
    for attempt in range(max_retries):
        try:
            return await call_claude(client, prompt)
        except httpx.HTTPStatusError as e:
            if e.response.status_code == 429 and attempt < max_retries - 1:
                await asyncio.sleep(2 ** attempt)
                continue
            raise

Use asyncio.sleep(), never time.sleep(), inside async functions — time.sleep() blocks the entire event loop, defeating the purpose of going async in the first place.

Streaming responses asynchronously

If you also need token-by-token streaming inside an async context, httpx.AsyncClient supports it with client.stream():

async def stream_claude(client, prompt):
    async with client.stream("POST", API_URL, json={...}, headers=HEADERS) as resp:
        async for line in resp.aiter_lines():
            if line.startswith("data:"):
                print(line)

This lets you run several streaming conversations concurrently, each yielding tokens independently — useful for chat backends serving multiple users at once.

Where SubToAPI fits

If you're building this against a hosted Claude proxy rather than raw Anthropic credentials, SubToAPI issues application keys (sub_live_...) that work with the same async patterns above — just swap the URL for https://api.subtoapi.app/v1/messages and the header for Authorization: Bearer $SUBTOAPI_KEY. It supports streaming and tool use over the same async client code, and gives each team member their own key with usage tracked centrally. See the quickstart and streaming docs for exact payload shapes, or check pricing if you need per-seat billing for a team.

FAQ

Does asyncio actually make Claude API calls faster?

Yes, for I/O-bound workloads like HTTP calls. asyncio doesn't speed up any single request, but it lets many requests wait on network responses at the same time instead of one after another, cutting total wall-clock time for batches significantly.

Should I use requests with threads instead of asyncio?

You can — ThreadPoolExecutor with requests also achieves concurrency and is simpler to reason about for small scripts. asyncio with httpx scales better to hundreds of concurrent calls with lower memory overhead, which matters for larger batch jobs.

How many concurrent Claude requests can I safely send?

It depends on your account's rate limits, which vary by plan and model. Start with a semaphore limit around 5–10, monitor for 429 responses, and increase gradually. Always implement exponential backoff regardless of your concurrency cap.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →