Building a Claude API Rate Limit Retry Queue
When your app sends more requests to the Claude API than your rate limit allows, you get back a 429 with a retry-after header. A retry queue is the mechanism that catches those rejected (or about-to-be-rejected) requests, holds them, and replays them at a pace the API will actually accept — instead of just failing the user's action or hammering the endpoint until it works.
This is different from simple retry logic. A single retry-with-backoff wrapper handles one request in isolation. A retry queue handles the traffic pattern: it controls how many requests are in flight at once, decides which request goes next when capacity frees up, and survives bursts without dropping work. If your app does more than a handful of concurrent Claude calls — background jobs, chat fan-out, batch summarization — you need the queue, not just the retry.
Why retries alone aren't enough
A naive approach looks like this: catch the 429, wait, retry. That works for occasional failures but breaks down under load for two reasons:
- Thundering herd: if 50 requests fail at once, they all retry at roughly the same time and hit the rate limit again.
- No prioritization: a background export job and a live user chat message end up competing for the same retry slot, with no way to say the chat message matters more.
A proper retry queue solves both by decoupling "when a request is created" from "when it is sent."
Core components of a Claude API retry queue
1. A concurrency-limited worker pool
Instead of firing requests as they arrive, push them into a queue and pull them out with a fixed number of workers (e.g., 3–5 concurrent calls). This alone prevents most rate-limit errors because you're never exceeding your known concurrency budget.
import PQueue from 'p-queue';
const queue = new PQueue({ concurrency: 4 });
async function enqueueClaudeRequest(payload) {
return queue.add(() => callClaude(payload));
}
2. Respect retry-after, don't guess
When a 429 does happen, Anthropic's API returns a retry-after header. Use it instead of a fixed backoff value — it tells you exactly when the window resets.
async function callClaude(payload) {
const res = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'x-api-key': process.env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01',
'content-type': 'application/json',
},
body: JSON.stringify(payload),
});
if (res.status === 429) {
const retryAfter = Number(res.headers.get('retry-after')) || 5;
await sleep(retryAfter * 1000);
return callClaude(payload); // re-queue, don't call directly in production
}
return res.json();
}
In production, don't recurse directly — push the failed job back into the queue with a delay so other workers can keep making progress while this one waits.
3. Jitter on top of backoff
If several requests fail around the same second, waking them all up at exactly the same moment just recreates the burst. Add a small random jitter (a few hundred milliseconds) to each retry delay so they spread out naturally.
function withJitter(ms) {
return ms + Math.random() * 300;
}
4. A priority field
Not all requests are equal. A queue item should carry a priority so interactive requests (a user waiting on a chat response) jump ahead of batch work (nightly summarization) when capacity is limited.
queue.add(() => callClaude(payload), { priority: isUserFacing ? 10 : 0 });
5. Persistence for anything that can't be lost
An in-memory queue is fine for low-stakes retries, but if your process restarts mid-backlog, those jobs vanish. For anything business-critical — billing summaries, async document processing — back the queue with Redis, SQS, or a database table so pending jobs survive a deploy or crash.
Where this gets expensive to maintain
Building the pieces above is a weekend project. Keeping it correct in production is the ongoing cost: tracking per-key rate limits as Anthropic adjusts them, handling the difference between request-rate limits and token-rate limits, adding dead-letter handling for jobs that fail repeatedly, and instrumenting enough logging to debug why a job sat in the queue for 40 seconds.
This is the kind of plumbing SubToAPI is built to remove. It sits in front of your Claude access and exposes it as a standard HTTPS API with its own application keys (sub_live_...), so your retry queue talks to a stable endpoint rather than managing Anthropic's rate-limit headers and token accounting directly. You still own the queue logic (concurrency, priority, persistence) — SubToAPI handles the request layer underneath it, including usage metadata per key so you can see which part of your system is actually generating the retries.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this report."}]
}'
If you're running multiple services or team members against the same Claude access, separate API keys per service also make it much easier to see which queue is actually causing rate-limit pressure — see /docs/quickstart for setup and /docs/messages for the request format. Plans start at €9/month for solo use, with team pricing at /pricing.
Practical checklist
- Cap concurrency below your known rate limit, don't discover it through errors.
- Always read and honor
retry-afterrather than using a fixed backoff. - Add jitter to avoid synchronized retry storms.
- Give queue items a priority so interactive requests don't wait behind batch jobs.
- Persist the queue if losing a pending job is unacceptable.
- Cap retry attempts and route permanently-failing jobs to a dead-letter queue instead of retrying forever.
Questions
Does a retry queue replace exponential backoff? No — it wraps it. Backoff decides how long a single request waits before its next attempt; the queue decides which request runs next and how many run concurrently across your whole app.
How many retries should I allow before giving up? Three to five is typical. After that, log the failure and move the job to a dead-letter queue for manual inspection rather than retrying indefinitely.
Can I use the same queue for streaming and non-streaming requests? Yes, but track them separately in your concurrency count — a streaming response holds a connection open much longer, so mixing them under one naive limit can starve non-streaming requests. See /docs/streaming for streaming-specific details.