Claude API Rate Limit Handling Best Practices
Rate limits on the Claude API exist to protect shared infrastructure, and every production app that calls Claude at scale will hit them eventually. The fix isn't to avoid rate limits entirely — that's impossible under real traffic — it's to handle them gracefully so a 429 response never becomes a user-facing error.
This guide covers the concrete patterns: reading rate limit headers, implementing exponential backoff with jitter, queueing requests client-side, and structuring your architecture so a burst of traffic degrades gracefully instead of failing outright.
Understand what's actually being limited
Claude API rate limits typically apply across a few dimensions at once:
- Requests per minute (RPM) — how many API calls you can make
- Tokens per minute (TPM) — input + output tokens processed
- Tokens per day (TPD) — a longer-window cap on some plans
You can hit any of these independently. A handful of requests with huge prompts can exhaust your token budget long before you hit your request count. Before writing retry logic, check which limit you're actually bumping into — the response headers tell you.
curl -i https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-3-5-sonnet-20241022","max_tokens":100,"messages":[{"role":"user","content":"hi"}]}'
Look for anthropic-ratelimit-requests-remaining, anthropic-ratelimit-tokens-remaining, and retry-after in the response headers. These give you the information to back off proactively instead of waiting for a 429.
Implement exponential backoff with jitter
A 429 response means "try again later," not "the request failed permanently." The standard pattern is exponential backoff with jitter: wait progressively longer between retries, and add randomness so concurrent clients don't all retry at the exact same moment.
async function callWithBackoff(fn, maxRetries = 5) {
let attempt = 0;
while (true) {
try {
return await fn();
} catch (err) {
if (err.status !== 429 || attempt >= maxRetries) throw err;
const retryAfter = err.headers?.['retry-after'];
const base = retryAfter ? Number(retryAfter) * 1000 : 2 ** attempt * 500;
const jitter = Math.random() * 300;
await new Promise(r => setTimeout(r, base + jitter));
attempt++;
}
}
}
A few things matter here:
- Respect
retry-afterwhen present. It's more accurate than guessing with a fixed exponential curve. - Cap your retries. Five attempts is usually enough — beyond that, surface a clear error to the caller rather than hanging.
- Don't retry non-429 errors the same way. A 400 (bad request) will fail identically every time; retrying wastes time and quota.
Queue requests instead of firing them all at once
If your app sends bursts of requests — batch processing, bulk document analysis, background jobs — a queue with a controlled concurrency limit prevents you from ever hitting the ceiling in the first place.
class RateLimitedQueue {
constructor(concurrency = 3) {
this.concurrency = concurrency;
this.running = 0;
this.queue = [];
}
async add(task) {
return new Promise((resolve, reject) => {
this.queue.push({ task, resolve, reject });
this._next();
});
}
async _next() {
if (this.running >= this.concurrency || this.queue.length === 0) return;
this.running++;
const { task, resolve, reject } = this.queue.shift();
try {
resolve(await task());
} catch (e) {
reject(e);
} finally {
this.running--;
this._next();
}
}
}
Tune concurrency against your actual RPM/TPM limits. If your plan allows 50 RPM and each request takes ~2 seconds, a concurrency of 3–5 keeps you comfortably under the ceiling without idling.
Reduce token usage to stay under TPM limits
Since token throughput is often the binding constraint, reducing tokens per request buys you more headroom than retry logic ever will:
- Trim system prompts and remove redundant instructions
- Use prompt caching for repeated context (tool definitions, long system prompts) where supported
- Cap
max_tokensto what the task actually needs — don't default to the maximum - Summarize or truncate long conversation history instead of sending the full thread every time
Build in graceful degradation
Even with backoff and queueing, you should design for the case where Claude is temporarily unavailable or rate-limited beyond what retries can absorb:
- Return cached or partial results if available rather than a hard failure
- Show users a "processing" state instead of blocking on a synchronous call
- Log rate limit events separately from other errors so you can see patterns (time of day, specific endpoints) and adjust capacity planning accordingly
Where SubToAPI fits in
If you're managing Claude access for a team or across multiple apps, rate limit handling gets more complicated — you're coordinating concurrency and quotas across every consumer hitting the same underlying account. SubToAPI turns your Claude access into a standard HTTPS API with its own application keys (sub_live_...), so each app or environment gets its own key and usage is visible per key in one dashboard. That makes it much easier to see which part of your system is actually driving rate limit pressure, rather than debugging blind against a single shared credential.
Getting started takes a few minutes — see the quickstart guide or the Messages API reference for request and response formats, and the streaming guide if you're building long-running completions where backoff timing matters even more. You can try it with a free trial at signup and compare plans here.
Summary checklist
- Read rate limit headers on every response, not just on 429s
- Implement exponential backoff with jitter and a retry cap
- Queue bulk or batch requests with a concurrency limit tuned to your plan
- Reduce token usage with trimmed prompts, caching, and sensible
max_tokens - Log rate limit events separately to spot patterns before they become outages
FAQ
What's the difference between RPM and TPM limits on the Claude API? RPM (requests per minute) caps how many API calls you can make; TPM (tokens per minute) caps total input and output tokens processed. Long prompts or large outputs can exhaust your token budget well before you hit your request count, so track both.
Should I retry every Claude API error the same way? No. Retry 429s and 5xx errors with exponential backoff, since they're typically transient. Don't retry 400-class errors like malformed requests — they'll fail identically every time and just waste quota.
How do I avoid rate limits when processing documents in bulk? Use a concurrency-limited queue instead of firing all requests at once, reduce token usage per request where possible, and spread large batch jobs over time rather than running them in a single burst.