Claude API Rate Limit Exceeded Error: How to Fix It
If you're seeing a 429 status code with a body like {"type":"error","error":{"type":"rate_limit_error","message":"Number of request tokens has exceeded your per-minute rate limit"}}, your app has hit either the requests-per-minute (RPM), tokens-per-minute (TPM), or concurrent request cap tied to your Anthropic account tier. The fix depends on which limit you're hitting, but the short version is: catch the error, read the retry-after header, back off, and retry — don't just fire the same request again immediately.
The rest of this article walks through diagnosing which limit you're actually hitting, the retry logic that fixes most cases, and structural changes (queuing, batching, caching) that stop the errors from recurring under real traffic.
Why you're hitting the limit
Anthropic assigns rate limits per organization based on usage tier, and they cover three separate dimensions:
- Requests per minute (RPM) — total number of API calls, regardless of size
- Input tokens per minute (ITPM) and output tokens per minute (OTPM) — total token throughput
- Concurrent requests — how many requests can be in flight at once
A 429 doesn't tell you which one you crossed unless you inspect the response headers. Check these on every response, not just failed ones:
curl -i https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-4-20250514","max_tokens":100,"messages":[{"role":"user","content":"hi"}]}'
Look for anthropic-ratelimit-requests-remaining, anthropic-ratelimit-tokens-remaining, and retry-after. If tokens-remaining hits zero while requests-remaining is healthy, you're token-bound — usually because prompts are too large or too many concurrent streams are consuming output tokens simultaneously. If requests-remaining hits zero with plenty of token budget left, you're making too many small calls too fast.
Fix 1: implement exponential backoff with jitter
This is the baseline fix and it solves the vast majority of transient 429s, especially bursty traffic that trips RPM limits for a few seconds.
async function callWithBackoff(fn, maxRetries = 5) {
for (let attempt = 0; attempt <= maxRetries; attempt++) {
try {
return await fn();
} catch (err) {
if (err.status !== 429 || attempt === maxRetries) throw err;
const retryAfter = Number(err.headers?.get("retry-after"));
const backoff = retryAfter
? retryAfter * 1000
: Math.min(1000 * 2 ** attempt, 30000);
const jitter = Math.random() * 300;
await new Promise((r) => setTimeout(r, backoff + jitter));
}
}
}
Key details that matter:
- Always respect
retry-afterif present. It's more accurate than guessing. - Cap your backoff (30–60 seconds is reasonable) so a stuck retry loop doesn't hang a request indefinitely.
- Add jitter. Without it, many concurrent workers retry at the exact same moment and immediately re-trigger the limit.
- Set a max retry count and surface a real error to the caller after that — silent infinite retries hide the problem from your monitoring.
Fix 2: queue and throttle client-side
Backoff handles occasional bursts, but if your steady-state traffic is simply higher than your tier's limit, retries alone won't help — you need to cap outbound request rate before hitting the API.
A simple token-bucket or concurrency-limited queue works well:
import pLimit from "p-limit";
const limit = pLimit(5); // max 5 concurrent requests
async function sendMessage(payload) {
return limit(() => callWithBackoff(() => client.messages.create(payload)));
}
Tune the concurrency number against your actual RPM/TPM budget, not a guess. If your tier allows 50 RPM and your average request takes 2 seconds, 5 concurrent workers is roughly the right ceiling — do the math for your own latency and limit combination.
Fix 3: reduce token volume per request
If you're token-bound rather than request-bound, backoff won't fix the root cause — you're just delaying the same overload. Cut token usage:
- Trim system prompts and remove redundant context on every call
- Use prompt caching for repeated large context (docs, tool schemas, few-shot examples) if your workload supports it
- Cap
max_tokensto what the task actually needs instead of leaving it high by default - Summarize or truncate long conversation histories instead of resending the full transcript every turn
Fix 4: request a tier increase
Anthropic increases rate limits automatically as usage and billing history grow, and you can also request a limit increase directly through the console for legitimate scaling needs. If your application has predictable, growing traffic, this is the actual fix — backoff and queuing are mitigations, not a substitute for enough capacity.
Fix 5: separate the noisy caller from the rest of your app
A common failure mode: one endpoint or background job spikes token usage and exhausts the shared rate limit budget, causing unrelated requests elsewhere in your app to fail. Give high-volume or bursty workloads (batch jobs, bulk imports) their own request queue and concurrency cap, separate from user-facing request paths, so a backfill script doesn't take down your live chat feature.
If you're distributing access across a team or app
If multiple services, environments, or team members are all calling the Claude API with the same underlying account, it gets hard to tell which caller is causing rate limit pressure and hard to enforce per-service limits. This is one of the practical reasons teams put a layer like SubToAPI in front: it issues separate sub_live_... keys per app or environment, so you can rate-limit, monitor, and debug at the key level instead of guessing which part of your stack is responsible for a spike. Usage metadata per key also makes it obvious which caller to fix first. See /docs/quickstart for setup and /pricing for plan details.
FAQ
Why do I get a rate limit error even with low traffic?
Check the response headers — you may be token-bound rather than request-bound. A single request with a huge prompt or high max_tokens can consume the same token budget as dozens of small requests, triggering a 429 even at low request counts.
Does retrying immediately after a 429 make things worse?
Yes. Immediate retries, especially from multiple concurrent workers, tend to hit the same limit window and extend the throttling period. Always back off with jitter and honor the retry-after header when present.
Will upgrading my Anthropic usage tier fix this permanently?
It raises your ceiling, but if your traffic pattern is bursty, you still need backoff and queuing — a higher tier delays when you hit the limit, it doesn't eliminate burst-related 429s entirely.