Claude API Rate Limit 429 Error: Causes and Fixes
A 429 Too Many Requests error from the Claude API means you've exceeded one of Anthropic's rate limits — requests per minute, tokens per minute, or tokens per day, depending on your usage tier. The fix depends on which limit you hit: for occasional spikes, implement exponential backoff and retry logic; for sustained high traffic, you need to either request a higher tier, batch your requests more efficiently, or distribute load across multiple keys/accounts.
This article walks through how to diagnose which limit you're hitting, how to implement proper retry handling, and the architectural changes that actually stop 429s from recurring — rather than just papering over them with a longer sleep() call.
Why You're Getting a 429
Anthropic rate limits are tiered based on your account's usage history and spend. Each tier has three separate caps:
- Requests per minute (RPM) — how many API calls you can make
- Input tokens per minute (ITPM) — total tokens sent across all requests
- Output tokens per minute (OTPM) — total tokens generated across all requests
Hitting any one of these triggers a 429, even if the other two are nowhere near their ceiling. A common mistake is assuming RPM is the bottleneck when it's actually ITPM — sending ten requests with 50,000-token prompts can exhaust your token budget long before you exhaust your request count.
The response headers on a successful (or rate-limited) call tell you exactly which limit is close to being hit:
anthropic-ratelimit-requests-remaining: 12
anthropic-ratelimit-tokens-remaining: 4500
anthropic-ratelimit-requests-reset: 2024-01-15T10:32:00Z
Check these headers before you add more backoff logic blindly — they tell you precisely how much headroom you have and when it resets.
Fix 1: Implement Exponential Backoff with Jitter
If 429s are occasional, the standard fix is retrying with exponential backoff instead of hammering the endpoint immediately:
async function callClaudeWithRetry(payload, maxRetries = 5) {
for (let attempt = 0; attempt < maxRetries; attempt++) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
if (res.status !== 429) return res.json();
const retryAfter = res.headers.get("retry-after");
const delayMs = retryAfter
? Number(retryAfter) * 1000
: Math.min(1000 * 2 ** attempt + Math.random() * 500, 30000);
await new Promise((r) => setTimeout(r, delayMs));
}
throw new Error("Max retries exceeded on 429");
}
Two details matter here: always respect the retry-after header if it's present — it's more accurate than a guessed exponential curve — and add jitter so that multiple concurrent workers don't retry in lockstep and cause a second wave of 429s.
Fix 2: Reduce Token Usage Per Request
If your bottleneck is ITPM or OTPM rather than RPM, backoff won't help much — you'll just keep hitting the same ceiling. Instead:
- Trim system prompts and few-shot examples to the minimum needed
- Cap
max_tokensto what the task actually requires instead of a generous default - Use prompt caching where supported to avoid re-sending large static context on every call
- Summarize or truncate long conversation histories instead of replaying the full transcript each turn
A 20% reduction in average prompt size often buys you a proportional increase in effective throughput without touching your rate limit tier at all.
Fix 3: Queue and Throttle Client-Side
Rather than firing requests as fast as your application generates them, add a queue that enforces your known RPM/TPM ceiling before requests ever leave your server:
class RateLimitedQueue {
constructor(maxPerMinute) {
this.maxPerMinute = maxPerMinute;
this.timestamps = [];
}
async acquire() {
const now = Date.now();
this.timestamps = this.timestamps.filter((t) => now - t < 60000);
if (this.timestamps.length >= this.maxPerMinute) {
const wait = 60000 - (now - this.timestamps[0]);
await new Promise((r) => setTimeout(r, wait));
}
this.timestamps.push(Date.now());
}
}
This converts bursty, unpredictable traffic into a smooth stream that stays under the ceiling — far more reliable than reacting to 429s after they happen.
Fix 4: Spread Load or Upgrade Your Tier
If you're consistently saturating your limits with legitimate traffic, the real fix is capacity, not code. You have two paths:
- Request a rate limit increase directly from Anthropic, which usually requires demonstrated usage history and sometimes a higher spend commitment.
- Route traffic through an API layer that handles key pooling, retries, and usage visibility for you, so you're not building and maintaining this infrastructure yourself.
This is where a service like SubToAPI fits in: it turns your Claude access into a clean HTTPS API with application-level keys (sub_live_...), per-key usage metadata, and team seats, so you can see exactly which key or team member is driving token consumption before it turns into a 429 storm. Check the quickstart to see how requests are structured, or look at streaming and tool use docs if your rate-limit pressure is coming from long-running or multi-step calls. Plans start with a free trial at signup, and tier details are on pricing.
Monitoring to Prevent Repeat 429s
Fixing one 429 incident doesn't prevent the next one. Add basic observability:
- Log the
anthropic-ratelimit-*-remainingheaders on every response - Alert when remaining capacity drops below 10-15% of the limit
- Track token usage per endpoint/feature so you know which part of your product is the heaviest consumer
- Review usage weekly, not just when something breaks
Rate limits are a capacity problem, not a bug — treating them like one you can monitor and plan around, rather than firefight, is the actual long-term fix.
questions
Does retrying immediately after a 429 make things worse? Yes. Immediate retries without backoff add more load to an already-saturated limit window and often trigger repeated 429s. Always wait at least until the retry-after value or reset timestamp before retrying.
Is a 429 the same as a 529 overloaded error? No. A 429 means you've exceeded your account's rate limit. A 529 means Anthropic's servers are temporarily overloaded regardless of your limit — backoff helps with both, but only a tier increase or usage reduction fixes recurring 429s.
Can I avoid 429s entirely by using a third-party API wrapper? A wrapper won't eliminate Anthropic's underlying limits, but tools like SubToAPI add visibility into per-key usage and streaming behavior, making it easier to spot and fix the specific request pattern causing the errors before it recurs.