Claude API Error Code 429: Troubleshooting Guide
A 429 response from the Claude API means you've hit a rate limit or quota ceiling, not that your request is malformed. The fix depends on which limit you tripped — requests per minute, tokens per minute, tokens per day, or concurrent connections — and that information is almost always in the response headers, not just the error body. This guide walks through how to identify the exact limit you hit and what to change so it stops happening.
What a 429 actually tells you
Anthropic's API returns 429 for several distinct situations that all look the same at the HTTP status level:
- Requests-per-minute (RPM) limit exceeded — you're sending too many individual API calls too fast.
- Tokens-per-minute (TPM) limit exceeded — your requests are large (long prompts, big context windows) even if the request count is low.
- Daily token quota exhausted — especially common on free tiers or newly created accounts.
- Concurrent request limit exceeded — too many in-flight requests at once, common with parallelized batch jobs.
The response body usually includes an error.type of rate_limit_error, but the body alone won't tell you which limit was hit. You need to inspect the headers.
Reading the rate limit headers
Every response from the Claude API includes headers describing your current limit state:
anthropic-ratelimit-requests-limit: 50
anthropic-ratelimit-requests-remaining: 0
anthropic-ratelimit-requests-reset: 2024-01-15T10:32:00Z
anthropic-ratelimit-tokens-limit: 40000
anthropic-ratelimit-tokens-remaining: 38210
anthropic-ratelimit-tokens-reset: 2024-01-15T10:32:00Z
If requests-remaining is 0 but tokens-remaining is healthy, you're hitting the RPM wall — the fix is to space out calls or batch work differently, not to shrink your prompts. If it's the reverse, your prompts or outputs are too large relative to your per-minute token budget, and spacing out requests won't help much.
Log these headers on every call. Without them you're guessing blind every time a 429 shows up in production.
Common causes and fixes
1. Bursty traffic patterns
Rate limits are per-minute windows, not smooth throttles. Sending 50 requests in the first 5 seconds of a minute and then nothing will still trip RPM limits even if your average load is low. Spread requests evenly instead of firing them all at once — a simple token-bucket or leaky-bucket limiter on your own side solves this.
2. Retrying without backoff
A naive retry loop that resends immediately on 429 makes the problem worse, because it adds more requests into a window that's already full. Always respect the anthropic-ratelimit-requests-reset timestamp (or at minimum use exponential backoff with jitter) before retrying. Retrying instantly is the single most common cause of cascading 429 storms in production systems.
3. Concurrent workers all hitting the API at once
If you run background jobs, webhooks, or queue workers that each call Claude independently, you can exceed concurrency or RPM limits even though no single worker is sending many requests. Add a shared rate limiter (Redis-based token bucket, for example) across all workers instead of limiting each process independently.
4. Underestimating token usage
Long system prompts, large conversation histories, or big tool definitions all count toward your token budget per minute — even on requests that return short answers. If you're hitting TPM limits, check whether you're re-sending the full conversation history on every turn instead of trimming it, and consider prompt caching for repeated large context blocks.
5. Tier or quota limits, not technical bugs
New accounts and some plan tiers have conservative default limits that increase automatically with usage history or require a support request to raise. If your headers show you're nowhere near your documented limit but still get 429s, check your account's usage tier — the limits are account-wide, not per-API-key.
A minimal retry pattern
async function callClaude(payload, attempt = 0) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
if (res.status === 429 && attempt < 5) {
const resetHeader = res.headers.get("anthropic-ratelimit-requests-reset");
const waitMs = resetHeader
? Math.max(0, new Date(resetHeader) - Date.now())
: 1000 * 2 ** attempt;
await new Promise((r) => setTimeout(r, waitMs + Math.random() * 300));
return callClaude(payload, attempt + 1);
}
return res.json();
}
This respects the actual reset window instead of guessing, and adds jitter so multiple workers don't retry in sync and re-trigger the limit together.
When the real problem is managing access, not retry logic
If your team is splitting a single Claude subscription across multiple apps, environments, or internal tools, 429s often come from everything sharing one undifferentiated quota with no visibility into who's consuming it. That's a structural problem retry logic can't fix.
This is where SubToAPI helps: it sits in front of your Claude access and issues separate application API keys (sub_live_...) per app or team, so you can see usage per key, set expectations per project, and isolate a noisy internal script from your production traffic instead of everything competing for the same limit invisibly. Usage metadata is included per request, so you can spot which key is driving token or request volume before it causes a 429 for everyone else.
Getting set up takes a few minutes — see the quickstart or the messages API reference for request formats, and streaming docs if your 429s are happening on long-running streamed responses. Plans start at Solo €9 with a free trial at signup, and pricing details are on the pricing page.
Checklist before you ship a fix
- Log
anthropic-ratelimit-*headers on every response, success or failure. - Confirm whether you're hitting RPM, TPM, or concurrency — don't assume.
- Add backoff that respects the reset timestamp, not a fixed delay.
- Centralize rate limiting across workers instead of per-process limits.
- Review whether prompt size, not request count, is the real driver.
questions
Does a 429 mean my API key is invalid? No. A 429 is purely a rate or quota limit response. An invalid or revoked key returns a 401, not a 429.
Will retrying immediately eventually succeed? Usually not — immediate retries add load to an already-full rate limit window and often make 429s more frequent. Wait until the reset timestamp in the response headers passes.
Can I increase my rate limits? Limits typically scale with account usage tier and can sometimes be raised through a support request. Check your current tier and headers first to confirm which specific limit you're hitting before requesting an increase.