Claude API Error 429 Too Many Requests: Causes & Fixes
What a 429 error means on the Claude API
A 429 Too Many Requests error from the Claude API means you've exceeded one of Anthropic's rate limits: requests per minute (RPM), input tokens per minute (ITPM), or output tokens per minute (OTPM). The limit is tied to your organization's usage tier, which is determined by your billing history and spend, not by how urgent your request is. When any one of these three limits is crossed, even for a single second, the API rejects the request with a 429 and a retry-after header telling you how long to wait.
The fix depends on which limit you're hitting. If it's RPM, you're sending requests too frequently (common with tight retry loops or parallel workers). If it's ITPM/OTPM, your prompts or outputs are too large relative to your tier, even if request count is low. Below is how to diagnose which one is the problem and the concrete changes that stop 429s from recurring, plus how a gateway layer like SubToAPI can absorb bursty traffic so your app doesn't crash on spikes.
Diagnosing which limit you hit
Check the response headers on the failed request — Anthropic returns rate limit metadata on every call, not just on errors:
curl -i https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-4-5","max_tokens":1024,"messages":[{"role":"user","content":"hi"}]}'
Look for:
anthropic-ratelimit-requests-remaininganthropic-ratelimit-tokens-remainingretry-after
If requests-remaining hits zero first, you're RPM-bound. If tokens-remaining hits zero while you still have request headroom, you're token-bound — usually because of large system prompts, long conversation history, or big max_tokens values on every call.
Fixing RPM (requests per minute) limits
RPM issues almost always come from one of these patterns:
- Tight retry loops — a failed request retried instantly in a
whileloop, multiplying calls per second. - Parallel fan-out — spinning up many concurrent workers (serverless functions, queue consumers) that all call the API at once with no coordination.
- Polling instead of streaming — repeatedly checking for a response instead of using a single streamed connection.
The standard fix is exponential backoff with jitter:
async function callClaude(payload, attempt = 0) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
if (res.status === 429 && attempt < 5) {
const retryAfter = Number(res.headers.get("retry-after")) || 2 ** attempt;
const jitter = Math.random() * 0.5;
await new Promise((r) => setTimeout(r, (retryAfter + jitter) * 1000));
return callClaude(payload, attempt + 1);
}
if (!res.ok) throw new Error(`Claude API error: ${res.status}`);
return res.json();
}
This respects the retry-after header when present, falls back to exponential delay otherwise, and caps retries so a persistent problem doesn't retry forever.
Fixing token-based (ITPM/OTPM) limits
If backoff doesn't help and you're still getting 429s, the problem is token volume, not request count. Common causes and fixes:
- Trim conversation history. Don't send the full chat log on every turn if older messages aren't needed — summarize or truncate.
- Cap
max_tokensrealistically. Settingmax_tokens: 4096on every call when you only need short answers reserves capacity you don't use but may still count toward throttling behavior in some configurations. - Use prompt caching for repeated system prompts. If you send the same large system prompt or document context repeatedly, caching avoids re-processing those tokens on every call — see /docs/messages for implementation details.
- Batch smarter. Instead of 100 tiny sequential calls, consider whether some work can be combined into fewer, larger calls (within reason) or spread across a longer time window.
Reducing concurrent request pressure
If your application has multiple services or team members all calling the Claude API under the same account, you can hit limits without any single service being "at fault" — it's the combined traffic. This is a common issue for teams that give each developer or each microservice its own raw API key pointed at the same org quota.
A practical way to manage this is to route all traffic through a single gateway that handles queuing, retries, and key issuance per application rather than per developer. SubToAPI sits in front of your Claude access and gives each app or team member its own sub_live_... API key, so you can see exactly which key is generating load, apply per-key limits, and avoid one noisy service starving the others. Streaming responses and tool use both work the same as calling Claude directly — see /docs/streaming and /docs/tools — and usage metadata per key makes it easy to spot which integration is causing 429s before it becomes a production incident. Setup takes a few minutes; start at /signup or check /pricing for the Solo, Team, and Scale plans.
Preventing 429s before they happen
- Monitor headroom, not just errors. Log the rate limit headers on every successful call so you see when you're approaching a limit, not just after you've hit it.
- Add a queue for bursty workloads. If traffic spikes (e.g., a batch job or a marketing push), queue requests and drain them at a controlled rate instead of firing them all at once.
- Separate keys by workload. Give background jobs and interactive user requests different keys/limits where possible so a batch job doesn't starve real-time traffic.
- Review your tier. If you consistently hit limits despite good client-side behavior, your usage has outgrown your current tier — check your Anthropic console for upgrade options.
Getting a quick start on the basics of request structure also helps avoid unnecessary retries from malformed requests — /docs/quickstart covers the essentials if you're new to the Messages API.
questions
Does a 429 error cost me tokens or money? No. A 429 response means the request was rejected before processing, so you are not billed for it. Only successfully processed requests consume tokens and incur cost.
Is 429 the same as a 529 "overloaded" error? No. 429 means you exceeded your own account's rate limit. 529 means Anthropic's infrastructure is temporarily overloaded regardless of your usage — the fix for 529 is simply to retry with backoff, since it's not tied to your limits.
Will upgrading my Anthropic usage tier fix 429 errors permanently? It raises your limits, but if request patterns are inefficient (tight retry loops, unnecessary token bloat), you'll eventually hit the new, higher ceiling too. Fix client-side behavior first, then consider a tier upgrade if volume genuinely requires it.