Claude API Rate Limit Best Practices Guide
If you're building anything beyond a demo on top of Claude, you will hit rate limits eventually — either the per-minute request cap, the tokens-per-minute cap, or both. The question isn't whether you'll get a 429, it's how your application behaves when it does. This guide covers the practical patterns that keep production apps stable under Claude's rate limits: retry strategies, request shaping, concurrency control, and monitoring.
The short answer: implement exponential backoff with jitter on 429 responses, keep a client-side token budget so you rarely hit the limit in the first place, batch or queue bursty traffic instead of firing requests in parallel, and monitor usage trends before they become incidents. The rest of this article walks through each of these in detail.
Understand what's actually being limited
Claude's API enforces limits on multiple axes simultaneously:
- Requests per minute (RPM) — a cap on how many calls you can make
- Tokens per minute (TPM) — a cap on total input + output tokens processed
- Tokens per day, on some plans/tiers
Hitting any one of these triggers a 429. A common mistake is only tracking request count and being surprised by a 429 caused by token volume — a handful of requests with huge context windows can exhaust your TPM budget faster than a hundred small ones. Before optimizing anything, log both request count and token usage (input + output) per minute so you know which limit you're actually approaching.
Exponential backoff with jitter
The baseline pattern for any 429 or 5xx response is retry with exponential backoff, plus jitter to avoid synchronized retry storms across concurrent workers.
async function callWithBackoff(fn, maxRetries = 5) {
let attempt = 0;
while (true) {
try {
return await fn();
} catch (err) {
const status = err.status || err.response?.status;
if (![429, 500, 502, 503, 529].includes(status) || attempt >= maxRetries) {
throw err;
}
const base = Math.min(1000 * 2 ** attempt, 30000);
const jitter = Math.random() * base * 0.3;
await new Promise((r) => setTimeout(r, base + jitter));
attempt++;
}
}
}
Respect the retry-after header when it's present rather than always relying on your own backoff curve — it tells you exactly when the limit window resets. Cap the number of retries; silently retrying forever hides real problems and can cause request pileups.
Don't rely on retries alone — shape your traffic
Backoff handles occasional spikes gracefully, but if your baseline traffic is consistently near the limit, retries just delay the inevitable. A few structural changes reduce how often you hit 429s at all:
Queue instead of fan-out. If a single user action triggers ten Claude calls (e.g., summarizing ten documents), don't fire them all with Promise.all. Use a concurrency-limited queue (p-limit, bottleneck, or a simple semaphore) so you control how many requests are in flight at once.
import pLimit from "p-limit";
const limit = pLimit(4); // max 4 concurrent Claude calls
const results = await Promise.all(
documents.map((doc) => limit(() => callClaude(doc)))
);
Batch where the task allows it. Instead of one request per short item, combine several items into a single prompt when the task logic permits it. This reduces request count without necessarily reducing total tokens, so track both.
Trim context aggressively. Long conversation histories or oversized RAG contexts are the most common cause of hitting TPM limits. Summarize or truncate older turns, cap retrieved chunks to what's actually relevant, and avoid re-sending system prompts that don't need to change per request.
Separate interactive and batch workloads. User-facing chat traffic is latency-sensitive and unpredictable; background jobs (report generation, bulk classification) are not. Route batch jobs through a lower-priority queue with its own throttling so a nightly job doesn't starve real-time users of rate-limit headroom.
Client-side token budgeting
Rather than discovering your TPM ceiling via 429 errors, estimate token usage before sending and throttle proactively. A rough character-to-token ratio (roughly 4 characters per token for English text) is enough for a soft budget check — it doesn't need to be exact, just conservative enough to leave headroom.
function estimateTokens(text) {
return Math.ceil(text.length / 4);
}
if (runningTokenCount + estimateTokens(prompt) > tpmBudget) {
await sleep(msUntilNextWindow);
}
This is especially useful for multi-tenant apps where one customer's burst traffic shouldn't degrade service for everyone else — track token usage per tenant, not just globally.
Monitor before you get paged
Set up alerting on 429 rate, not just error rate in general — a spike in 429s specifically indicates you're pushing against a ceiling, which is a capacity problem, not a bug. Track:
- 429 count per minute, per endpoint
- Average and p95 token usage per request
- Queue depth if you're using a request queue
If you're seeing 429s during normal (not anomalous) traffic, that's a signal to either upgrade your rate limit tier, reduce token usage per request, or spread load more evenly across the day.
Simplify by moving rate limit logic out of your app
A lot of teams end up rebuilding the same retry/queue/monitoring stack around every LLM provider they use. If you'd rather not maintain that layer yourself, SubToAPI sits between your app and Claude, giving you application-scoped API keys (sub_live_...), streaming, and usage metadata through a single dashboard — useful if you want centralized visibility into token and request volume across multiple apps or team members without wiring up your own monitoring for each one. See the quickstart or the messages API reference for details, and pricing for plan options starting at Solo for individual use up to Scale for larger teams.
Practical checklist
- Implement exponential backoff with jitter and respect
retry-after - Track both request count and token volume, not just one
- Use a concurrency limiter for bursty or parallel workloads
- Trim conversation history and RAG context to reduce TPM pressure
- Separate interactive traffic from background batch jobs
- Alert on 429 rate specifically, before it becomes user-facing latency
FAQs
What causes a 429 from the Claude API? Either the requests-per-minute or tokens-per-minute limit for your account tier has been exceeded. Large prompts or long conversation histories can trigger token-based limits even with relatively few requests.
Does retrying automatically fix rate limit errors? Retries with backoff handle short bursts, but if your baseline traffic is consistently near the limit, retries just add latency rather than solving the underlying capacity issue. Reduce token usage or throttle proactively instead.
How do I avoid rate limits with multiple concurrent users? Use a concurrency-limited queue rather than firing all requests in parallel, track token usage per user or tenant so one burst doesn't starve others, and separate real-time traffic from background batch jobs so each can be throttled independently.