Claude API Rate Limits Per Minute Explained
Claude API rate limits per minute are enforced across three dimensions: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Anthropic assigns your account to a usage tier based on billing history and account age, and each tier has fixed limits for these three metrics. Hit any one of them and the API returns a 429 Too Many Requests error, even if the other two limits still have headroom.
The practical answer most developers need: limits are not a single number, they're a combination of request count and token volume, and the token limits usually bind before the request count does. A tier that allows 50 requests per minute but only 40,000 input tokens per minute will throttle you on tokens if you're sending long prompts, long conversation history, or large documents, well before you hit 50 calls.
How Claude's rate limit tiers work
Anthropic organizes accounts into numbered usage tiers (Tier 1 through Tier 4, plus custom enterprise tiers). You move up tiers automatically as your spend and account history increase — there's no manual request process for the standard tiers. Each tier defines separate limits per model, meaning Claude Opus, Sonnet, and Haiku don't share the same pool.
The three metrics that matter:
- RPM (requests per minute) — how many individual API calls you can make
- ITPM (input tokens per minute) — total tokens across the prompt, system message, and any tool definitions you send
- OTPM (output tokens per minute) — total tokens Claude generates back to you
Streaming requests count against the same limits as non-streaming ones. A long-running streamed response still consumes OTPM budget as tokens are generated, not just at the start of the call.
Where the actual numbers come from
Anthropic does not publish a single static table that applies forever — limits change as they roll out new tiers and models, and enterprise agreements have custom numbers. The reliable way to check your current limits is the response headers on any API call:
anthropic-ratelimit-requests-limit: 50
anthropic-ratelimit-requests-remaining: 47
anthropic-ratelimit-tokens-limit: 40000
anthropic-ratelimit-tokens-remaining: 31500
anthropic-ratelimit-tokens-reset: 2024-01-01T00:00:15Z
Reading these headers on every response is the only way to know your real-time budget, because limits differ by model, by tier, and by whether you're on the standard API or a reseller/proxy layer.
Why you hit 429s even under "reasonable" usage
Three common patterns trigger rate limit errors that feel surprising:
- Long system prompts on every call. If your system prompt is 2,000 tokens and you send it on every request in a tight loop, you burn ITPM fast even with modest request counts.
- Conversation history growth. Multi-turn chat apps that resend the full message history each turn scale token usage linearly with conversation length, not with request count.
- Parallel workers without coordination. If you fan out requests from multiple background jobs or serverless functions without a shared rate limiter, you can blow past RPM in bursts even though your average usage looks fine.
Practical ways to stay under the limit
Implement exponential backoff. On a 429, read the retry-after header if present, or back off with jitter (start at 1s, double up to a cap, add randomness so parallel workers don't retry in lockstep).
async function callWithBackoff(fn, maxRetries = 5) {
for (let attempt = 0; attempt < maxRetries; attempt++) {
try {
return await fn();
} catch (err) {
if (err.status !== 429 || attempt === maxRetries - 1) throw err;
const delay = Math.min(1000 * 2 ** attempt, 30000) + Math.random() * 500;
await new Promise(r => setTimeout(r, delay));
}
}
}
Trim what you send. Summarize or truncate conversation history instead of resending everything. Cache system prompts where the API supports prompt caching so repeated tokens don't count at full price against your budget every call.
Queue instead of fan out. Route requests through a single queue with a token bucket that mirrors your actual tier limits, rather than letting every background job call the API independently.
Batch non-urgent work. If you're doing bulk classification, summarization, or embedding-adjacent tasks, spread the calls over time instead of firing them all at once.
Where SubToAPI fits
If you're building on top of your existing Claude access rather than a direct Anthropic enterprise contract, SubToAPI gives you a standard HTTPS API with your own sub_live_... application keys, so you don't have to hand-roll retry logic and key management for every project separately. It supports streaming, tool use, and returns usage metadata per request so you can watch token consumption in the dashboard instead of parsing headers manually across services. Plans start at Solo €9, with Team and Scale tiers for shared seats — see /pricing. The quickstart covers setup, and the messages and streaming docs cover request shapes if you're migrating existing integration code.
None of this changes the underlying Claude rate limits — those are set by your account tier — but centralizing your calls through one API layer makes it much easier to see when you're approaching them and to apply backoff consistently across every app that uses your Claude access.
questions
What counts toward the tokens-per-minute limit — input, output, or both? Both, but they're tracked as separate limits (ITPM and OTPM). A long prompt with a short answer can hit the input limit while leaving output headroom, and vice versa.
Do streaming responses count differently against rate limits? No. Streaming and non-streaming requests draw from the same RPM/ITPM/OTPM budget — streaming just delivers the output incrementally rather than all at once.
How do I increase my Claude API rate limits? Standard tiers increase automatically based on account age and cumulative spend. For limits beyond the standard tiers, you need to contact Anthropic sales for a custom enterprise agreement.