← Blog

Anthropic API Rate Limits Explained

2026-10-05 · 5 min read · SubToAPI Team

Anthropic's API rate limits are usage caps tied to your account tier that restrict how many requests and tokens you can send per minute (and sometimes per day). They exist to protect Anthropic's infrastructure from sudden spikes and to allocate capacity fairly across customers as usage scales. If you've hit a 429 Too Many Requests error or seen your requests slow down under load, you've run into one of these limits.

Anthropic measures limits across three dimensions simultaneously: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). You can be rate limited by whichever one you hit first, even if the other two are nowhere near their cap. This is the detail that trips up most developers — a low-traffic app with very long prompts can get rate limited on tokens long before it hits a request count ceiling.

How the tier system works

Anthropic assigns accounts to usage tiers that unlock progressively higher limits. Tiers typically advance based on cumulative spend and account age, not a manual request. A brand new account starts at the lowest tier with conservative limits — enough for development and testing, but not for production traffic at scale. As your billing history grows, Anthropic automatically raises your tier and your RPM/TPM ceilings along with it.

This matters for planning: if you're building a product that expects real usage on day one, budget time for your account to "warm up" through the tiers, or plan around the limits you'll actually have at launch rather than the ones you hope to reach.

The three limits that matter

Requests per minute (RPM) — a hard cap on how many API calls you can make in a 60-second window, regardless of size.

Input tokens per minute (ITPM) — the total tokens across all request bodies (including system prompts, message history, and any tool definitions) you send within a minute.

Output tokens per minute (OTPM) — the total tokens Claude generates in responses within a minute. This one is trickier to predict because you don't fully control output length unless you cap it with max_tokens.

All three reset on a rolling basis, not a fixed clock minute, so bursts right at a reset boundary don't give you a free pass.

Reading rate limit headers

Every API response includes headers that tell you exactly where you stand:

anthropic-ratelimit-requests-limit: 50
anthropic-ratelimit-requests-remaining: 42
anthropic-ratelimit-requests-reset: 2024-06-01T12:01:00Z
anthropic-ratelimit-tokens-limit: 40000
anthropic-ratelimit-tokens-remaining: 31500
anthropic-ratelimit-tokens-reset: 2024-06-01T12:01:00Z

Checking these on every response — not just when you get a 429 — lets you throttle proactively instead of reactively. A simple pattern: if remaining drops below some threshold (say 10% of limit), start adding delay between requests before you actually get rejected.

What happens when you exceed a limit

You get a 429 status code with a retry-after header indicating how many seconds to wait. The correct response is exponential backoff with jitter, not a tight retry loop — hammering the API immediately after a 429 just extends the time you're locked out.

async function callWithBackoff(fn, maxRetries = 5) {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await fn();
    } catch (err) {
      if (err.status !== 429 || attempt === maxRetries - 1) throw err;
      const retryAfter = Number(err.headers?.['retry-after']) || 2 ** attempt;
      const jitter = Math.random() * 0.3 * retryAfter;
      await new Promise((r) => setTimeout(r, (retryAfter + jitter) * 1000));
    }
  }
}

This is a small amount of code, but it's the difference between a rate limit being an occasional slow response and it being a production outage.

Practical ways to work within limits

Where a gateway layer helps

A common pattern for teams with multiple services or multiple developers hitting Claude is that nobody has a clear view of combined usage until a rate limit error shows up in production — often because two services, two developers, or two environments are drawing from the same account limits without anyone coordinating.

This is one of the reasons we built SubToAPI: it sits between your applications and Claude, issuing scoped sub_live_... keys per app or per team member, with usage metadata visible in one dashboard instead of scattered across logs. You still work against the same request and token model described above — see the messages docs and streaming docs for the request shapes — but you get visibility into who's consuming what before it turns into a shared rate-limit problem. Plans start at Solo for individual use, with Team and Scale tiers adding per-seat API keys; check /pricing for details, or run through the quickstart to see it against your own Claude access.

Designing for limits, not around them

Rate limits aren't a bug to route around — they're a capacity signal. The teams that handle them well treat token and request budgets as a first-class part of their system design: they set max_tokens intentionally, they backoff properly, and they know which of RPM, ITPM, or OTPM is their actual bottleneck instead of guessing. That last part is easy to find out — just read the headers.

questions

Why am I rate limited when my request volume looks low? You're likely hitting the token limit, not the request limit. Long system prompts, large message histories, or generous max_tokens settings consume ITPM/OTPM budget fast even with few requests.

Do rate limits reset instantly at a fixed time? No — they're rolling windows, not fixed-clock minutes. A burst right before what you think is a reset boundary can still get throttled.

Can I request a higher rate limit tier manually? Tiers generally advance automatically based on account spend and history. Sustained, legitimate usage over time is what moves you up, rather than a one-time request.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →