← Blog

Claude API Rate Limit Exceeded Error: How to Fix It

2026-09-28 · 5 min read · SubToAPI Team

If you're seeing a 429 status code with a body like {"type":"error","error":{"type":"rate_limit_error","message":"Number of request tokens has exceeded your per-minute rate limit"}}, your app has hit either the requests-per-minute (RPM), tokens-per-minute (TPM), or concurrent request cap tied to your Anthropic account tier. The fix depends on which limit you're hitting, but the short version is: catch the error, read the retry-after header, back off, and retry — don't just fire the same request again immediately.

The rest of this article walks through diagnosing which limit you're actually hitting, the retry logic that fixes most cases, and structural changes (queuing, batching, caching) that stop the errors from recurring under real traffic.

Why you're hitting the limit

Anthropic assigns rate limits per organization based on usage tier, and they cover three separate dimensions:

A 429 doesn't tell you which one you crossed unless you inspect the response headers. Check these on every response, not just failed ones:

curl -i https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-4-20250514","max_tokens":100,"messages":[{"role":"user","content":"hi"}]}'

Look for anthropic-ratelimit-requests-remaining, anthropic-ratelimit-tokens-remaining, and retry-after. If tokens-remaining hits zero while requests-remaining is healthy, you're token-bound — usually because prompts are too large or too many concurrent streams are consuming output tokens simultaneously. If requests-remaining hits zero with plenty of token budget left, you're making too many small calls too fast.

Fix 1: implement exponential backoff with jitter

This is the baseline fix and it solves the vast majority of transient 429s, especially bursty traffic that trips RPM limits for a few seconds.

async function callWithBackoff(fn, maxRetries = 5) {
  for (let attempt = 0; attempt <= maxRetries; attempt++) {
    try {
      return await fn();
    } catch (err) {
      if (err.status !== 429 || attempt === maxRetries) throw err;

      const retryAfter = Number(err.headers?.get("retry-after"));
      const backoff = retryAfter
        ? retryAfter * 1000
        : Math.min(1000 * 2 ** attempt, 30000);
      const jitter = Math.random() * 300;

      await new Promise((r) => setTimeout(r, backoff + jitter));
    }
  }
}

Key details that matter:

Fix 2: queue and throttle client-side

Backoff handles occasional bursts, but if your steady-state traffic is simply higher than your tier's limit, retries alone won't help — you need to cap outbound request rate before hitting the API.

A simple token-bucket or concurrency-limited queue works well:

import pLimit from "p-limit";

const limit = pLimit(5); // max 5 concurrent requests

async function sendMessage(payload) {
  return limit(() => callWithBackoff(() => client.messages.create(payload)));
}

Tune the concurrency number against your actual RPM/TPM budget, not a guess. If your tier allows 50 RPM and your average request takes 2 seconds, 5 concurrent workers is roughly the right ceiling — do the math for your own latency and limit combination.

Fix 3: reduce token volume per request

If you're token-bound rather than request-bound, backoff won't fix the root cause — you're just delaying the same overload. Cut token usage:

Fix 4: request a tier increase

Anthropic increases rate limits automatically as usage and billing history grow, and you can also request a limit increase directly through the console for legitimate scaling needs. If your application has predictable, growing traffic, this is the actual fix — backoff and queuing are mitigations, not a substitute for enough capacity.

Fix 5: separate the noisy caller from the rest of your app

A common failure mode: one endpoint or background job spikes token usage and exhausts the shared rate limit budget, causing unrelated requests elsewhere in your app to fail. Give high-volume or bursty workloads (batch jobs, bulk imports) their own request queue and concurrency cap, separate from user-facing request paths, so a backfill script doesn't take down your live chat feature.

If you're distributing access across a team or app

If multiple services, environments, or team members are all calling the Claude API with the same underlying account, it gets hard to tell which caller is causing rate limit pressure and hard to enforce per-service limits. This is one of the practical reasons teams put a layer like SubToAPI in front: it issues separate sub_live_... keys per app or environment, so you can rate-limit, monitor, and debug at the key level instead of guessing which part of your stack is responsible for a spike. Usage metadata per key also makes it obvious which caller to fix first. See /docs/quickstart for setup and /pricing for plan details.

FAQ

Why do I get a rate limit error even with low traffic?

Check the response headers — you may be token-bound rather than request-bound. A single request with a huge prompt or high max_tokens can consume the same token budget as dozens of small requests, triggering a 429 even at low request counts.

Does retrying immediately after a 429 make things worse?

Yes. Immediate retries, especially from multiple concurrent workers, tend to hit the same limit window and extend the throttling period. Always back off with jitter and honor the retry-after header when present.

Will upgrading my Anthropic usage tier fix this permanently?

It raises your ceiling, but if your traffic pattern is bursty, you still need backoff and queuing — a higher tier delays when you hit the limit, it doesn't eliminate burst-related 429s entirely.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →