← Blog

Claude API Rate Limit 429 Error: Causes and Fixes

2026-10-10 · 5 min read · SubToAPI Team

A 429 Too Many Requests error from the Claude API means you've exceeded one of Anthropic's rate limits — requests per minute, tokens per minute, or tokens per day, depending on your usage tier. The fix depends on which limit you hit: for occasional spikes, implement exponential backoff and retry logic; for sustained high traffic, you need to either request a higher tier, batch your requests more efficiently, or distribute load across multiple keys/accounts.

This article walks through how to diagnose which limit you're hitting, how to implement proper retry handling, and the architectural changes that actually stop 429s from recurring — rather than just papering over them with a longer sleep() call.

Why You're Getting a 429

Anthropic rate limits are tiered based on your account's usage history and spend. Each tier has three separate caps:

Hitting any one of these triggers a 429, even if the other two are nowhere near their ceiling. A common mistake is assuming RPM is the bottleneck when it's actually ITPM — sending ten requests with 50,000-token prompts can exhaust your token budget long before you exhaust your request count.

The response headers on a successful (or rate-limited) call tell you exactly which limit is close to being hit:

anthropic-ratelimit-requests-remaining: 12
anthropic-ratelimit-tokens-remaining: 4500
anthropic-ratelimit-requests-reset: 2024-01-15T10:32:00Z

Check these headers before you add more backoff logic blindly — they tell you precisely how much headroom you have and when it resets.

Fix 1: Implement Exponential Backoff with Jitter

If 429s are occasional, the standard fix is retrying with exponential backoff instead of hammering the endpoint immediately:

async function callClaudeWithRetry(payload, maxRetries = 5) {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    const res = await fetch("https://api.anthropic.com/v1/messages", {
      method: "POST",
      headers: {
        "x-api-key": process.env.ANTHROPIC_API_KEY,
        "anthropic-version": "2023-06-01",
        "content-type": "application/json",
      },
      body: JSON.stringify(payload),
    });

    if (res.status !== 429) return res.json();

    const retryAfter = res.headers.get("retry-after");
    const delayMs = retryAfter
      ? Number(retryAfter) * 1000
      : Math.min(1000 * 2 ** attempt + Math.random() * 500, 30000);

    await new Promise((r) => setTimeout(r, delayMs));
  }
  throw new Error("Max retries exceeded on 429");
}

Two details matter here: always respect the retry-after header if it's present — it's more accurate than a guessed exponential curve — and add jitter so that multiple concurrent workers don't retry in lockstep and cause a second wave of 429s.

Fix 2: Reduce Token Usage Per Request

If your bottleneck is ITPM or OTPM rather than RPM, backoff won't help much — you'll just keep hitting the same ceiling. Instead:

A 20% reduction in average prompt size often buys you a proportional increase in effective throughput without touching your rate limit tier at all.

Fix 3: Queue and Throttle Client-Side

Rather than firing requests as fast as your application generates them, add a queue that enforces your known RPM/TPM ceiling before requests ever leave your server:

class RateLimitedQueue {
  constructor(maxPerMinute) {
    this.maxPerMinute = maxPerMinute;
    this.timestamps = [];
  }

  async acquire() {
    const now = Date.now();
    this.timestamps = this.timestamps.filter((t) => now - t < 60000);
    if (this.timestamps.length >= this.maxPerMinute) {
      const wait = 60000 - (now - this.timestamps[0]);
      await new Promise((r) => setTimeout(r, wait));
    }
    this.timestamps.push(Date.now());
  }
}

This converts bursty, unpredictable traffic into a smooth stream that stays under the ceiling — far more reliable than reacting to 429s after they happen.

Fix 4: Spread Load or Upgrade Your Tier

If you're consistently saturating your limits with legitimate traffic, the real fix is capacity, not code. You have two paths:

  1. Request a rate limit increase directly from Anthropic, which usually requires demonstrated usage history and sometimes a higher spend commitment.
  2. Route traffic through an API layer that handles key pooling, retries, and usage visibility for you, so you're not building and maintaining this infrastructure yourself.

This is where a service like SubToAPI fits in: it turns your Claude access into a clean HTTPS API with application-level keys (sub_live_...), per-key usage metadata, and team seats, so you can see exactly which key or team member is driving token consumption before it turns into a 429 storm. Check the quickstart to see how requests are structured, or look at streaming and tool use docs if your rate-limit pressure is coming from long-running or multi-step calls. Plans start with a free trial at signup, and tier details are on pricing.

Monitoring to Prevent Repeat 429s

Fixing one 429 incident doesn't prevent the next one. Add basic observability:

Rate limits are a capacity problem, not a bug — treating them like one you can monitor and plan around, rather than firefight, is the actual long-term fix.

questions

Does retrying immediately after a 429 make things worse? Yes. Immediate retries without backoff add more load to an already-saturated limit window and often trigger repeated 429s. Always wait at least until the retry-after value or reset timestamp before retrying.

Is a 429 the same as a 529 overloaded error? No. A 429 means you've exceeded your account's rate limit. A 529 means Anthropic's servers are temporarily overloaded regardless of your limit — backoff helps with both, but only a tier increase or usage reduction fixes recurring 429s.

Can I avoid 429s entirely by using a third-party API wrapper? A wrapper won't eliminate Anthropic's underlying limits, but tools like SubToAPI add visibility into per-key usage and streaming behavior, making it easier to spot and fix the specific request pattern causing the errors before it recurs.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →