← Blog

Claude API Rate Limiting: Best Practices Guide

2026-10-11 · 6 min read · SubToAPI Team

Rate limiting is the single most common reason production apps built on Claude break under real traffic. You test locally with a handful of requests, ship to production, and suddenly users are seeing failed completions during peak hours because you hit a 429 and had no plan for it.

This guide covers how Claude's rate limits actually work (requests per minute, tokens per minute, and tiered limits by usage history), and the concrete patterns you should implement: exponential backoff, request queuing, token budgeting, and multi-key load distribution. None of this is theoretical — these are the same patterns you need whether you're calling Anthropic directly or routing through a gateway.

How Claude API rate limits work

Anthropic enforces limits along several axes simultaneously:

The practical implication: you can hit a 429 even if you're nowhere near your RPM cap, simply because a few long-context requests consumed your token budget for that minute. Rate limiting isn't just about request count — token volume is usually the binding constraint for anything beyond short chat messages.

Read the response headers, always

Every response (success or 429) includes headers that tell you exactly where you stand:

anthropic-ratelimit-requests-remaining
anthropic-ratelimit-tokens-remaining
anthropic-ratelimit-requests-reset
anthropic-ratelimit-tokens-reset
retry-after

If you're not parsing these, you're flying blind. A well-behaved client should track remaining capacity and throttle proactively, rather than waiting to get rejected.

const res = await fetch(url, options);
const remaining = Number(res.headers.get("anthropic-ratelimit-tokens-remaining"));
const resetAt = res.headers.get("anthropic-ratelimit-tokens-reset");

if (remaining < 2000) {
  // slow down before you get a 429, not after
  await sleep(1000);
}

Implement exponential backoff with jitter

When you do get a 429, the naive retry-immediately approach just adds load to an already-throttled window and makes things worse. Use exponential backoff with jitter:

async function callWithBackoff(fn, maxRetries = 5) {
  let attempt = 0;
  while (attempt < maxRetries) {
    try {
      return await fn();
    } catch (err) {
      if (err.status !== 429 && err.status !== 529) throw err;
      const retryAfter = Number(err.headers?.get("retry-after")) || 0;
      const backoff = retryAfter * 1000 || Math.min(1000 * 2 ** attempt, 30000);
      const jitter = Math.random() * 300;
      await sleep(backoff + jitter);
      attempt++;
    }
  }
  throw new Error("Max retries exceeded");
}

Key details that matter in production:

Queue requests instead of firing them all at once

If your app sends bursts — batch jobs, bulk document processing, a cron that wakes up and processes 500 queued items — don't fire them concurrently. Use a queue with a concurrency limit and a token-aware rate limiter:

import pLimit from "p-limit";

const limit = pLimit(5); // max 5 concurrent requests

const results = await Promise.all(
  items.map((item) => limit(() => callWithBackoff(() => callClaude(item))))
);

For token-heavy workloads, concurrency alone isn't enough — you also need to estimate token usage per request and throttle based on your ITPM budget, not just request count. A simple leaky-bucket counter that tracks estimated tokens sent in the last 60 seconds works well here.

Separate latency-sensitive traffic from batch traffic

A common failure mode: a background job processing a CSV of 10,000 rows shares the same API key and rate limit pool as your live chat feature. When the batch job saturates your token budget, live users start getting 429s.

Fixes:

This is exactly the kind of operational problem that gets harder to manage as a team grows, which is one reason teams move to a layer that handles key issuance and quota separation per application rather than hand-rolling it. SubToAPI issues separate sub_live_... application keys per project so your batch pipeline and your production chat feature aren't competing for the same bucket, and gives you usage metadata per key so you can see which one is actually hitting limits.

Cache and dedupe before you even make the call

The cheapest way to avoid rate limits is to not make the request. If users frequently ask near-identical questions, or your pipeline reprocesses the same documents, cache responses keyed on a normalized hash of the input. This reduces both RPM and token pressure simultaneously, and it's worth doing before you invest in sophisticated retry logic.

Monitor before you hit the wall

Set up alerting on your token-remaining headers, not just on 429 counts. If you only react after errors start, you've already degraded user experience. Track:

If you're integrating Claude through SubToAPI, streaming responses and usage metadata are available per request, which makes it straightforward to build this kind of dashboard without parsing raw headers yourself — see the streaming docs and messages API reference for the request/response shape.

Quick checklist

FAQ

What's the difference between a 429 and a 529 from the Claude API? A 429 means you've exceeded your own account's rate limit (RPM, ITPM, or OTPM). A 529 means the service itself is overloaded, independent of your limits. Both should trigger a retry with backoff, but only 429s are something you can fix by throttling your own traffic.

Does streaming count differently against rate limits? Streamed output tokens still count toward your OTPM limit as they're generated, not just at the end of the response. Long streaming completions can consume your token budget for the minute even though it's a single request — plan concurrency limits with that in mind.

Can I get higher rate limits without changing providers? Anthropic increases limits as account spend and usage history grow (usage tiers). If you need predictable, separated limits per project or team sooner than that, issuing distinct API keys per application — as supported in SubToAPI — lets you isolate quota pressure without waiting on tier upgrades.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →