← Blog

Claude API Rate Limit Headers Explained

2026-09-29 · 4 min read · SubToAPI Team

Every response from the Claude API carries a set of headers that tell you exactly how much of your rate limit you've used and how much is left. Most developers ignore them until they hit a 429, but reading these headers proactively lets you throttle your own request rate before you ever get rejected.

This article walks through each header Claude returns, what the numbers actually mean, and how to build simple client-side logic around them so your app degrades gracefully instead of crashing into a wall.

Why Claude uses header-based rate limiting

Claude enforces limits per organization and per model, measured across three dimensions: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Rather than making you guess where you stand, the API returns your current limit, remaining quota, and reset time on every single response — successful or not. That's the whole point of the headers: you don't need a separate endpoint to check your usage, and you don't need to keep your own counter in sync with Anthropic's servers.

The headers, one by one

Claude returns a family of headers prefixed anthropic-ratelimit-, split by the three limit types:

anthropic-ratelimit-requests-limit: 50
anthropic-ratelimit-requests-remaining: 47
anthropic-ratelimit-requests-reset: 2024-06-01T12:34:56Z

anthropic-ratelimit-input-tokens-limit: 40000
anthropic-ratelimit-input-tokens-remaining: 38210
anthropic-ratelimit-input-tokens-reset: 2024-06-01T12:34:56Z

anthropic-ratelimit-output-tokens-limit: 8000
anthropic-ratelimit-output-tokens-remaining: 7650
anthropic-ratelimit-output-tokens-reset: 2024-06-01T12:34:56Z

retry-after: 12

Here's what each one is doing:

Note that three separate limits exist simultaneously. You can have plenty of requests-remaining but run out of output-tokens-remaining if you're generating long completions. A single large request can burn through your token budget while barely touching your request-count budget — so don't assume checking one header is enough.

Reading headers in practice

A minimal check after each call looks like this:

const res = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
    "content-type": "application/json"
  },
  body: JSON.stringify({
    model: "claude-3-5-sonnet-20241022",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Hello" }]
  })
});

const remaining = res.headers.get("anthropic-ratelimit-output-tokens-remaining");
const reset = res.headers.get("anthropic-ratelimit-output-tokens-reset");

if (Number(remaining) < 500) {
  console.warn(`Low on output tokens (${remaining} left), resets at ${reset}`);
  // slow down, queue, or switch to a smaller max_tokens value
}

The pattern that works well in production: track the lowest -remaining value across the three dimensions after every call, and if any of them drops below a threshold (say, 10% of the limit), start adding artificial delay between requests. This is cheaper and more predictable than waiting for a 429 and reacting after the fact.

Limits are per-model and per-org, not per-request

A subtlety worth knowing: these headers describe the limit for the model you just called, under the org tied to your API key. If you're calling Claude 3.5 Sonnet and Claude 3 Opus from the same key, they have independent budgets — checking headers from a Sonnet response tells you nothing about your Opus quota. If you run multiple models or workloads, track each one's headers separately rather than assuming a single global counter.

Also worth noting: limits scale with usage tier, which increases automatically as your account has a track record of successful billing and traffic. If you're seeing tight limits early on, they typically loosen over the first weeks of steady usage — no action needed on your part beyond normal billing.

Where this matters for multi-key or team setups

If you're distributing Claude access across multiple applications, services, or team members using a single underlying Anthropic account, rate limit headers get complicated fast — every internal consumer is drawing from the same shared pool, and there's no built-in way to see who's using what. This is one of the problems SubToAPI solves: it sits between your apps and your Claude access, issuing separate sub_live_... keys per application or team member, so each consumer has its own visibility into usage instead of everyone fighting over one shared header set. You still get the same streaming and tool-use behavior described in the messages and tools docs, just with per-key usage metadata layered on top. Check the quickstart or pricing if that's a problem you're running into.

Practical takeaways

FAQ

Do rate limit headers count against my usage? No. Reading response headers costs nothing extra — they're included on every API response regardless of outcome, successful or failed.

Why do I see different -remaining values for input and output tokens on the same request? Input and output tokens are billed and limited separately. A request with a short prompt but a long generated response will consume output-token budget much faster than input-token budget.

Can I increase my rate limits directly through these headers? No, the headers are read-only reporting. Limits increase based on account tier and usage history, or through a direct request to Anthropic for enterprise accounts.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →