Claude API Rate Limit Headers Explained
Every response from the Claude API carries a set of headers that tell you exactly how much of your rate limit you've used and how much is left. Most developers ignore them until they hit a 429, but reading these headers proactively lets you throttle your own request rate before you ever get rejected.
This article walks through each header Claude returns, what the numbers actually mean, and how to build simple client-side logic around them so your app degrades gracefully instead of crashing into a wall.
Why Claude uses header-based rate limiting
Claude enforces limits per organization and per model, measured across three dimensions: requests per minute (RPM), input tokens per minute (ITPM), and output tokens per minute (OTPM). Rather than making you guess where you stand, the API returns your current limit, remaining quota, and reset time on every single response — successful or not. That's the whole point of the headers: you don't need a separate endpoint to check your usage, and you don't need to keep your own counter in sync with Anthropic's servers.
The headers, one by one
Claude returns a family of headers prefixed anthropic-ratelimit-, split by the three limit types:
anthropic-ratelimit-requests-limit: 50
anthropic-ratelimit-requests-remaining: 47
anthropic-ratelimit-requests-reset: 2024-06-01T12:34:56Z
anthropic-ratelimit-input-tokens-limit: 40000
anthropic-ratelimit-input-tokens-remaining: 38210
anthropic-ratelimit-input-tokens-reset: 2024-06-01T12:34:56Z
anthropic-ratelimit-output-tokens-limit: 8000
anthropic-ratelimit-output-tokens-remaining: 7650
anthropic-ratelimit-output-tokens-reset: 2024-06-01T12:34:56Z
retry-after: 12
Here's what each one is doing:
-limit— the ceiling for that dimension in the current window. This reflects your account tier or negotiated limit, not something you set.-remaining— how much of that ceiling is left before the window resets. This is the number you should actually watch.-reset— an ISO 8601 timestamp for when the counter refills. Windows are rolling, not fixed-clock, so this timestamp moves with every request.retry-after— only present when you've been rate limited (HTTP 429). It's a number of seconds, not a timestamp, telling you the minimum wait before retrying.
Note that three separate limits exist simultaneously. You can have plenty of requests-remaining but run out of output-tokens-remaining if you're generating long completions. A single large request can burn through your token budget while barely touching your request-count budget — so don't assume checking one header is enough.
Reading headers in practice
A minimal check after each call looks like this:
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello" }]
})
});
const remaining = res.headers.get("anthropic-ratelimit-output-tokens-remaining");
const reset = res.headers.get("anthropic-ratelimit-output-tokens-reset");
if (Number(remaining) < 500) {
console.warn(`Low on output tokens (${remaining} left), resets at ${reset}`);
// slow down, queue, or switch to a smaller max_tokens value
}
The pattern that works well in production: track the lowest -remaining value across the three dimensions after every call, and if any of them drops below a threshold (say, 10% of the limit), start adding artificial delay between requests. This is cheaper and more predictable than waiting for a 429 and reacting after the fact.
Limits are per-model and per-org, not per-request
A subtlety worth knowing: these headers describe the limit for the model you just called, under the org tied to your API key. If you're calling Claude 3.5 Sonnet and Claude 3 Opus from the same key, they have independent budgets — checking headers from a Sonnet response tells you nothing about your Opus quota. If you run multiple models or workloads, track each one's headers separately rather than assuming a single global counter.
Also worth noting: limits scale with usage tier, which increases automatically as your account has a track record of successful billing and traffic. If you're seeing tight limits early on, they typically loosen over the first weeks of steady usage — no action needed on your part beyond normal billing.
Where this matters for multi-key or team setups
If you're distributing Claude access across multiple applications, services, or team members using a single underlying Anthropic account, rate limit headers get complicated fast — every internal consumer is drawing from the same shared pool, and there's no built-in way to see who's using what. This is one of the problems SubToAPI solves: it sits between your apps and your Claude access, issuing separate sub_live_... keys per application or team member, so each consumer has its own visibility into usage instead of everyone fighting over one shared header set. You still get the same streaming and tool-use behavior described in the messages and tools docs, just with per-key usage metadata layered on top. Check the quickstart or pricing if that's a problem you're running into.
Practical takeaways
- Watch all three
-remainingheaders, not just one — a token-heavy request can exhaust your budget while your request count looks fine. - Use
-resettimestamps to schedule retries precisely instead of guessing wait times. - Treat
retry-afteras authoritative when present — it accounts for server-side state you can't compute client-side. - Log header values over time if you're debugging intermittent 429s; a graph of
-remainingover a day usually reveals the actual bottleneck immediately.
FAQ
Do rate limit headers count against my usage? No. Reading response headers costs nothing extra — they're included on every API response regardless of outcome, successful or failed.
Why do I see different -remaining values for input and output tokens on the same request? Input and output tokens are billed and limited separately. A request with a short prompt but a long generated response will consume output-token budget much faster than input-token budget.
Can I increase my rate limits directly through these headers? No, the headers are read-only reporting. Limits increase based on account tier and usage history, or through a direct request to Anthropic for enterprise accounts.