Claude API Usage Limits Per Plan: What You Need to Know
If you're asking about Claude API usage limits per plan, you're probably trying to figure out one of two things: how much you can send before getting throttled, or why your app suddenly started returning 429 errors. The short answer is that Claude's usage limits are tiered by account type (free, Build/Pay-as-you-go, and Scale/Enterprise) and are measured in requests per minute (RPM), tokens per minute (TPM), and sometimes tokens per day (TPD) — not a flat "number of calls per month."
This matters because limits aren't static. They scale with your usage tier, your spend history, and the specific model you're calling (Opus, Sonnet, and Haiku each have separate limit pools). Below is a practical breakdown of how these limits work, what typically trips people up, and how to architect around them.
How Claude API limits are actually structured
Anthropic's API limits are organized around three dimensions:
- RPM (requests per minute) — the number of API calls you can make in a 60-second window
- TPM (tokens per minute) — the combined input + output token volume per minute
- TPD (tokens per day) — a daily ceiling on some lower tiers
Each model family (Claude Opus, Sonnet, Haiku) has its own independent limit bucket. Burning through your Sonnet TPM doesn't touch your Haiku allowance, which is useful if you route cheap/fast tasks to Haiku and reserve Opus for complex reasoning.
Limits also increase automatically as you spend more and build account history. New accounts start conservative; after sustained usage and successful billing cycles, Anthropic raises your tier without you asking. Enterprise customers can negotiate custom limits directly.
Why this trips people up in production
Three recurring problems:
- Bursty traffic — a feature launch or marketing push spikes RPM instantly, even though your daily token volume is fine.
- Token miscounting — long system prompts, tool definitions, and conversation history all count toward TPM, not just the user's message.
- No visibility — the standard API gives you headers like
anthropic-ratelimit-requests-remaining, but most teams don't log or alert on them until they're already getting 429s in production.
Reading rate limit headers correctly
Every response from the API includes headers that tell you exactly where you stand:
anthropic-ratelimit-requests-limit: 1000
anthropic-ratelimit-requests-remaining: 842
anthropic-ratelimit-tokens-limit: 100000
anthropic-ratelimit-tokens-remaining: 76210
The practical move is to parse these on every response and back off proactively — don't wait for a 429. A simple client-side throttle that checks remaining against a threshold (say, 10% of limit) and adds delay before firing the next batch of requests will save you from cascading failures during traffic spikes.
Planning capacity: a checklist
Before you ship a feature that depends on Claude, estimate:
- Average tokens per request (input + output, including system prompt and any tool schemas)
- Peak concurrent requests during your busiest hour, not your daily average
- Which model you're calling for which workload — don't default everything to Opus if Haiku or Sonnet can handle it
- Retry behavior — exponential backoff with jitter, capped at a sane number of attempts, so a rate-limit event doesn't turn into a request storm
If your peak RPM regularly exceeds your current tier, you have two paths: request a limit increase (which requires sustained billing history) or add a queue/throttle layer in front of your own API so requests smooth out before they hit Claude.
Working around limits with a managed layer
A lot of teams building internal tools or customer-facing features don't want to manage rate-limit headers, retry logic, and multi-model routing themselves. That's the gap SubToAPI fills: it turns your existing Claude access into a clean HTTPS API with its own application keys (sub_live_...), so you get predictable request handling, streaming, tool use, and usage metadata per key without re-implementing backoff logic for every service that calls Claude.
Instead of every microservice independently tracking its own rate-limit state, you point them at your SubToAPI endpoint and manage capacity centrally. A basic call looks like this:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this ticket."}]
}'
Each application key gets its own usage metadata, so if one internal tool is eating your token budget, you see it per-key in the dashboard instead of guessing from aggregate logs. Team and Scale plans add seats so multiple engineers or services can each have scoped keys instead of sharing one credential and one limit pool blindly. See /docs/quickstart to get a key running in a few minutes, or /docs/messages for the full request format.
Streaming and tool use under rate limits
Streaming responses don't change your token accounting — you're still billed and limited on total tokens, just delivered incrementally. The benefit is perceived latency, not reduced limit pressure. If you're building chat-style UIs, check /docs/streaming for how to handle partial tokens without blocking on the full response.
Tool use (function calling) adds tokens for the tool schema itself on every request, which is easy to underestimate when budgeting TPM. If you have five tools defined with verbose descriptions, that schema gets sent and counted every single call. Trim tool definitions to what's actually needed; see /docs/tools for guidance on structuring tool calls efficiently.
Choosing a plan based on real usage patterns
If you're evaluating whether to manage Claude's native limits yourself or route through a layer like SubToAPI, the deciding factor is usually team size and visibility needs, not raw volume. A solo developer prototyping an internal tool can get by with direct API access and careful backoff logic. A team shipping a customer-facing feature across multiple services usually wants per-key usage breakdowns and centralized key rotation — that's where /pricing plans like Solo, Team, and Scale start making sense, since they map naturally to how many people or services need independent, trackable access.
Check /signup for a free trial if you want to test request handling and usage metadata before committing to a plan.
Questions
Do Claude API limits reset daily or per minute? Most limits are measured per minute (RPM and TPM), though some account tiers also have a token-per-day ceiling. Check the rate-limit response headers on each call for your exact current limits.
Does streaming reduce how many tokens count against my limit? No. Streaming changes delivery speed and perceived latency, not token accounting — the same total input and output tokens count against your TPM limit either way.
Can I increase my Claude API limits without contacting sales? Yes, for Build-tier accounts, limits typically rise automatically based on sustained usage and billing history. Enterprise-level increases usually require direct negotiation.