How to Set a Claude API Budget Cap Per User
If you're building a product on top of Claude and multiple users or team members share one Anthropic account, you've probably hit the same wall: Anthropic's API lets you set spend limits at the organization level, but not per individual user. There's no dashboard toggle that says "cap Alice at €20/month" and "cap Bob at €5/month." If you need that, you have to build it yourself.
This is a common requirement for SaaS products with per-seat Claude features, internal tools shared across a team, or agencies reselling AI assistants to clients. Below is how to actually implement a per-user budget cap on the Claude API, plus a simpler route if you'd rather not maintain the plumbing yourself.
Why Anthropic's native limits aren't enough
Anthropic's console gives you:
- Organization-wide rate limits (requests per minute, tokens per minute)
- A single spend/usage dashboard for the whole account
- One pool of API keys, typically shared across your backend
None of this is scoped to an individual end user of your product. If ten people share one sk-ant-... key, Anthropic has no concept of "user 7 spent €14 this month." That tracking has to happen in your own application layer.
The core pattern: track spend per user, enforce before the call
A working per-user budget cap needs three pieces:
- A per-user usage ledger — every Claude call is logged with input tokens, output tokens, model, and cost
- A cap check that runs before the request, not after
- A reset window (daily, monthly, or rolling) so caps don't become permanent locks
1. Log every call with cost, not just tokens
Token counts alone aren't useful for a budget cap — you need actual cost, since input and output tokens are priced differently per model.
function estimateCost(model, inputTokens, outputTokens) {
const prices = {
"claude-sonnet": { input: 3 / 1e6, output: 15 / 1e6 },
"claude-haiku": { input: 0.8 / 1e6, output: 4 / 1e6 },
};
const p = prices[model];
return inputTokens * p.input + outputTokens * p.output;
}
Store this per call, tagged with a userId, in whatever database you already use.
2. Check the cap before you call Claude
This is the part teams get wrong most often — they log usage after the response comes back, which means a user can blow past their cap on the last request before the system catches up. Check the cumulative spend first.
async function callClaudeWithBudgetCheck(userId, payload) {
const spentThisMonth = await getUserSpend(userId); // from your DB
const cap = await getUserCap(userId); // e.g. 5.00 (EUR)
if (spentThisMonth >= cap) {
throw new Error("Budget cap reached for this user");
}
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
const data = await response.json();
const cost = estimateCost(
payload.model,
data.usage.input_tokens,
data.usage.output_tokens
);
await recordUsage(userId, cost);
return data;
}
This is a soft cap (one request could slightly overshoot if it's a long completion), but for practical purposes — stopping runaway usage from a single user — it works well. If you need a hard cap, estimate the max possible cost from max_tokens before sending the request and reject upfront if the worst case would exceed the remaining budget.
3. Pick a reset window that matches your billing
Monthly is the most common choice if you bill users monthly, but if you're protecting against abuse rather than billing, a rolling 24-hour window is often safer — it prevents someone from front-loading an entire month's budget in one afternoon and then complaining they're locked out for 29 days.
A simpler path: one API key per user
The approach above works, but it means you're maintaining a ledger, a reset cron job, and reconciliation logic forever. An alternative is to give each user or team member their own application API key instead of sharing one Anthropic key across your whole user base.
This is exactly what SubToAPI is for: it turns your existing Claude access into application API keys (sub_live_...) that you can issue per user, per team seat, or per environment, each with its own usage metadata. Instead of tracking spend in your own database from scratch, you pull per-key usage from the dashboard or API and apply your cap logic against that — one key per user means the accounting boundary already matches the thing you're trying to limit.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this ticket."}]
}'
Team plans add seats (so each person on your team gets their own key under one billing relationship), which makes "cap per user" a much more natural operation — you're capping per key, and each key already maps to a person. See /docs/quickstart for setup and /docs/messages for the request format. If your budget-cap logic needs to react to partial output before cutting a user off mid-stream, check /docs/streaming; if you're capping agentic workflows that call tools, /docs/tools covers tool-use requests specifically, since a single user turn can trigger several model calls.
Practical tips regardless of approach
- Warn before you block. A hard cutoff with no warning generates support tickets. Send a notification at 80% of the cap.
- Separate "cap reached" from "rate limited." Users need different messaging and different retry behavior for each.
- Account for retries and tool loops. A single user action can trigger multiple Claude calls (tool use, retries on errors). Make sure your ledger captures all of them, not just the first call per request.
- Decide what happens at zero. Hard block, downgrade to a cheaper model, or queue until reset — pick one deliberately instead of letting it be undefined behavior.
FAQ
Does Anthropic's API support per-user spend limits natively?
No. Anthropic's console offers organization-level rate limits and usage dashboards, but not limits scoped to individual end users. Per-user caps have to be implemented in your application layer or by issuing separate keys per user.
What's the difference between a rate limit and a budget cap?
A rate limit controls requests or tokens per minute to prevent bursts; a budget cap controls total spend over a period (daily, monthly) to control cost. You typically need both — rate limits to prevent abuse in the moment, budget caps to control the bill over time.
Is giving every user their own API key a reasonable strategy?
Yes, and it's often simpler than building a shared-key usage ledger. Per-user keys let you track and cap spend per key directly instead of reconciling logs against user IDs. See /pricing for plans that support per-seat keys.