Claude API Billing API for SaaS Products: A Guide
If you're building a SaaS product on top of Claude and searching for a "billing API," you're almost certainly running into the same wall every team hits: Anthropic bills your organization as a whole, but you need to bill your customers individually, based on how much of Claude they actually use. There is no built-in Claude feature that splits usage by end customer, tracks cost per feature, or exposes a clean API for your own billing system to consume.
The practical answer is to put a billing/metering layer between your product and the Claude API — either built in-house from the token counts Claude returns on every response, or via a service that already does this, like SubToAPI, which issues per-application API keys with usage metadata so you can attribute cost without building the tracking infrastructure yourself.
Why Claude API Billing Is Your Problem, Not Anthropic's
Anthropic's console shows you aggregate usage and cost for your account. That's useful for your own finance team, but it tells you nothing about which customer, workspace, or feature inside your product generated those tokens. If you're charging customers per seat, per usage tier, or metered by consumption, you need to build that attribution layer yourself.
The core data you need is in every API response: the usage object with input_tokens and output_tokens. Pricing differs by model and by input vs. output, and now also by cache reads/writes if you use prompt caching. None of that is automatically tied to "customer A used $3.40 this month" — you have to do that math.
What a Real Billing Layer Needs
For a SaaS product to bill Claude usage correctly, you need at minimum:
- Per-customer or per-key attribution — every request needs to be traceable to an account, not just your org
- Token-to-cost conversion — applied per model, since prices differ between model tiers
- Aggregation over a billing period — daily or monthly rollups you can push to Stripe or your own invoicing
- Visibility into tool use and streaming — tool calls and long streamed responses still consume tokens and need to count
- Alerting or hard limits — so one customer's runaway script doesn't blow your margin on a flat-rate plan
Missing any of these turns "add AI billing" into a multi-week engineering project instead of a config change.
Building It Yourself
If you're calling the Claude API directly, the minimal version looks like this: log the usage field from every response, tag it with the customer ID from your own session, and store it.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this contract."}]
}'
{
"usage": {
"input_tokens": 512,
"output_tokens": 180
}
}
You then need a cron job or event stream that converts those numbers into dollars per model, groups them by customer, and feeds a usage record to Stripe's metered billing API (or whatever invoicing system you run). Streaming responses complicate this further because token counts arrive at the end of the stream, and tool-use loops mean one "user action" can trigger several API calls, each with its own usage.
This is entirely doable, but it's infrastructure you have to maintain: a logging pipeline, a pricing table that you update every time Anthropic changes rates, and reconciliation logic for retries and errors.
Using a Metering Layer Instead
SubToAPI exists specifically to remove this layer of work. Instead of calling Claude directly with one shared organization key, you issue a sub_live_... key per customer, per app, or per environment from your dashboard. Every request made with that key carries usage metadata automatically — input/output tokens, model, and timing — without you writing a logging pipeline.
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{ role: "user", content: "Summarize this contract." }],
}),
});
const data = await res.json();
console.log(data.usage); // per-request token data tied to this key
Because each application key maps to a customer or tenant in your own system, pulling usage for billing becomes a matter of querying that key's activity rather than parsing raw logs across your whole fleet. This works the same way whether you're calling messages, streaming, or tool use endpoints — the usage metadata is attached regardless of request type.
For teams, SubToAPI also supports seat-based access, so you can mirror your own pricing model (per-seat SaaS tiers billed to you, at €9 Solo, €19/seat Team, or €49/seat Scale) against the way you're already billing your customers, instead of maintaining a separate internal accounting system just to track Claude spend. Check /pricing for the current plan breakdown, or start with a free trial at /signup.
A Practical Architecture
A common pattern for SaaS teams:
- Issue one SubToAPI key per customer account (or per internal service, if you bill per feature rather than per customer)
- Route all Claude calls for that customer through their key
- Pull usage metadata on a schedule (daily is usually enough) from the dashboard or your logging
- Convert to a Stripe usage record or your invoicing line item
- Set soft alerts on customers approaching unusual usage, so support can reach out before a surprise invoice
This keeps your billing code focused on "convert usage events into invoice lines" rather than "figure out how many tokens Claude actually used and who they belong to." If you're starting fresh, the quickstart walks through getting a key issued and your first request made in a few minutes.
Questions
Does Anthropic provide per-customer billing natively? No. The Claude API bills your organization as a whole based on total token usage. Splitting that by customer, feature, or tenant is something you have to build or get from a layer like SubToAPI.
What's the simplest way to estimate cost per request? Multiply input_tokens and output_tokens from the response's usage object by the current per-token price for the model you called, then sum per customer over your billing period.
Can I bill customers for streaming or tool-use requests the same way? Yes — streamed and tool-enabled requests still return standard usage data once complete; you just need to make sure your logging captures the final usage object after the stream or tool loop finishes, not partway through.