Claude API Cost Per Request Calculator (Formula)
If you're trying to figure out what a single Claude API call actually costs, the answer isn't a fixed number — it depends entirely on how many tokens go in and how many come out. There's no flat "per request" fee. Instead, cost is a function of input tokens × input price + output tokens × output price, and both prices vary by model.
This article gives you the exact formula, a working calculator you can drop into any project, and the variables that actually move your bill: system prompts, tool use, caching, and streaming.
The core formula
Every Claude API response includes a usage object with input_tokens and output_tokens for that specific call. Multiply each by its per-million-token price, divide by a million, and add them together:
cost = (input_tokens / 1,000,000) * input_price_per_million
+ (output_tokens / 1,000,000) * output_price_per_million
That's it. There's no separate charge for the request itself, for headers, or for the model "thinking" — only for tokens processed and tokens generated. Input and output are priced differently (output is typically several times more expensive than input), so a request with a short prompt and a long answer costs more than one with a long prompt and a short answer, even if the total token count is the same.
Because pricing changes over time and differs across model tiers (fast/cheap models vs. large/capable ones), plug in the current per-million rates for the model you're using rather than relying on a hardcoded number — this calculator is about the method, not a fixed price table.
A working cost calculator
Here's a small function you can use in a Node script, a serverless function, or a browser console to estimate cost per request:
function estimateCost({
inputTokens,
outputTokens,
inputPricePerMillion,
outputPricePerMillion,
}) {
const inputCost = (inputTokens / 1_000_000) * inputPricePerMillion;
const outputCost = (outputTokens / 1_000_000) * outputPricePerMillion;
return {
inputCost: Number(inputCost.toFixed(6)),
outputCost: Number(outputCost.toFixed(6)),
totalCost: Number((inputCost + outputCost).toFixed(6)),
};
}
const result = estimateCost({
inputTokens: 1200,
outputTokens: 400,
inputPricePerMillion: 3, // replace with current input price
outputPricePerMillion: 15, // replace with current output price
});
console.log(result);
// { inputCost: 0.0036, outputCost: 0.006, totalCost: 0.0096 }
Swap in the token counts and prices for your specific model, and you have an accurate per-request cost. To scale this to a monthly estimate, multiply by your expected daily request volume and by 30.
Where to get real token counts
Don't guess at token counts — pull them from the actual API response. Every Claude API call (and any proxy built on top of it, including SubToAPI) returns usage data with the response:
{
"usage": {
"input_tokens": 842,
"output_tokens": 213
}
}
If you're building your own dashboard, log this usage object alongside a timestamp and request ID for every call. That log is your real cost-per-request calculator — far more accurate than estimating tokens from character counts (roughly 4 characters per token in English text, but this varies a lot with code, JSON, and non-English languages).
If you're routing requests through SubToAPI, the same usage metadata is returned on every call and surfaced per API key in the dashboard, so you don't need to build this logging yourself — see /docs/messages for the response shape.
Variables that change the real cost
The formula is simple, but several things inflate token counts in ways people forget to account for:
- System prompts — a long system prompt is billed as input tokens on every single request, not once per session. A 2,000-token system prompt sent 10,000 times a day adds up fast.
- Conversation history — if you resend prior turns for multi-turn context, each turn's tokens are counted again as input on every follow-up request.
- Tool definitions — when using tool use / function calling, the tool schemas you pass in count as input tokens on every call. See /docs/tools for how tool definitions are structured.
- Streaming — streaming responses (see /docs/streaming) don't change the token count or price; they only change when you receive the tokens, not how many you're billed for.
- Prompt caching — if the model supports caching repeated prefixes (like a large system prompt or long context), cached input tokens are typically billed at a lower rate than fresh input tokens. If you're sending the same large context repeatedly, this is the single biggest lever for reducing cost per request.
- Retries — a failed request that you retry doesn't refund the original attempt's tokens if it partially processed. Build idempotency and backoff carefully so retries don't silently double your bill.
Estimating cost before you build
If you're scoping a feature and want a rough number before writing code, work backward from a typical interaction:
- Estimate average input tokens (system prompt + user message + any retrieved context).
- Estimate average output tokens (a short answer vs. a long structured response).
- Multiply by your expected daily volume.
- Run it through the formula above with current pricing for the model you plan to use.
This is also where model choice matters as much as prompt design — a smaller, faster model can cost a fraction of a larger one for the same task, so it's worth benchmarking accuracy against cost before committing to a tier for a high-volume feature.
Tracking cost across a team
A calculator tells you what should happen. In practice, teams lose track of cost because usage is scattered across personal API keys, side projects, and staging environments with no central view. If you're distributing access across a team, having per-key usage broken down by request — not just a monthly total — makes it much easier to catch a runaway loop or an expensive prompt before it shows up on an invoice. SubToAPI's dashboard shows usage per application key alongside your pricing plan, which is useful once you have more than one service or teammate calling the API. Start with a free trial at /signup if you want to see this without committing.
Questions
Does Claude API charge per request or per token? Per token. There's no flat request fee — you pay for input tokens (what you send) and output tokens (what the model generates), each at its own rate.
Why do two requests with the same word count cost differently? Token count isn't the same as word count, and input/output tokens are priced differently. A response with more output tokens (a longer answer) costs more than one with more input tokens, even at similar total length.
How do I get exact token counts for a past request instead of estimating? Read the usage.input_tokens and usage.output_tokens fields returned with every API response — see /docs/messages — and log them per request rather than estimating from text length.