Claude API Cost Per Request: A Full Breakdown
What determines the cost of a single Claude API request
The cost of a single Claude API request is calculated from three things: the model you call, the number of input tokens you send, and the number of output tokens the model generates. There's no flat "per request" fee — Anthropic bills per token, in two separate buckets with different rates, and the final number per request can range from a fraction of a cent to several dollars depending on what you're doing.
That means "cost per request" isn't a fixed figure you can look up — it's a formula you apply to your own traffic. Below is the exact breakdown of what goes into that formula, where the hidden costs hide, and how to estimate it before you ship a feature.
The core formula
Every request costs:
request_cost = (input_tokens / 1,000,000 × input_price_per_million)
+ (output_tokens / 1,000,000 × output_price_per_million)
Input and output are priced differently — output tokens are consistently more expensive than input tokens across every Claude model, usually by a factor of 4–5x. This matters more than most people assume: a chatbot that reads a 2,000-token document and replies with 50 tokens costs almost nothing per call. A code generator that reads 500 tokens and writes 3,000 tokens of code can cost 10x more per call even though the input is far smaller.
What counts as "input tokens"
Input tokens aren't just the user's message. They include:
- The system prompt, every time, on every request
- Full conversation history if you're not using prompt caching
- Tool definitions you pass in the request
- Any documents, retrieved context, or RAG chunks you inject
- Few-shot examples baked into your prompt
This is the part that catches teams off guard. If your system prompt is 800 tokens and you have 10 tool definitions adding another 600 tokens, every single request pays for 1,400 tokens before the user has typed a word.
What counts as "output tokens"
Output tokens are everything the model generates in its response — including the content inside tool calls (the function name and arguments Claude generates when using tools), and any reasoning or intermediate text depending on the model and settings you use.
A worked example
Say a model is priced at $3 per million input tokens and $15 per million output tokens (illustrative numbers — always check current pricing for the exact model you're using).
Simple Q&A request:
- Input: 200 tokens (system prompt + user question)
- Output: 150 tokens (short answer)
- Cost: (200/1,000,000 × 3) + (150/1,000,000 × 15) = $0.0006 + $0.00225 = $0.00285
Document summarization:
- Input: 8,000 tokens (full document + instructions)
- Output: 300 tokens (summary)
- Cost: (8,000/1,000,000 × 3) + (300/1,000,000 × 15) = $0.024 + $0.0045 = $0.0285
Long-form generation with tool use:
- Input: 1,500 tokens (prompt + tool schema + history)
- Output: 2,000 tokens (generated content + tool call arguments)
- Cost: (1,500/1,000,000 × 3) + (2,000/1,000,000 × 15) = $0.0045 + $0.03 = $0.0345
The spread here — $0.003 to $0.035 — is a 12x difference between the cheapest and most expensive call, all within one app that mixes task types. This is why a single "cost per request" number is almost always misleading unless you break it down by endpoint or feature.
Costs most people forget to count
Conversation history growth. In a multi-turn chat, every turn re-sends the entire history as input. Turn 10 of a conversation costs far more in input tokens than turn 1, even if the user's message length stays constant.
Retries on errors or rate limits. A failed request that you retry still counts as two billed attempts if the first one partially completed or if you resend the same payload. Build retry logic that doesn't duplicate full-context calls unnecessarily.
Streaming vs non-streaming. Streaming doesn't change the token cost, but it does change how early you can cancel an unwanted generation — cutting a stream short when you detect a bad response saves output tokens that a blocking call would have paid for in full.
System prompts and tool schemas on every call. These are static costs you pay repeatedly. Trimming an 800-token system prompt to 300 tokens, multiplied across hundreds of thousands of requests, is often the single biggest lever for reducing average cost per request.
How to actually calculate your per-request cost
- Log token usage per request. Every response includes usage metadata with exact input and output token counts — don't estimate with a character count, use the real numbers.
- Segment by feature, not globally. A chat feature and a summarization feature have wildly different cost profiles. Average them together and you'll misforecast both.
- Multiply by expected volume. Cost per request only matters once you multiply it by requests per day/month. A $0.03 request run 50,000 times a day is $1,500/day — know this before launch, not after the invoice.
- Re-check after every prompt change. Adding a few examples to a system prompt, or adding a new tool definition, shifts your baseline input cost on every single call going forward.
If you're building on top of Claude and want this usage tracked automatically instead of parsing it out of raw API responses, SubToAPI gives you per-key usage metadata in the dashboard so you can see cost per request broken down by application key without building your own logging pipeline. It sits on top of your existing Claude access and exposes it as a standard HTTPS API with streaming and tool use support — see the quickstart for setup.
Reducing cost per request without changing the model
- Trim system prompts — every static token is a static cost, forever
- Cap max output tokens for tasks that don't need long responses
- Summarize or truncate conversation history instead of sending full transcripts
- Use prompt caching where available to avoid re-billing static context on every call
- Route by task complexity — not every request needs your most capable (and most expensive) model
Questions
Is Claude API priced per request or per token? Per token, not per request. There's no flat fee per call — you're billed separately for input tokens (what you send) and output tokens (what the model generates), at different rates.
Why does my cost per request vary so much between calls? Because input and output lengths vary by task. Long documents increase input cost; long generated responses increase output cost — and output tokens are priced several times higher than input tokens, so generation-heavy tasks cost disproportionately more.
How can I see the exact cost of each request I make? Check the usage metadata returned with every API response, which includes exact input and output token counts. Multiply by the current per-million pricing for your model, or use a dashboard like SubToAPI's that surfaces this automatically per API key.