← Blog

Claude API Cost Per Request: A Full Breakdown

2026-10-09 · 5 min read · SubToAPI Team

What determines the cost of a single Claude API request

The cost of a single Claude API request is calculated from three things: the model you call, the number of input tokens you send, and the number of output tokens the model generates. There's no flat "per request" fee — Anthropic bills per token, in two separate buckets with different rates, and the final number per request can range from a fraction of a cent to several dollars depending on what you're doing.

That means "cost per request" isn't a fixed figure you can look up — it's a formula you apply to your own traffic. Below is the exact breakdown of what goes into that formula, where the hidden costs hide, and how to estimate it before you ship a feature.

The core formula

Every request costs:

request_cost = (input_tokens / 1,000,000 × input_price_per_million)
              + (output_tokens / 1,000,000 × output_price_per_million)

Input and output are priced differently — output tokens are consistently more expensive than input tokens across every Claude model, usually by a factor of 4–5x. This matters more than most people assume: a chatbot that reads a 2,000-token document and replies with 50 tokens costs almost nothing per call. A code generator that reads 500 tokens and writes 3,000 tokens of code can cost 10x more per call even though the input is far smaller.

What counts as "input tokens"

Input tokens aren't just the user's message. They include:

This is the part that catches teams off guard. If your system prompt is 800 tokens and you have 10 tool definitions adding another 600 tokens, every single request pays for 1,400 tokens before the user has typed a word.

What counts as "output tokens"

Output tokens are everything the model generates in its response — including the content inside tool calls (the function name and arguments Claude generates when using tools), and any reasoning or intermediate text depending on the model and settings you use.

A worked example

Say a model is priced at $3 per million input tokens and $15 per million output tokens (illustrative numbers — always check current pricing for the exact model you're using).

Simple Q&A request:

Document summarization:

Long-form generation with tool use:

The spread here — $0.003 to $0.035 — is a 12x difference between the cheapest and most expensive call, all within one app that mixes task types. This is why a single "cost per request" number is almost always misleading unless you break it down by endpoint or feature.

Costs most people forget to count

Conversation history growth. In a multi-turn chat, every turn re-sends the entire history as input. Turn 10 of a conversation costs far more in input tokens than turn 1, even if the user's message length stays constant.

Retries on errors or rate limits. A failed request that you retry still counts as two billed attempts if the first one partially completed or if you resend the same payload. Build retry logic that doesn't duplicate full-context calls unnecessarily.

Streaming vs non-streaming. Streaming doesn't change the token cost, but it does change how early you can cancel an unwanted generation — cutting a stream short when you detect a bad response saves output tokens that a blocking call would have paid for in full.

System prompts and tool schemas on every call. These are static costs you pay repeatedly. Trimming an 800-token system prompt to 300 tokens, multiplied across hundreds of thousands of requests, is often the single biggest lever for reducing average cost per request.

How to actually calculate your per-request cost

  1. Log token usage per request. Every response includes usage metadata with exact input and output token counts — don't estimate with a character count, use the real numbers.
  2. Segment by feature, not globally. A chat feature and a summarization feature have wildly different cost profiles. Average them together and you'll misforecast both.
  3. Multiply by expected volume. Cost per request only matters once you multiply it by requests per day/month. A $0.03 request run 50,000 times a day is $1,500/day — know this before launch, not after the invoice.
  4. Re-check after every prompt change. Adding a few examples to a system prompt, or adding a new tool definition, shifts your baseline input cost on every single call going forward.

If you're building on top of Claude and want this usage tracked automatically instead of parsing it out of raw API responses, SubToAPI gives you per-key usage metadata in the dashboard so you can see cost per request broken down by application key without building your own logging pipeline. It sits on top of your existing Claude access and exposes it as a standard HTTPS API with streaming and tool use support — see the quickstart for setup.

Reducing cost per request without changing the model

Questions

Is Claude API priced per request or per token? Per token, not per request. There's no flat fee per call — you're billed separately for input tokens (what you send) and output tokens (what the model generates), at different rates.

Why does my cost per request vary so much between calls? Because input and output lengths vary by task. Long documents increase input cost; long generated responses increase output cost — and output tokens are priced several times higher than input tokens, so generation-heavy tasks cost disproportionately more.

How can I see the exact cost of each request I make? Check the usage metadata returned with every API response, which includes exact input and output token counts. Multiply by the current per-million pricing for your model, or use a dashboard like SubToAPI's that surfaces this automatically per API key.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →