Claude API Pricing Per Token Breakdown Explained
Claude API pricing is calculated per token, not per request, and the rate you pay depends on three things: which model you call (Opus, Sonnet, or Haiku), whether the tokens are input or output, and whether you're using features like prompt caching. Input tokens — the text you send in — are always cheaper than output tokens, usually by a factor of 3 to 5x, because generating text costs more compute than reading it.
If you're trying to estimate what a Claude-powered feature will actually cost you each month, you need to understand four variables: model tier, input token volume, output token volume, and caching behavior. This article breaks down each one so you can do the math yourself instead of guessing from a monthly invoice.
How Token-Based Pricing Actually Works
Anthropic prices Claude models in dollars per million tokens (sometimes written as "per MTok"). A token is roughly 3/4 of an English word — so 1,000 tokens is about 750 words. Every API call you make consumes:
- Input tokens: your system prompt + conversation history + user message + any tool definitions you pass
- Output tokens: everything Claude generates in response, including any tool-use JSON it emits
Both are counted and billed separately, and the counts are returned in the usage field of every API response, so you don't have to estimate — you get exact numbers per call.
{
"usage": {
"input_tokens": 512,
"output_tokens": 184
}
}
Multiply those counts by the per-token rate for the model you used, and you have the exact cost of that single call.
The Model Tiers Price Differently
Claude is offered in multiple tiers, and the pricing gap between them is large — often 10x or more between the cheapest and most expensive model:
- Haiku — the fastest, cheapest tier. Best for classification, extraction, routing, and high-volume simple tasks.
- Sonnet — the balanced tier. Good default for most product features: chat, summarization, coding assistance.
- Opus — the most capable and most expensive tier. Reserved for complex reasoning, long multi-step tasks, or cases where output quality directly drives revenue.
The exact current rates per million tokens change over time as Anthropic updates its model lineup, so always check the official Anthropic pricing page for the live numbers before budgeting. What matters architecturally is the ratio between tiers: a well-designed app routes simple requests to Haiku and reserves Opus for the subset of requests that genuinely need it, which can cut blended cost per request dramatically without hurting output quality where it counts.
Input vs Output: Why the Split Matters
Because output tokens cost more than input tokens, two apps with identical total token counts can have very different bills depending on the ratio of input to output.
- A summarization tool (large input, short output) will skew cheap — you're mostly paying input rates.
- A content generation tool (short input, long output) will skew expensive — you're mostly paying the higher output rate.
- A chat agent with long conversation history re-sent on every turn pays input costs repeatedly for the same context, which is where costs quietly balloon.
This last point is the one teams miss most often: if you're not trimming or caching conversation history, you're re-billing the same input tokens on every single turn of a long conversation.
Prompt Caching Lowers Repeated-Context Costs
If your application sends the same large system prompt, document, or tool definitions on every call — which is extremely common in RAG pipelines and coding assistants — prompt caching lets you mark that content as cacheable. Cached input tokens are billed at a significantly reduced rate on subsequent calls within the cache window, while only the new, changed portion of the prompt is billed at full price.
This is the single highest-leverage optimization for input-heavy workloads: a 10,000-token system prompt reused across thousands of requests per day is exactly the scenario caching was built for.
A Worked Example (Illustrative)
To see how the pieces combine, imagine a support chatbot built on a mid-tier model where:
- System prompt + context: 1,200 input tokens per call
- User message: 80 input tokens
- Response: 220 output tokens
- 50,000 calls per month
That's 64,000,000 input tokens and 11,000,000 output tokens per month. Since output tokens are billed at several times the input rate, the output side can end up contributing a disproportionate share of the total bill even though it's a much smaller token count — which is exactly why tracking input/output separately (not just "total tokens") matters when you're forecasting spend.
Managing Per-Token Costs in Practice
A few practical levers, regardless of which model tier you use:
- Trim conversation history — summarize or drop old turns instead of resending full context every call.
- Cache stable context — system prompts, long documents, and tool schemas that don't change call-to-call.
- Route by task complexity — Haiku for routing/classification, Sonnet for general work, Opus only where needed.
- Cap output length — set
max_tokensdeliberately instead of leaving generous defaults. - Watch tool-use overhead — tool definitions are sent as input tokens on every call unless cached, and tool-call JSON counts as output tokens.
If you're building a product on top of Claude and want token usage visible per user or per team rather than buried in a single Anthropic invoice, SubToAPI sits on top of your existing Claude access and gives you a standard HTTPS API with per-key usage metadata, so you can see exactly which application, user, or feature is driving token spend. Plans start at €9/month on the Solo tier, with Team and Scale tiers for multi-seat setups — see /pricing for details, or get started at /signup.
questions
Are input and output tokens priced the same? No. Output tokens cost more than input tokens, typically 3–5x, because generation requires more compute than reading. Check your usage object per call to see the exact split.
Why does the same conversation get more expensive over time? Because most Claude integrations resend full conversation history on every turn, input token cost compounds as a chat gets longer. Trimming history or using prompt caching reduces this.
Does tool use add extra token cost? Yes. Tool definitions are counted as input tokens on each call unless cached, and any tool-call output Claude generates counts as output tokens — see /docs/tools for how this is structured in requests and responses.