← Blog

Claude API Pricing Per Token Breakdown Explained

2026-10-11 · 5 min read · SubToAPI Team

Claude API pricing is calculated per token, not per request, and the rate you pay depends on three things: which model you call (Opus, Sonnet, or Haiku), whether the tokens are input or output, and whether you're using features like prompt caching. Input tokens — the text you send in — are always cheaper than output tokens, usually by a factor of 3 to 5x, because generating text costs more compute than reading it.

If you're trying to estimate what a Claude-powered feature will actually cost you each month, you need to understand four variables: model tier, input token volume, output token volume, and caching behavior. This article breaks down each one so you can do the math yourself instead of guessing from a monthly invoice.

How Token-Based Pricing Actually Works

Anthropic prices Claude models in dollars per million tokens (sometimes written as "per MTok"). A token is roughly 3/4 of an English word — so 1,000 tokens is about 750 words. Every API call you make consumes:

Both are counted and billed separately, and the counts are returned in the usage field of every API response, so you don't have to estimate — you get exact numbers per call.

{
  "usage": {
    "input_tokens": 512,
    "output_tokens": 184
  }
}

Multiply those counts by the per-token rate for the model you used, and you have the exact cost of that single call.

The Model Tiers Price Differently

Claude is offered in multiple tiers, and the pricing gap between them is large — often 10x or more between the cheapest and most expensive model:

The exact current rates per million tokens change over time as Anthropic updates its model lineup, so always check the official Anthropic pricing page for the live numbers before budgeting. What matters architecturally is the ratio between tiers: a well-designed app routes simple requests to Haiku and reserves Opus for the subset of requests that genuinely need it, which can cut blended cost per request dramatically without hurting output quality where it counts.

Input vs Output: Why the Split Matters

Because output tokens cost more than input tokens, two apps with identical total token counts can have very different bills depending on the ratio of input to output.

This last point is the one teams miss most often: if you're not trimming or caching conversation history, you're re-billing the same input tokens on every single turn of a long conversation.

Prompt Caching Lowers Repeated-Context Costs

If your application sends the same large system prompt, document, or tool definitions on every call — which is extremely common in RAG pipelines and coding assistants — prompt caching lets you mark that content as cacheable. Cached input tokens are billed at a significantly reduced rate on subsequent calls within the cache window, while only the new, changed portion of the prompt is billed at full price.

This is the single highest-leverage optimization for input-heavy workloads: a 10,000-token system prompt reused across thousands of requests per day is exactly the scenario caching was built for.

A Worked Example (Illustrative)

To see how the pieces combine, imagine a support chatbot built on a mid-tier model where:

That's 64,000,000 input tokens and 11,000,000 output tokens per month. Since output tokens are billed at several times the input rate, the output side can end up contributing a disproportionate share of the total bill even though it's a much smaller token count — which is exactly why tracking input/output separately (not just "total tokens") matters when you're forecasting spend.

Managing Per-Token Costs in Practice

A few practical levers, regardless of which model tier you use:

  1. Trim conversation history — summarize or drop old turns instead of resending full context every call.
  2. Cache stable context — system prompts, long documents, and tool schemas that don't change call-to-call.
  3. Route by task complexity — Haiku for routing/classification, Sonnet for general work, Opus only where needed.
  4. Cap output length — set max_tokens deliberately instead of leaving generous defaults.
  5. Watch tool-use overhead — tool definitions are sent as input tokens on every call unless cached, and tool-call JSON counts as output tokens.

If you're building a product on top of Claude and want token usage visible per user or per team rather than buried in a single Anthropic invoice, SubToAPI sits on top of your existing Claude access and gives you a standard HTTPS API with per-key usage metadata, so you can see exactly which application, user, or feature is driving token spend. Plans start at €9/month on the Solo tier, with Team and Scale tiers for multi-seat setups — see /pricing for details, or get started at /signup.

questions

Are input and output tokens priced the same? No. Output tokens cost more than input tokens, typically 3–5x, because generation requires more compute than reading. Check your usage object per call to see the exact split.

Why does the same conversation get more expensive over time? Because most Claude integrations resend full conversation history on every turn, input token cost compounds as a chat gets longer. Trimming history or using prompt caching reduces this.

Does tool use add extra token cost? Yes. Tool definitions are counted as input tokens on each call unless cached, and any tool-call output Claude generates counts as output tokens — see /docs/tools for how this is structured in requests and responses.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →