Claude API Pricing Per Token Explained
Claude API pricing works on a per-token basis, meaning you pay separately for the text you send (input tokens) and the text Claude generates back (output tokens), with rates varying by model. A token is roughly 3.5-4 characters of English text, so a 1,000-word prompt is typically around 1,300-1,500 tokens. Output tokens almost always cost more than input tokens — often 4-5x more — because generation is the computationally expensive part.
This matters because the same prompt can cost wildly different amounts depending on which Claude model you pick, how much context you send with every request, and whether you're reusing long system prompts across calls. Understanding the per-token math lets you estimate costs before you build, and catch runaway spend before it shows up on an invoice.
How Per-Token Pricing Actually Works
Anthropic prices Claude models in dollars per million tokens, split into two rates:
- Input price: cost per million tokens sent to the model (your prompt, system instructions, conversation history, tool definitions)
- Output price: cost per million tokens the model generates back
Both rates scale with model tier. The most capable models (used for complex reasoning, coding, long documents) charge more per token than lighter, faster models built for simple classification or short replies. As a rule of thumb across the Claude lineup:
- Lightweight/fast models: cheapest per token, best for high-volume simple tasks
- Mid-tier models: balanced cost and capability, the default choice for most apps
- Top-tier models: most expensive per token, reserved for tasks that genuinely need the extra reasoning quality
Choosing the right tier for the job is the single biggest lever you have over your bill — using a flagship model for a task a lighter model could handle is the most common way teams overspend.
Calculating a Real Cost Estimate
The formula is simple:
cost = (input_tokens / 1,000,000 * input_price) + (output_tokens / 1,000,000 * output_price)
For example, if a model charges $3 per million input tokens and $15 per million output tokens, and a single request sends 2,000 input tokens and receives 500 output tokens back:
input cost = 2,000 / 1,000,000 * 3 = $0.006
output cost = 500 / 1,000,000 * 15 = $0.0075
total = $0.0135 per request
That looks negligible until you multiply by volume. At 50,000 requests a month, that's roughly $675. The math is why teams building chat features, summarization tools, or agents with tool calls need to model their expected token volume before shipping, not after the first invoice.
What Actually Drives Up Your Token Count
Per-token pricing means every part of the request body counts, not just the user's visible message:
- System prompts: a long, detailed system prompt is sent on every single call, so its cost multiplies by your request volume
- Conversation history: multi-turn chat apps resend prior messages each turn unless you manage context carefully, so cost grows with conversation length
- Tool definitions: if you're using tool use, the schema for every available tool is counted as input tokens on each request
- Retrieved context: RAG pipelines that stuff retrieved documents into the prompt can balloon input tokens fast
Output tokens are usually smaller in volume but priced higher, so verbose responses (long explanations instead of concise answers) cost more than you'd expect. Prompting the model to be concise, or setting a reasonable max_tokens cap, is a legitimate cost control, not just a style choice.
Tracking Spend in Practice
Raw token math is useful for estimation, but most teams struggle with visibility once an app is in production across multiple features or multiple developers. If you're integrating Claude directly, you need to build your own usage logging around every request to know which feature, user, or endpoint is driving cost.
This is one of the practical reasons teams put SubToAPI in front of their Claude usage: it turns your existing Claude access into a standard HTTPS API with per-key usage metadata, so you can see token consumption broken down by application key without building that tracking yourself. You generate scoped sub_live_... keys per app or per team member from the dashboard, and usage shows up alongside your existing plan — Solo at €9, Team at €19/seat, or Scale at €49/seat — with a free trial at signup to test it against your actual workload.
A basic request through SubToAPI looks like a normal Messages call:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4",
"max_tokens": 500,
"messages": [{"role": "user", "content": "Summarize this changelog."}]
}'
See /docs/quickstart to get a key running, and /docs/messages for the full request and response shape, including where usage data appears in the response.
Practical Ways to Reduce Per-Token Cost
- Match the model to the task. Don't route simple classification or extraction jobs through your most expensive model — see /docs for available models.
- Trim system prompts. Audit what's actually needed in every call; remove instructions the model doesn't need for that specific request.
- Cap output length. Set
max_tokensto a sane ceiling instead of leaving it unbounded. - Manage conversation history. Summarize or truncate older turns instead of resending full history indefinitely.
- Use streaming for long responses. Streaming doesn't change token cost, but it improves perceived latency and lets you cut off generation early if needed — see /docs/streaming.
- Batch and cache where possible. Avoid re-sending identical large contexts across repeated calls when the content hasn't changed.
Pricing plans and current per-token rates are listed at /pricing.
FAQ
Is Claude API pricing the same for input and output tokens? No. Output tokens cost more than input tokens on every Claude model, typically 4-5x the input rate, because generation requires more compute than processing the prompt.
Does a longer system prompt increase my bill even if the user's message is short? Yes. Every token in the system prompt, conversation history, and tool definitions counts as input tokens on each request, so a verbose system prompt sent thousands of times adds up quickly.
How can I estimate my monthly Claude API cost before launching? Estimate average input and output tokens per request, multiply by your expected monthly request volume and the model's per-token rates, then add a buffer for retries and longer-than-average conversations.