Claude API Pricing Tiers Explained Simply
Claude API pricing is not a single number — it's a combination of which model you call, how many tokens you send and receive, and whether you use features like prompt caching or batch processing. If you've looked at Anthropic's pricing page and felt confused, you're not alone: it's structured around usage, not seats, which is a different mental model than most SaaS pricing.
This article breaks down exactly how Claude API pricing works, what the model tiers mean for your wallet, and how to think about costs when you're building a real product instead of just testing prompts in a console.
The Core Idea: You Pay Per Token, Not Per Request
Unlike a flat "per API call" fee, Claude charges based on tokens — small chunks of text (roughly ¾ of a word in English). Every request has two token costs:
- Input tokens: the text you send (system prompt, conversation history, documents, tool definitions)
- Output tokens: the text Claude generates back
Output tokens almost always cost more per unit than input tokens, because generating text is more compute-intensive than reading it. This matters in practice: a chatbot that reads a long document and gives a short answer is cheap. An agent that reads a short prompt but writes a long report is comparatively expensive.
The Three Model Tiers
Anthropic offers Claude in three broad capability tiers, and the pricing tiers map directly to them:
Haiku — fast and cheap
The smallest, fastest model. Priced the lowest of the three tiers. Good for high-volume, low-complexity tasks: classification, simple extraction, short replies, autocomplete-style features. If you're processing thousands of requests per hour and the task doesn't need deep reasoning, this is where you start.
Sonnet — the balanced default
Mid-tier pricing, mid-tier latency, strong general reasoning. This is what most production apps end up using for chat features, summarization, coding assistants, and RAG-style question answering. It's the tier most people mean when they say "Claude" without specifying a model.
Opus — top capability
The most expensive tier, reserved for tasks that genuinely need the strongest reasoning: complex multi-step agents, hard coding problems, long-form analysis where accuracy matters more than speed or cost.
The pricing gap between tiers is intentional — it lets you route easy tasks to Haiku and only pay Opus rates when the task actually requires it. Many production systems mix models: Haiku for triage, Sonnet for the main workload, Opus only when a task is flagged as complex.
Features That Change Your Effective Price
A few Claude API features can meaningfully lower your cost per request beyond the base per-token rate:
- Prompt caching: if you repeatedly send the same large system prompt or reference document, caching lets you avoid paying full input price on every call.
- Batch processing: for non-real-time workloads (bulk classification, overnight jobs), batch requests are typically discounted compared to synchronous calls.
- Context window usage: sending unnecessarily long conversation history inflates input token costs on every single turn — trimming or summarizing history is a real cost lever, not just a performance one.
None of these change which "tier" you're on, but they change your actual bill, sometimes significantly.
Where Token Pricing Gets Complicated for Teams
Token-based pricing makes sense for a single developer running experiments. It gets harder once you have a team: multiple people need access, you want per-project usage visibility, and finance wants one predictable invoice instead of a variable Anthropic bill tied to raw API keys shared over Slack.
This is the gap SubToAPI is built for. Instead of managing raw provider keys and reconciling token usage across a team, you get application-level API keys (sub_live_...), streaming and tool use support, usage metadata per key, and team seats — all behind flat monthly pricing: Solo €9, Team €19/seat, and Scale €49/seat, with a free trial at signup.
The API surface stays familiar — same message structure, streaming, and tool calling patterns you'd expect — see /docs/quickstart, /docs/messages, /docs/streaming, and /docs/tools for the specifics. The difference is you're issuing scoped keys per app or environment and billing is predictable per seat rather than fluctuating with token volume across a shared account.
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet",
"max_tokens": 512,
"messages": [
{"role": "user", "content": "Summarize this pricing model in one sentence."}
]
}'
If you're comparing raw token pricing against a seat-based plan, check /pricing against your expected monthly token spend — for a small team making frequent Sonnet-tier calls, a flat per-seat price is often easier to forecast than variable usage billing.
How to Actually Choose a Tier
A practical approach:
- Start on Sonnet for your core feature. It's the safest default for reasoning quality vs. cost.
- Downgrade to Haiku for any sub-task that's simple, repetitive, or high-volume (tagging, routing, short replies).
- Reserve Opus for a narrow set of high-stakes tasks — complex agent steps, code review, anything where a wrong answer is expensive.
- Measure before optimizing. Log token usage per request type before assuming you need a cheaper model — often the fix is shorter prompts, not a different tier.
FAQ
Is Claude API pricing based on subscriptions or usage? It's usage-based — you pay per input and output token, per model tier. There's no flat subscription fee from Anthropic directly; costs scale with how much you send and generate.
Which Claude tier should I use for a chatbot? Sonnet is the standard starting point for most chat products — strong reasoning at moderate cost. Use Haiku for simple intents and Opus only for genuinely hard queries.
Why do output tokens cost more than input tokens? Generating text requires more computation per token than reading it, so Anthropic prices output higher across all three model tiers to reflect that cost difference.