Claude API Token Counting Library for Python
If you're building anything with Claude in Python — a chatbot, a RAG pipeline, a billing system — you'll eventually need to know how many tokens a prompt or response will consume. The short answer: Anthropic's official Python SDK includes a token counting method, and it's the most accurate option available. Third-party libraries like tiktoken will give you the wrong numbers because they use OpenAI's tokenizer, not Claude's.
This article covers what actually works for counting Claude tokens in Python, when to use the official method versus estimation, and why for most production apps you don't need a separate counting step at all — you can just read token usage off the API response.
Why Token Counting Matters for Claude
Token counts drive three things in any Claude integration:
- Cost estimation — you're billed per input and output token, so knowing the count before you send a request lets you predict spend.
- Context window management — Claude models have fixed context limits, and you need to know if a prompt (plus history, plus tool definitions) fits.
- Truncation and chunking — for RAG or long-document workflows, you need to split text into chunks that stay under a token budget.
None of these are solved reliably by guessing based on word or character count. Token boundaries depend on the model's specific tokenizer, and Claude's tokenizer is not the same as GPT's.
The Official Way: Anthropic's Token Counting Endpoint
Anthropic exposes a count_tokens capability through its Python SDK that sends your message payload to the API and returns the exact input token count Claude would use — including system prompts, tool definitions, and multi-turn history.
import anthropic
client = anthropic.Anthropic(api_key="sk-ant-...")
response = client.messages.count_tokens(
model="claude-sonnet-4-20250514",
messages=[
{"role": "user", "content": "Explain token counting in one paragraph."}
],
)
print(response.input_tokens)
This is the most accurate method because it's the actual tokenizer, not an approximation. The tradeoff is that it's a network call — you're paying a small latency cost to get an exact number before you send the real request.
Why tiktoken and Other Estimators Fall Short
tiktoken is OpenAI's tokenizer library, built for GPT models. It's fast and works offline, which makes it tempting to reach for when estimating Claude token counts too. The problem is that Claude uses a different vocabulary and byte-pair encoding scheme, so tiktoken counts will be off — sometimes by 10-20%, sometimes more depending on the text (code, non-English text, and unusual formatting tend to diverge the most).
If you need a rough, offline estimate and can tolerate some error margin, a common heuristic is:
def estimate_tokens(text: str) -> int:
# Rough approximation: ~4 characters per token for English text
return len(text) // 4
This is fine for a progress bar or a "this might be long" warning. It is not fine for billing calculations or hard context-limit enforcement — use the official counting endpoint for anything where accuracy matters.
A Simple Chunking Pattern Using Token Counts
For RAG pipelines, you often need to split documents into chunks that fit a token budget. Combine the official counter with a splitting loop:
import anthropic
client = anthropic.Anthropic(api_key="sk-ant-...")
def count_tokens(text: str, model: str) -> int:
result = client.messages.count_tokens(
model=model,
messages=[{"role": "user", "content": text}],
)
return result.input_tokens
def chunk_by_tokens(text: str, max_tokens: int, model: str) -> list[str]:
words = text.split()
chunks, current = [], []
for word in words:
current.append(word)
if count_tokens(" ".join(current), model) >= max_tokens:
current.pop()
chunks.append(" ".join(current))
current = [word]
if current:
chunks.append(" ".join(current))
return chunks
Calling the counting endpoint per word is slow — in practice you'd batch this (check every N words, or estimate with the character heuristic first and only verify near the boundary) to cut down on API calls.
Skipping the Counting Step Entirely
For a lot of applications, the real goal isn't counting tokens in advance — it's knowing how many tokens a request actually used, for cost tracking or usage dashboards. If that's your use case, you don't need a token counting library at all: every Claude API response already includes exact input_tokens and output_tokens in its usage metadata.
This is where SubToAPI is useful if you're already routing requests through it. Every response from /v1/messages includes usage metadata alongside the completion, so you get exact per-request token counts without a separate counting call:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 512,
"messages": [{"role": "user", "content": "Summarize this in 3 bullets."}]
}'
The response includes token usage per call, which you can log and aggregate for per-user or per-team billing without maintaining a separate counting library. Combined with per-application API keys (sub_live_...), this makes it straightforward to see exactly how many tokens each part of your product is consuming. See /docs/messages for the full response shape, or /docs/quickstart to get set up.
Choosing an Approach
- Need exact pre-send counts (billing estimates, hard context limits): use the official
count_tokensmethod from Anthropic's SDK. - Need a fast, rough estimate (UI hints, non-critical warnings): use a character-based heuristic like
len(text) // 4. - Need post-hoc usage tracking (dashboards, per-customer billing): read
usagefields off the actual API response instead of counting anything in advance — SubToAPI surfaces this automatically per request.
FAQs
Does tiktoken work for counting Claude tokens accurately? No. tiktoken is built for OpenAI's tokenizer and will produce inaccurate counts for Claude text, sometimes off by a meaningful margin. Use Anthropic's official count_tokens method for exact numbers.
How do I count tokens for a multi-turn conversation, not just one message? Pass the full messages array (all turns, plus any system prompt) to the count_tokens call — it returns the total input token count for the whole payload, not just the last message.
Can I count tokens without making a network call? Not exactly — there's no official offline tokenizer library for Claude. You can approximate with a character-per-token heuristic for non-critical use cases, but for accurate counts you need the API-based count_tokens method or the usage metadata returned with an actual request.