← Blog

Claude API Token Counting Library for Python

2026-09-29 · 5 min read · SubToAPI Team

If you're building anything with Claude in Python — a chatbot, a RAG pipeline, a billing system — you'll eventually need to know how many tokens a prompt or response will consume. The short answer: Anthropic's official Python SDK includes a token counting method, and it's the most accurate option available. Third-party libraries like tiktoken will give you the wrong numbers because they use OpenAI's tokenizer, not Claude's.

This article covers what actually works for counting Claude tokens in Python, when to use the official method versus estimation, and why for most production apps you don't need a separate counting step at all — you can just read token usage off the API response.

Why Token Counting Matters for Claude

Token counts drive three things in any Claude integration:

None of these are solved reliably by guessing based on word or character count. Token boundaries depend on the model's specific tokenizer, and Claude's tokenizer is not the same as GPT's.

The Official Way: Anthropic's Token Counting Endpoint

Anthropic exposes a count_tokens capability through its Python SDK that sends your message payload to the API and returns the exact input token count Claude would use — including system prompts, tool definitions, and multi-turn history.

import anthropic

client = anthropic.Anthropic(api_key="sk-ant-...")

response = client.messages.count_tokens(
    model="claude-sonnet-4-20250514",
    messages=[
        {"role": "user", "content": "Explain token counting in one paragraph."}
    ],
)

print(response.input_tokens)

This is the most accurate method because it's the actual tokenizer, not an approximation. The tradeoff is that it's a network call — you're paying a small latency cost to get an exact number before you send the real request.

Why tiktoken and Other Estimators Fall Short

tiktoken is OpenAI's tokenizer library, built for GPT models. It's fast and works offline, which makes it tempting to reach for when estimating Claude token counts too. The problem is that Claude uses a different vocabulary and byte-pair encoding scheme, so tiktoken counts will be off — sometimes by 10-20%, sometimes more depending on the text (code, non-English text, and unusual formatting tend to diverge the most).

If you need a rough, offline estimate and can tolerate some error margin, a common heuristic is:

def estimate_tokens(text: str) -> int:
    # Rough approximation: ~4 characters per token for English text
    return len(text) // 4

This is fine for a progress bar or a "this might be long" warning. It is not fine for billing calculations or hard context-limit enforcement — use the official counting endpoint for anything where accuracy matters.

A Simple Chunking Pattern Using Token Counts

For RAG pipelines, you often need to split documents into chunks that fit a token budget. Combine the official counter with a splitting loop:

import anthropic

client = anthropic.Anthropic(api_key="sk-ant-...")

def count_tokens(text: str, model: str) -> int:
    result = client.messages.count_tokens(
        model=model,
        messages=[{"role": "user", "content": text}],
    )
    return result.input_tokens

def chunk_by_tokens(text: str, max_tokens: int, model: str) -> list[str]:
    words = text.split()
    chunks, current = [], []
    for word in words:
        current.append(word)
        if count_tokens(" ".join(current), model) >= max_tokens:
            current.pop()
            chunks.append(" ".join(current))
            current = [word]
    if current:
        chunks.append(" ".join(current))
    return chunks

Calling the counting endpoint per word is slow — in practice you'd batch this (check every N words, or estimate with the character heuristic first and only verify near the boundary) to cut down on API calls.

Skipping the Counting Step Entirely

For a lot of applications, the real goal isn't counting tokens in advance — it's knowing how many tokens a request actually used, for cost tracking or usage dashboards. If that's your use case, you don't need a token counting library at all: every Claude API response already includes exact input_tokens and output_tokens in its usage metadata.

This is where SubToAPI is useful if you're already routing requests through it. Every response from /v1/messages includes usage metadata alongside the completion, so you get exact per-request token counts without a separate counting call:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-20250514",
    "max_tokens": 512,
    "messages": [{"role": "user", "content": "Summarize this in 3 bullets."}]
  }'

The response includes token usage per call, which you can log and aggregate for per-user or per-team billing without maintaining a separate counting library. Combined with per-application API keys (sub_live_...), this makes it straightforward to see exactly how many tokens each part of your product is consuming. See /docs/messages for the full response shape, or /docs/quickstart to get set up.

Choosing an Approach

FAQs

Does tiktoken work for counting Claude tokens accurately? No. tiktoken is built for OpenAI's tokenizer and will produce inaccurate counts for Claude text, sometimes off by a meaningful margin. Use Anthropic's official count_tokens method for exact numbers.

How do I count tokens for a multi-turn conversation, not just one message? Pass the full messages array (all turns, plus any system prompt) to the count_tokens call — it returns the total input token count for the whole payload, not just the last message.

Can I count tokens without making a network call? Not exactly — there's no official offline tokenizer library for Claude. You can approximate with a character-per-token heuristic for non-critical use cases, but for accurate counts you need the API-based count_tokens method or the usage metadata returned with an actual request.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →