Claude API Token Counting Before Request: Full Guide
If you're building anything that sends variable-length content to Claude — long documents, chat history, retrieved context — you eventually hit the same problem: you don't know how many tokens a request will use until you've already sent it. By then it's too late to trim the payload, and you've either hit a context-length error or paid for tokens you didn't need to send.
Counting tokens before a request means estimating or exactly computing the token count of your prompt (system message, messages, tool definitions) before calling the Messages endpoint, so you can enforce limits, truncate history, or choose a cheaper model when the input is small. Anthropic exposes a dedicated token-counting endpoint for exactly this purpose, and there are reliable approximation methods you can use client-side when you don't want an extra network round trip.
Why count tokens before sending a request
A few concrete reasons this matters in production:
- Context window limits. Claude models have a fixed context window (input + output tokens combined). If your conversation history plus a new user message exceeds that, the request fails. Catching this before the call lets you summarize or drop old messages gracefully instead of surfacing an API error to the user.
- Cost control. Token counting lets you estimate cost per request before you spend it, which is essential if you're exposing a chat feature with a per-user budget or rate limit.
- Model routing. Some teams route short requests to a faster/cheaper model and only use a larger model when the input is large or complex. You need a token count to make that decision.
- Streaming UX. If you want to show "this request will use ~X tokens" in your UI before the user hits send, you need a pre-flight count.
Using Anthropic's token counting endpoint
Anthropic provides a count_tokens endpoint that accepts the same messages, system, and tools structure as the Messages API and returns an exact input token count without generating a completion. This is the most accurate method because it uses the same tokenizer the model will actually use.
curl https://api.anthropic.com/v1/messages/count_tokens \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-3-5-sonnet-20241022",
"messages": [
{"role": "user", "content": "Summarize this contract in three bullet points."}
]
}'
The response returns an input_tokens field. Note that this count only covers input — the final cost and context usage also depends on the output tokens Claude generates, which you can't know until after the response completes.
Counting tokens with tools and system prompts included
If your requests include tool definitions or a system prompt, include them in the count request exactly as you would in the real call — tool schemas can add a non-trivial number of tokens, especially with several tools and detailed parameter descriptions.
const response = await fetch("https://api.anthropic.com/v1/messages/count_tokens", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-3-5-sonnet-20241022",
system: "You are a support assistant for an e-commerce platform.",
messages: conversationHistory,
tools: toolDefinitions,
}),
});
const { input_tokens } = await response.json();
if (input_tokens > MAX_CONTEXT_BUDGET) {
conversationHistory = truncateHistory(conversationHistory);
}
Approximate counting without an API call
Calling count_tokens is accurate but adds latency and an extra request. For quick client-side estimates — before deciding whether to even attempt a call — a rough rule of thumb for English text is:
- ~4 characters per token
- ~0.75 tokens per word
This is good enough for UI feedback ("approximately 1,200 tokens") but not reliable enough for hard limit enforcement, since actual tokenization varies with punctuation, code, non-English text, and JSON structure. For anything where exceeding the limit causes a failed request, use the real count_tokens endpoint or a proper tokenizer library instead of character-based estimates.
Building a pre-flight check into your request flow
A practical pattern looks like this:
- Assemble the full request payload (system, messages, tools) as you normally would.
- Call
count_tokenswith that exact payload. - If
input_tokensplus your expected max output tokens exceeds the model's context window, trim the oldest messages or summarize them. - Send the actual Messages request.
async function sendWithBudgetCheck(payload, maxContext) {
const countRes = await fetch("https://api.anthropic.com/v1/messages/count_tokens", {
method: "POST",
headers: { /* auth headers */ },
body: JSON.stringify(payload),
});
const { input_tokens } = await countRes.json();
if (input_tokens + payload.max_tokens > maxContext) {
payload.messages = trimOldestMessages(payload.messages);
}
return fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: { /* auth headers */ },
body: JSON.stringify(payload),
});
}
This same logic applies regardless of which client you build against. If you're routing Claude access through SubToAPI, your application still talks to a standard Messages-style endpoint, so the same pre-flight pattern works — you build the request payload, check it against your own budget logic, then send it to your sub_live_... key-authenticated endpoint. See /docs/messages for the exact request shape and /docs/quickstart if you're setting this up for the first time.
Token counting and usage metadata after the fact
Pre-flight counting handles input tokens, but you'll also want to track actual usage (input + output) per request for billing and rate-limit purposes. The Messages API response includes a usage object with input_tokens and output_tokens for every completed request, which is what you should log for accurate cost tracking rather than relying on pre-flight estimates alone. SubToAPI surfaces this same usage metadata per API key in the dashboard, so teams can see per-key and per-seat consumption without building their own logging pipeline — useful if multiple people on a /pricing Team or Scale plan are sharing the same underlying Claude access through separate sub_live_ keys.
questions
Does counting tokens before a request cost anything? The count_tokens endpoint itself doesn't generate a completion, so it's typically free or negligible compared to an actual Messages call, but it still counts as an API request against your rate limits.
Is character-based token estimation accurate enough for production limits? No. Character or word-based heuristics (roughly 4 characters per token) are fine for rough UI estimates but can be off by 10–20%, which is too imprecise for hard context-window enforcement — use the real tokenizer or count_tokens endpoint for that.
Does token counting include the output Claude will generate? No. Pre-flight counting only measures input tokens (system, messages, tools). Output tokens are only known after the response completes and are reported in the response's usage field.