Claude API max_tokens Parameter Explained
What max_tokens actually controls
max_tokens is a required parameter in the Claude API's messages endpoint that sets a hard ceiling on how many tokens the model is allowed to generate in its response. It does not control the length of your input, it does not set a target length, and it is not a minimum. It's purely an upper bound on the output.
If Claude finishes its answer naturally before hitting that limit, the response stops on its own and you get a normal completion. If Claude is still generating when it reaches the limit, the response is cut off mid-stream — potentially mid-sentence, mid-code-block, or mid-JSON-object. Understanding this distinction is the whole point of this article, because getting max_tokens wrong is one of the most common sources of broken output in production Claude integrations.
Why max_tokens exists
Every Claude model has a maximum context window (input + output combined) and a separate maximum output length per model. max_tokens lets you cap generation within those limits for two practical reasons:
- Cost control. Output tokens are billed, and in most pricing tiers they cost more per token than input tokens. A runaway generation with no cap could produce an unexpectedly large bill.
- Latency control. Generation time scales roughly linearly with output length. Capping tokens caps worst-case response time, which matters for chat UIs and anything with a timeout.
Because it's required on every request, you can't accidentally omit it and let the model generate indefinitely — Anthropic forces you to think about it.
Example request
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Summarize this changelog in 3 bullet points."}
]
}'
Here, max_tokens: 1024 means Claude can generate up to 1024 tokens of output. For a 3-bullet summary, it will almost certainly finish well under that cap and stop naturally — the limit is just a safety ceiling in this case.
How to tell if a response was truncated
Every response includes a stop_reason field. This is the field you should actually be checking, not just the content length:
"end_turn"— Claude finished naturally. Everything is complete."max_tokens"— Claude hit the cap and was cut off. The content is incomplete."stop_sequence"— Claude hit a custom stop sequence you defined."tool_use"— Claude stopped to call a tool.
{
"id": "msg_01...",
"content": [{"type": "text", "text": "Here is the summary..."}],
"stop_reason": "max_tokens",
"usage": {"input_tokens": 142, "output_tokens": 1024}
}
If you see stop_reason: "max_tokens", treat that response as incomplete and either raise the limit and retry, or handle the truncation explicitly (e.g., ask Claude to continue). Never assume a response is complete just because it returned a 200 status.
Choosing a sensible value
There's no universal "right" number — it depends on the task:
- Short classification or extraction tasks (sentiment labels, yes/no answers, single JSON fields): 50–200 tokens is usually plenty.
- Chat responses and summaries: 500–1500 tokens covers most conversational replies.
- Long-form content, code generation, or structured documents: 2000–8000+ tokens, depending on the model's max output limit.
- Tool use / JSON output: give enough headroom for the full schema plus any explanatory text, or you'll get truncated JSON that fails to parse — a very common bug.
A good rule of thumb: set max_tokens generously above what you expect to need, since it doesn't cost anything extra if the model stops early. You're only billed for tokens actually generated, not for the cap itself. The cap only matters when the model actually reaches it.
max_tokens vs stop_sequences
These two are often confused. max_tokens is a blunt numeric ceiling. stop_sequences lets you define specific strings that, when generated, cause Claude to stop immediately — useful for structured formats like custom delimiters or role-play transcripts. You can use both together: stop_sequences for clean, intentional stopping points, and max_tokens as the backstop that guarantees the request never runs away.
max_tokens when streaming or building an API layer
If you're streaming responses — directly from Anthropic or through a layer like SubToAPI — max_tokens behaves the same way: it caps the total tokens across the whole stream, and the final message_stop event will carry the same stop_reason semantics. If you're building internal tooling on top of Claude access, this is also where usage metadata becomes useful for cost tracking and debugging truncated responses across a team. SubToAPI's messages endpoint and streaming guide expose the same stop_reason and token usage fields so you can monitor truncation issues without writing your own logging layer. See the quickstart for a working example, or check pricing if you're evaluating it for a team.
Practical checklist
- Always check
stop_reason, not just whether the request succeeded. - Set
max_tokensbased on task type, not a single global default across your app. - For JSON/tool-use responses, leave extra headroom to avoid truncated, unparseable output.
- Combine with
stop_sequencesfor predictable, clean stopping points. - Remember it's a ceiling, not a target — Claude won't pad output to reach it.
Questions
Does a higher max_tokens value cost more if Claude stops early? No. You're billed only for tokens actually generated, reported in the usage.output_tokens field. Setting a high cap costs nothing extra unless the model actually generates that many tokens.
What happens if my response gets cut off by max_tokens? The stop_reason field will read "max_tokens" and the content will be incomplete, possibly mid-sentence or mid-JSON. You should detect this and either retry with a higher limit or send a follow-up request asking Claude to continue.
Is there a maximum value I can set for max_tokens? Yes, each Claude model has a maximum output token limit that max_tokens cannot exceed, separate from the total context window. Check the current model documentation for the exact ceiling, since it varies by model version.