← Blog

Claude API max_tokens Parameter Explained

2026-10-10 · 5 min read · SubToAPI Team

What max_tokens actually controls

max_tokens is a required parameter in the Claude API's messages endpoint that sets a hard ceiling on how many tokens the model is allowed to generate in its response. It does not control the length of your input, it does not set a target length, and it is not a minimum. It's purely an upper bound on the output.

If Claude finishes its answer naturally before hitting that limit, the response stops on its own and you get a normal completion. If Claude is still generating when it reaches the limit, the response is cut off mid-stream — potentially mid-sentence, mid-code-block, or mid-JSON-object. Understanding this distinction is the whole point of this article, because getting max_tokens wrong is one of the most common sources of broken output in production Claude integrations.

Why max_tokens exists

Every Claude model has a maximum context window (input + output combined) and a separate maximum output length per model. max_tokens lets you cap generation within those limits for two practical reasons:

  1. Cost control. Output tokens are billed, and in most pricing tiers they cost more per token than input tokens. A runaway generation with no cap could produce an unexpectedly large bill.
  2. Latency control. Generation time scales roughly linearly with output length. Capping tokens caps worst-case response time, which matters for chat UIs and anything with a timeout.

Because it's required on every request, you can't accidentally omit it and let the model generate indefinitely — Anthropic forces you to think about it.

Example request

curl https://api.anthropic.com/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Summarize this changelog in 3 bullet points."}
    ]
  }'

Here, max_tokens: 1024 means Claude can generate up to 1024 tokens of output. For a 3-bullet summary, it will almost certainly finish well under that cap and stop naturally — the limit is just a safety ceiling in this case.

How to tell if a response was truncated

Every response includes a stop_reason field. This is the field you should actually be checking, not just the content length:

{
  "id": "msg_01...",
  "content": [{"type": "text", "text": "Here is the summary..."}],
  "stop_reason": "max_tokens",
  "usage": {"input_tokens": 142, "output_tokens": 1024}
}

If you see stop_reason: "max_tokens", treat that response as incomplete and either raise the limit and retry, or handle the truncation explicitly (e.g., ask Claude to continue). Never assume a response is complete just because it returned a 200 status.

Choosing a sensible value

There's no universal "right" number — it depends on the task:

A good rule of thumb: set max_tokens generously above what you expect to need, since it doesn't cost anything extra if the model stops early. You're only billed for tokens actually generated, not for the cap itself. The cap only matters when the model actually reaches it.

max_tokens vs stop_sequences

These two are often confused. max_tokens is a blunt numeric ceiling. stop_sequences lets you define specific strings that, when generated, cause Claude to stop immediately — useful for structured formats like custom delimiters or role-play transcripts. You can use both together: stop_sequences for clean, intentional stopping points, and max_tokens as the backstop that guarantees the request never runs away.

max_tokens when streaming or building an API layer

If you're streaming responses — directly from Anthropic or through a layer like SubToAPI — max_tokens behaves the same way: it caps the total tokens across the whole stream, and the final message_stop event will carry the same stop_reason semantics. If you're building internal tooling on top of Claude access, this is also where usage metadata becomes useful for cost tracking and debugging truncated responses across a team. SubToAPI's messages endpoint and streaming guide expose the same stop_reason and token usage fields so you can monitor truncation issues without writing your own logging layer. See the quickstart for a working example, or check pricing if you're evaluating it for a team.

Practical checklist

Questions

Does a higher max_tokens value cost more if Claude stops early? No. You're billed only for tokens actually generated, reported in the usage.output_tokens field. Setting a high cap costs nothing extra unless the model actually generates that many tokens.

What happens if my response gets cut off by max_tokens? The stop_reason field will read "max_tokens" and the content will be incomplete, possibly mid-sentence or mid-JSON. You should detect this and either retry with a higher limit or send a follow-up request asking Claude to continue.

Is there a maximum value I can set for max_tokens? Yes, each Claude model has a maximum output token limit that max_tokens cannot exceed, separate from the total context window. Check the current model documentation for the exact ceiling, since it varies by model version.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →