Claude API Temperature and Top P Settings Explained
When you call the Claude API, two parameters control how predictable or varied the output is: temperature and top_p. Temperature adjusts how much randomness is injected into token selection — low values make Claude pick the most likely next word almost every time, high values let it take more creative risks. Top_p (nucleus sampling) limits the pool of candidate tokens to the smallest set whose combined probability reaches a threshold, then samples from that reduced set.
For most use cases, you only need to touch one of these, not both. If you want deterministic, focused answers (code generation, data extraction, classification), set temperature low — around 0 to 0.3. If you want more varied, creative output (brainstorming, copywriting, dialogue), raise it to 0.7–1.0. top_p is usually left at its default (1.0) unless you have a specific reason to cap the sampling pool; adjusting both at once makes behavior harder to predict and debug.
How Temperature Actually Works
Temperature scales the probability distribution over possible next tokens before sampling.
- temperature = 0: Claude almost always picks the highest-probability token. Output becomes repeatable and deterministic (though not perfectly deterministic — there's still some floating-point and system-level variance).
- temperature = 1 (default in most SDKs): The model samples according to its natural probability distribution, striking a balance between coherence and variety.
- temperature > 1: Rarely useful in practice. The distribution flattens further, increasing the chance of selecting low-probability tokens, which often produces incoherent or off-topic text.
A simple mental model: temperature controls how much the model is willing to gamble on less likely words.
How Top_p Works
top_p restricts sampling to a cumulative probability mass. If top_p = 0.9, Claude only considers tokens that together make up the top 90% of probability mass, then samples from that reduced pool (optionally combined with temperature scaling).
- top_p = 1.0: No restriction — all tokens with nonzero probability stay eligible.
- top_p = 0.5: Only a narrow, high-confidence set of tokens is eligible, which tends to produce safer, more conventional phrasing.
- top_p < 0.3: Very restrictive. Useful for tightly constrained outputs like structured data or short classification labels, but can make prose feel stilted.
Top_p and temperature interact: a high temperature with a low top_p still constrains the candidate pool, just with more randomness within it. In practice, most teams pick one lever and leave the other at default rather than tuning both simultaneously.
Practical Settings by Task
| Task | temperature | top_p | |---|---|---| | Code generation / SQL | 0–0.2 | 1.0 (default) | | Data extraction / classification | 0–0.3 | 1.0 (default) | | Technical writing / docs | 0.3–0.5 | 1.0 (default) | | Customer support replies | 0.4–0.6 | 1.0 (default) | | Brainstorming / ideation | 0.8–1.0 | 1.0 (default) | | Creative writing / fiction | 0.9–1.0 | 0.9–1.0 |
These are starting points, not rules — test against your own prompts and evaluate output quality directly.
Setting Temperature and Top_p in Requests
Here's a basic example calling the Claude Messages API directly:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-4",
"max_tokens": 500,
"temperature": 0.2,
"messages": [
{"role": "user", "content": "Extract the invoice total from this text: ..."}
]
}'
If you're routing through SubToAPI instead of managing raw Anthropic credentials, the request shape is nearly identical — you just swap the endpoint and auth header:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-4",
"max_tokens": 500,
"temperature": 0.2,
"messages": [
{"role": "user", "content": "Extract the invoice total from this text: ..."}
]
}'
This matters for teams because sub_live_ keys let you set per-application defaults for temperature without touching the underlying Anthropic account, and usage metadata in the dashboard shows which temperature settings are being used across requests — useful for spotting when a team member accidentally left a creative-writing config on a production extraction endpoint. See the Messages API docs for the full parameter list, or the quickstart if you're setting up a new key.
Common Mistakes
Setting temperature to 0 and expecting perfect determinism. Temperature 0 drastically reduces variance but doesn't guarantee byte-identical outputs across calls, especially with longer generations or tool use involved.
Cranking temperature up to "fix" repetitive or generic answers. If output feels bland, the real issue is usually prompt specificity, not randomness. Try adding examples or more detail to your system prompt before touching temperature.
Adjusting top_p and temperature together without testing each in isolation. This makes it hard to know which change actually affected the output. Pick one, run a batch of test prompts, then adjust the other if needed.
Using high temperature in streaming applications without handling inconsistency. If you're streaming responses to users (see streaming docs), higher temperature means more variance in response length and structure, which can complicate UI logic expecting consistent formatting.
FAQ
Should I use temperature or top_p for more consistent output?
Use temperature. It's the more intuitive and widely tested lever for controlling randomness. Reserve top_p adjustments for cases where you specifically want to narrow the candidate token pool while keeping some sampling variety — most teams never need to touch it.
What temperature should I use for code generation with Claude?
Start at 0 to 0.2. Code benefits from determinism and correctness over creativity, and low temperature reduces the chance of syntactically odd or inconsistent output across repeated calls.
Does changing temperature affect token usage or cost?
No. Temperature and top_p affect which tokens are sampled, not how many are generated or billed. Cost is still based on input and output token counts regardless of these settings — see pricing for how SubToAPI plans handle usage.