Claude API Request Timeout Configuration Guide
Claude API requests can take anywhere from a few hundred milliseconds to over a minute, depending on model size, output length, and whether you're streaming. If your HTTP client's default timeout is too aggressive, you'll see connection aborts on perfectly healthy requests — especially for long completions, large tool-use loops, or extended thinking. The fix is almost always to configure the timeout explicitly at the client level, and in most cases, switch to streaming so you get partial output before any timeout has a chance to fire.
This guide covers how to set request timeouts correctly across the common HTTP clients and SDKs used with Claude, what timeout values make sense for different workloads, and how to distinguish a real timeout from a stalled connection or rate limit.
Why Claude requests time out
There are three common causes:
- Client-side default timeout is too short. Many HTTP libraries default to 10–30 seconds. A Claude response generating 4,000 output tokens on a larger model can easily exceed that.
- Long output without streaming. Non-streaming requests wait for the full response body before returning anything. If generation takes 90 seconds, your client waits 90 seconds with zero visibility into progress.
- Network-level idle timeouts. Load balancers, reverse proxies, or corporate firewalls sometimes kill idle connections after a fixed window, independent of your application code.
The practical fix depends on which of these applies, but the baseline recommendation is: set an explicit timeout on every Claude request, size it to your expected output length, and prefer streaming for anything that might run long.
Setting timeouts with curl
By default, curl has no timeout — it will wait indefinitely, which is usually not what you want either. Use --max-time for a hard ceiling on the whole request, and --connect-timeout for the initial connection:
curl https://api.anthropic.com/v1/messages \
--max-time 120 \
--connect-timeout 10 \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-opus-4-6-20251001",
"max_tokens": 4096,
"messages": [{"role": "user", "content": "Summarize this document."}]
}'
--max-time 120 gives the full request-response cycle two minutes. For large max_tokens values (8k+), bump this to 180–300 seconds unless you're streaming.
Setting timeouts in JavaScript (fetch / axios)
Native fetch doesn't support a timeout option directly — you need AbortController:
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 90_000);
try {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-opus-4-6-20251001",
max_tokens: 4096,
messages: [{ role: "user", content: "Summarize this document." }],
}),
signal: controller.signal,
});
clearTimeout(timeoutId);
} catch (err) {
if (err.name === "AbortError") {
console.error("Request timed out");
}
throw err;
}
With axios, pass timeout in milliseconds directly in the config object:
const res = await axios.post(
"https://api.anthropic.com/v1/messages",
{ model: "claude-opus-4-6-20251001", max_tokens: 4096, messages },
{
headers: { "x-api-key": process.env.ANTHROPIC_API_KEY },
timeout: 90_000,
}
);
Note that axios's timeout applies to the response — for streaming responses it behaves differently, so test it against your actual streaming setup rather than assuming defaults.
Timeouts in the official SDKs
Both the Python and TypeScript Anthropic SDKs expose a timeout parameter at the client level and per-request:
const client = new Anthropic({
apiKey: process.env.ANTHROPIC_API_KEY,
timeout: 120_000, // 2 minutes, client default
});
// Override per request
const message = await client.messages.create(
{ model: "claude-opus-4-6-20251001", max_tokens: 4096, messages },
{ timeout: 300_000 }
);
The SDKs also implement automatic retries with backoff on transient failures, but a hard timeout is still treated as a failure that surfaces to your code — retries don't extend the timeout window, they just re-attempt the whole request.
Recommended timeout values by use case
| Workload | Suggested timeout | |---|---| | Short completions (chat replies, classification) | 30–60s | | Long-form generation (reports, code files) | 120–180s | | Streaming responses | 300s+ (measure idle time, not total time) | | Tool-use loops with multiple round trips | Set per-call timeout, not one timeout for the whole loop | | Extended thinking / complex reasoning | 180–300s |
For anything in the 60-second-plus range, streaming is the better fix rather than just raising the timeout — see the streaming guide for chunk-handling details.
Streaming avoids most timeout problems
The most reliable way to eliminate timeout issues on long generations is to stream the response. With stream: true, you receive tokens as they're generated, so your connection stays active and you can set a much shorter idle timeout (time since the last chunk) instead of a total-duration timeout:
const stream = await client.messages.stream({
model: "claude-opus-4-6-20251001",
max_tokens: 4096,
messages,
});
for await (const event of stream) {
// process each chunk as it arrives
}
If you're building a production API layer on top of Claude — handling retries, timeouts, and streaming reconnects for a team rather than a single script — SubToAPI wraps this into a hosted endpoint with its own sane defaults, so you don't have to tune timeout logic per client. Requests go through sub_live_... API keys with usage metadata attached, and the streaming docs cover chunk handling in detail if you're migrating existing timeout-sensitive code.
Handling timeout errors gracefully
When a timeout fires, don't assume the request failed on Claude's side — the model may have completed generation while your client gave up waiting. Build your error handling around three responses:
- Retry with backoff for genuine network failures, using exponential delay (1s, 2s, 4s...).
- Log the timeout duration and prompt length so you can distinguish "timeout too short" from "actually slow model."
- Never retry blindly on a POST that has side effects (e.g., triggering a tool call with a payment action) without idempotency handling — a timed-out request may still have executed server-side.
If you're testing timeout configuration against SubToAPI's endpoint instead of calling Anthropic directly, the request/response shape is identical — see the messages guide for the full parameter reference, and the quickstart to get a sub_live_... key set up in under five minutes.
FAQs
What's a safe default timeout for Claude API requests? For non-streaming requests, 60–120 seconds covers most short-to-medium completions. If you regularly generate long outputs (4k+ tokens), either raise it to 180–300 seconds or switch to streaming and use an idle timeout instead.
Does the Claude API have its own server-side timeout? Yes, the API will eventually close very long-running requests server-side, but the practical limit you'll hit first is almost always your own client's timeout setting, not Anthropic's infrastructure.
Should I set one timeout for an entire tool-use loop? No. Set a timeout per individual API call within the loop, not one timeout spanning multiple round trips — a multi-step tool-use conversation can legitimately take several minutes across calls even if each call is fast.