Claude API Streaming Response Timeout Fix
If your Claude API streaming requests are dying mid-response with a timeout error, the problem is almost never Claude itself — it's a timeout configured somewhere in the chain between your code and the model: your HTTP client's default timeout, a reverse proxy's idle connection limit, or a load balancer that doesn't understand long-lived Server-Sent Events connections. This article walks through where those timeouts come from and exactly how to configure around them.
The short version: streaming responses can legitimately take 30 seconds to several minutes depending on output length, and most default timeout settings (fetch, axios, nginx, API gateways) are tuned for short request/response cycles, not long-lived streams. Fixing the issue means raising the right timeout, resetting it on every received chunk instead of once per request, and handling disconnects gracefully.
Why streaming timeouts happen
A streaming call to Claude keeps a single HTTP connection open while tokens arrive incrementally as Server-Sent Events. Several layers can decide that connection has been open "too long" and kill it:
- Your HTTP client —
fetch,axios, or your language's HTTP library often has a default request timeout (sometimes as low as 10–30 seconds) that applies to the whole request, including the time spent streaming. - Reverse proxies — nginx, HAProxy, or similar have an idle timeout (
proxy_read_timeoutin nginx) that closes a connection if no data is sent for a window of time. - Load balancers and API gateways — many cloud load balancers (ALB, Cloud Load Balancing, API Gateway) have hard timeout ceilings, sometimes non-configurable below a certain plan tier.
- Corporate networks and VPNs — some networks silently drop idle TCP connections, which looks identical to a timeout from the client's perspective.
- Long time-to-first-token — complex prompts or large context windows can delay the first streamed chunk, which some clients misinterpret as a hung connection and abort early.
The fix depends on which layer is actually cutting the connection, so the first step is isolating where the timeout occurs.
Step 1: Separate the "no data at all" timeout from the "total duration" timeout
The most common mistake is setting a single fixed timeout on the entire streaming request. A 60-second total timeout will kill a perfectly healthy stream that's still actively sending tokens at second 61. What you actually want is an idle timeout that resets every time a chunk arrives, plus a generous (or no) cap on total duration.
function createIdleTimeout(ms, onTimeout) {
let timer = setTimeout(onTimeout, ms);
return {
reset() {
clearTimeout(timer);
timer = setTimeout(onTimeout, ms);
},
clear() {
clearTimeout(timer);
},
};
}
Use this instead of a single AbortController timeout fired once at request start.
Step 2: Raise HTTP client defaults for streaming calls
If you're using fetch with AbortController, don't apply a short default timeout to streaming requests:
const controller = new AbortController();
const idle = createIdleTimeout(30000, () => controller.abort());
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4",
stream: true,
max_tokens: 1024,
messages: [{ role: "user", content: "Write a 1500 word essay on Roman roads." }],
}),
signal: controller.signal,
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
idle.reset(); // keep the connection alive as long as data keeps flowing
const chunk = decoder.decode(value, { stream: true });
process.stdout.write(chunk);
}
idle.clear();
This pattern aborts only if 30 seconds pass with zero data, not if the total stream runs longer than 30 seconds. For most client libraries built on axios or node-fetch, look for separate timeout and socketTimeout-style options, or implement the same idle-reset logic manually around the stream reader.
Step 3: Fix proxy and load balancer idle timeouts
If you're running Claude API calls through your own backend before relaying to the browser, check your infrastructure's idle timeout settings:
- nginx: increase
proxy_read_timeoutandproxy_send_timeout(e.g.300s), and disable buffering for SSE withproxy_buffering off;. - AWS ALB: increase the idle timeout attribute on the load balancer (default is often 60 seconds).
- API Gateway / serverless functions: many have hard execution limits (e.g. 30s on some API Gateway configurations) that are incompatible with long streams — route streaming traffic around these, not through them.
- Cloudflare / CDNs in front of your API: confirm streaming/SSE passthrough is enabled and buffering is disabled for the relevant route.
A single misconfigured proxy layer is the most common cause of "it works locally but times out in production" reports.
Step 4: Handle disconnects with resumption, not blind retries
Even with correct timeouts, networks drop connections. Don't retry a streaming request from scratch by default — that duplicates partial output and wastes tokens. Instead:
- Track how much content has already been streamed.
- On disconnect, decide whether to discard and restart cleanly (simplest, safest for most apps) or prompt the model to continue from where it left off (more complex, rarely worth it for chat UIs).
- Apply backoff only to the retry attempt itself, not to the idle timeout logic.
Step 5: Consider removing a layer entirely
A lot of streaming timeout pain comes from maintaining your own proxy between the browser and the model provider — one more place for idle timeouts, buffering, and TLS renegotiation to go wrong. If your app just needs a reliable streaming endpoint without managing that infrastructure, SubToAPI exposes Claude as a standard HTTPS API with streaming support already configured correctly end to end — no proxy timeouts to debug on your side. The streaming docs show the exact event format, and the quickstart gets a working key in a couple of minutes via signup.
Checklist before you ship
- Idle timeout resets on every chunk, not a single total-duration timeout
- Reverse proxy
read_timeoutraised above your longest expected generation time - Buffering disabled on any proxy/CDN sitting in front of the stream
- Serverless/API Gateway execution limits don't apply to the streaming path
- Reconnect logic doesn't blindly re-run the full prompt
FAQ
Why does my stream work for short responses but time out on long ones? Short responses finish before any idle or total-duration timeout kicks in. Long responses expose a timeout set too low for the full generation time, or a proxy buffering output until it has "enough" data, which looks like a hang to the client.
Should I set a timeout at all for streaming requests? Yes, but make it an idle timeout that resets on each received chunk, not a single fixed duration applied to the whole request. A stream still sending tokens shouldn't be killed just because it's running long.
Is this a Claude API limitation? No — it's almost always a client, proxy, or load balancer timeout configured for short-lived requests. Correctly configuring idle timeouts and disabling intermediate buffering resolves it in the vast majority of cases.