Claude API Streaming: Chunked Response Handling Guide
Why chunked response handling trips people up
When you set "stream": true on a Claude API request, you don't get back one JSON blob — you get a chunked HTTP response made up of multiple Server-Sent Events (SSE), each carrying a fragment of the final message. The chunked part matters because TCP packets and HTTP chunks don't align with event boundaries: a single data: line can arrive split across two reads, or two full events can arrive in one read. If you parse naively (e.g. JSON.parse() on every chunk you receive from the socket), you will eventually hit malformed JSON errors in production, usually under load or on slow networks.
This guide covers how Claude's streaming format is structured, how to buffer and parse it correctly, and the common failure modes (partial JSON, out-of-order tool-use deltas, dropped connections) along with working code.
The shape of a Claude streaming response
A streaming request returns Content-Type: text/event-stream with a sequence of named events, each as a event: line followed by a data: line containing JSON:
event: message_start
data: {"type":"message_start","message":{"id":"msg_01...","role":"assistant",...}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":12}}
event: message_stop
data: {"type":"message_stop"}
Each event is separated by a blank line. That blank line is your real delimiter — not the chunk boundary from the transport layer. This distinction is the root of almost every "my stream parser crashes randomly" bug.
The correct buffering strategy
Treat the incoming bytes as a continuous stream and accumulate them in a string buffer. Split on double newlines (\n\n) to extract complete events, and keep any trailing partial event in the buffer for the next read.
async function readStream(response, onEvent) {
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split("\n\n");
buffer = parts.pop(); // keep incomplete tail
for (const part of parts) {
const lines = part.split("\n");
const dataLine = lines.find((l) => l.startsWith("data:"));
if (!dataLine) continue;
const json = dataLine.slice(5).trim();
if (json === "[DONE]") continue;
try {
onEvent(JSON.parse(json));
} catch (err) {
console.error("Failed to parse event chunk:", json);
}
}
}
}
Three details matter here:
- Never
JSON.parseraw socket chunks. Always split on the event delimiter first. - Keep the tail. The last element of
split("\n\n")is often incomplete — don't process it until more data arrives. - Use
{ stream: true }on the decoder. Without it,TextDecodercan corrupt multi-byte UTF-8 characters that get split across chunk boundaries (this is a real issue with emoji or non-Latin text in responses).
Reassembling the final message
content_block_delta events only give you fragments. To reconstruct the full text, accumulate deltas by index:
let blocks = {};
function onEvent(evt) {
if (evt.type === "content_block_start") {
blocks[evt.index] = evt.content_block.text ?? "";
}
if (evt.type === "content_block_delta" && evt.delta.type === "text_delta") {
blocks[evt.index] += evt.delta.text;
}
}
For tool use, deltas arrive as input_json_delta fragments that need to be concatenated into a single JSON string before parsing — don't try to parse each fragment individually, since a tool call's arguments are usually streamed as partial, syntactically invalid JSON until the final chunk.
Handling dropped or stalled connections
Chunked streams can stall mid-response — a proxy timeout, a mobile network drop, a load balancer idle-connection kill. Build in:
- A read timeout per chunk, not just per request. If no bytes arrive for N seconds, abort and retry rather than hanging indefinitely.
- Idempotent retry logic that re-sends the same request rather than resuming mid-stream — Claude's API does not support resuming a stream from a given offset.
- Graceful partial-output handling on your UI side: if a stream dies after
content_block_startbut beforemessage_stop, you should still render what you have and mark the response as incomplete rather than discarding it.
Where this gets harder in production
This buffering logic is straightforward for a single request in a demo, but it gets noticeably more complex once you're running it across many concurrent requests, multiple API keys, and client types (browser fetch, server-to-server, mobile). You end up re-implementing the same SSE buffering, reconnection, and usage-tracking logic in every service that talks to Claude.
This is one of the reasons teams put a layer like SubToAPI between their apps and their Claude access: streaming is already handled correctly end to end, so your client code just consumes clean SSE events without worrying about chunk boundaries, UTF-8 splitting, or stalled-connection retries. See the streaming docs for the exact event format, or the quickstart to get an API key running in a few minutes. Pricing starts at €9/month on the Solo plan, with team seats on the pricing page.
A minimal correct consumer, end to end
curl -N https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 512,
"stream": true,
"messages": [{"role": "user", "content": "Explain SSE in two sentences."}]
}'
curl -N disables output buffering so you see events as they arrive — useful for debugging whether your issue is in the server response or your client-side parsing. Full request/response field details are in the messages docs.
questions
Why do I get "Unexpected end of JSON input" errors when parsing Claude's stream? You're almost certainly parsing raw chunk data instead of buffering until a full event (ending in \n\n) has arrived. Split on the double-newline delimiter first, keep incomplete tails in a buffer, and only parse complete data: lines.
Can I resume a Claude stream after a connection drop? No — there's no offset-based resume mechanism. The standard approach is to detect the stall with a read timeout, discard the partial response, and retry the full request.
How do I stream tool-use arguments correctly? Tool inputs arrive as input_json_delta fragments tied to a content block index. Concatenate all fragments for that index into one string and only call JSON.parse once the block's content_block_stop event arrives — intermediate fragments are not valid JSON on their own.