Claude API Streaming Response Chunk Parsing Guide
When you stream a response from the Claude API, you don't get one clean JSON object back — you get a sequence of Server-Sent Events (SSE), each carrying a small piece of the final answer. Parsing them correctly means handling event types, partial JSON fragments, and chunk boundaries that don't always line up with a full event.
This article walks through the exact structure of Claude's streaming payload, shows working parsing code for Node.js and the browser, and covers the edge cases that trip people up: split JSON across TCP packets, multiple events in one chunk, and tool-use deltas that arrive as fragments you have to reassemble yourself.
How Claude structures a streaming response
A streaming request returns text/event-stream content. Each event looks like this:
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
The event types you'll see, in order, are:
message_start— metadata about the response (model, id, initial usage)content_block_start— a new content block begins (text or tool_use)content_block_delta— the actual incremental content, sent repeatedlycontent_block_stop— the block is completemessage_delta— top-level changes likestop_reasonand final usagemessage_stop— the stream is doneping— keep-alive, safe to ignore
The important part is that a "chunk" from your HTTP client is not the same thing as an "event." A single chunk from fetch or a raw socket read can contain zero, one, or several events, and it can also end in the middle of one. Any parser that assumes one chunk equals one event will silently drop or corrupt data under load.
A correct parsing loop
The safe pattern is: accumulate raw text in a buffer, split on double newlines (\n\n, which separates SSE events), process each complete event, and keep the leftover partial text in the buffer for the next read.
async function streamClaude(response, onDelta) {
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const events = buffer.split("\n\n");
buffer = events.pop(); // keep the incomplete tail
for (const raw of events) {
const line = raw.split("\n").find((l) => l.startsWith("data:"));
if (!line) continue;
const json = line.replace("data:", "").trim();
if (json === "[DONE]") return;
const event = JSON.parse(json);
if (event.type === "content_block_delta" && event.delta?.type === "text_delta") {
onDelta(event.delta.text);
}
}
}
}
The two details that matter here:
decoder.decode(value, { stream: true })— withoutstream: true, a multi-byte UTF-8 character split across two chunks (common with emoji or accented text) will corrupt the output.buffer = events.pop()— always assume the last "event" you split out might be incomplete, and hold it for the next iteration rather than parsing it immediately.
Parsing tool use deltas
Tool calls stream differently from text. The arguments arrive as input_json_delta fragments that are partial JSON strings, not parseable JSON on their own — you concatenate them and only parse the full string once the block stops:
let toolInputBuffer = "";
function handleEvent(event) {
if (event.type === "content_block_start" && event.content_block.type === "tool_use") {
toolInputBuffer = "";
}
if (event.type === "content_block_delta" && event.delta.type === "input_json_delta") {
toolInputBuffer += event.delta.partial_json;
}
if (event.type === "content_block_stop") {
const args = JSON.parse(toolInputBuffer); // now safe to parse
}
}
Trying to JSON.parse each individual partial_json fragment will throw constantly — it's not meant to be valid JSON until fully reassembled.
Parsing with curl for debugging
curl is useful for inspecting raw events without any client-side parsing logic hiding what's happening on the wire:
curl -N https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 512,
"stream": true,
"messages": [{"role": "user", "content": "Count to five"}]
}'
The -N flag disables curl's output buffering so you see events as they arrive rather than all at once at the end — critical for confirming your server is actually streaming and not just buffering internally before responding.
Where this gets brittle in production
- Proxies and load balancers that buffer responses. Some reverse proxies buffer the full response before forwarding it, which defeats streaming entirely even if your parser is correct. Nginx needs
proxy_buffering offfor SSE endpoints, for example. - Reconnection logic. If a connection drops mid-stream, you need to decide whether to resume, retry from scratch, or surface a partial result — Claude's streaming API doesn't have resumable cursors, so most implementations just retry the full request.
- Usage accounting. Token usage for streamed responses only arrives in the final
message_deltaevent, so any cost tracking or rate limiting logic has to wait for the stream to close rather than estimating from deltas as they arrive.
If you're building this parsing layer yourself across multiple services, it's worth checking whether you actually need to. SubToAPI exposes Claude through a standard HTTPS API with the same SSE event structure described here, plus per-key usage metadata attached to the final event, so your client code doesn't have to reconcile usage separately. The streaming docs cover the exact event sequence, and the quickstart has working examples in curl and JavaScript. Plans start at €9/month for solo use with team seats from €19, and a free trial is available at signup.
Questions
Why does my JSON.parse fail on some chunks but not others? You're almost certainly parsing raw chunk boundaries instead of SSE event boundaries. Buffer incoming text and only parse after splitting on \n\n, keeping incomplete tails for the next read.
Do I need to reconstruct tool_use arguments manually? Yes. input_json_delta events send partial_json fragments that must be concatenated into a single string and parsed only once content_block_stop fires — parsing each fragment individually will throw errors.
Why does my proxy seem to buffer the whole stream before sending it? Many reverse proxies buffer HTTP responses by default. For Nginx, set proxy_buffering off; for other proxies, check for equivalent settings, otherwise clients receive the entire response at once regardless of correct client-side parsing.