Claude API Streaming Chunk Parsing Guide
What this guide answers
If you're building a chat UI, CLI tool, or backend proxy on top of Claude, you've probably hit a wall trying to parse the streamed response. The Claude API streams responses as Server-Sent Events (SSE), where each event carries a JSON payload describing a small piece of the final message. Parsing that stream correctly means handling multiple event types, accumulating partial text and partial tool-call JSON, and dealing with chunks that arrive split across TCP packets.
This guide walks through the actual event structure, shows working parsing code, and covers the edge cases that trip people up: incomplete lines, ping events, tool-use JSON deltas, and error events mid-stream. If you're consuming Claude through a gateway like SubToAPI instead of calling Anthropic directly, the same SSE format applies — see /docs/streaming for the exact event shapes returned.
The SSE event types you'll see
When you set "stream": true on a Messages API request, the response body is a sequence of event: / data: pairs. Each data: line contains a JSON object, and the event: line names its type. In order, a typical stream looks like this:
message_start— the message shell (id, model, empty content, usage so far)content_block_start— a new content block begins (text or tool_use)content_block_delta— repeated many times, each carrying a fragment of text or JSONcontent_block_stop— the current block is completemessage_delta— updates to top-level fields likestop_reasonand final usagemessage_stop— the stream is doneping— periodic keep-alive, carries no content
A multi-block message (text followed by a tool call) will repeat steps 2–4 for each block before moving to message_delta and message_stop.
Parsing raw SSE lines
Each chunk from the server looks like:
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
Note the blank line separating events — that's the SSE boundary. A naive split('\n') on every network chunk will break because TCP doesn't respect those boundaries; a single data: line can arrive split across two reads, or two full events can arrive in one read.
The safe approach is to buffer incoming bytes and only process complete lines ending in \n:
let buffer = "";
function handleChunk(rawChunk, onEvent) {
buffer += rawChunk;
const lines = buffer.split("\n");
buffer = lines.pop(); // keep the last (possibly incomplete) line
let currentEvent = null;
for (const line of lines) {
if (line.startsWith("event:")) {
currentEvent = line.slice(6).trim();
} else if (line.startsWith("data:")) {
const json = JSON.parse(line.slice(5).trim());
onEvent(currentEvent, json);
}
}
}
If you're using fetch with a ReadableStream, feed each decoded chunk into handleChunk. This is the core logic regardless of whether you're calling Anthropic's endpoint or a compatible gateway.
Accumulating text deltas
For plain text responses, you only care about content_block_delta events where delta.type === "text_delta". Concatenate delta.text in order:
let fullText = "";
function onEvent(event, data) {
if (event === "content_block_delta" && data.delta?.type === "text_delta") {
fullText += data.delta.text;
renderPartial(fullText); // update your UI incrementally
}
if (event === "message_stop") {
finalize(fullText);
}
}
That's it for simple cases. The complexity goes up once tool use enters the picture.
Parsing tool_use deltas
When Claude calls a tool mid-response, the content block type is tool_use, and its input JSON arrives incrementally as input_json_delta fragments — not as complete JSON per chunk. You cannot JSON.parse each delta individually; you must concatenate the raw string fragments and parse only once the block is closed.
let toolInputBuffer = "";
let toolName = null;
function onEvent(event, data) {
if (event === "content_block_start" && data.content_block.type === "tool_use") {
toolName = data.content_block.name;
toolInputBuffer = "";
}
if (event === "content_block_delta" && data.delta?.type === "input_json_delta") {
toolInputBuffer += data.delta.partial_json;
}
if (event === "content_block_stop" && toolName) {
const toolInput = JSON.parse(toolInputBuffer);
runTool(toolName, toolInput);
toolName = null;
}
}
This pattern matters because the partial JSON fragments are frequently not valid JSON on their own — a fragment might end mid-string or mid-key. Trying to parse each one will throw. See /docs/tools for how tool-call payloads are structured end to end.
Handling errors and disconnects mid-stream
Two failure modes are common in production:
- An
errorevent arrives mid-stream instead ofmessage_stop. Always check forevent === "error"and surface or retry rather than assuming the stream always ends cleanly. - The connection drops before
message_stop. Track whether you've received a terminal event; if the HTTP stream closes without one, treat the response as incomplete and retry the request rather than trusting a partial accumulation.
Also don't ignore ping events — they're harmless, but if your parser throws on unrecognized event types instead of skipping them, a keep-alive will crash your client.
Using usage data from the stream
The message_start event includes initial usage (input tokens), and message_delta includes the final output_tokens once generation finishes. If you're tracking per-request cost or building usage dashboards, read both rather than estimating from character counts — token counts from estimation are unreliable. SubToAPI's dashboard does this accounting automatically per API key if you'd rather not parse usage fields yourself; see /docs/messages for the full field reference, or check /pricing if you're evaluating plans. A free trial is available at /signup.
Putting it together
A production-grade parser needs: line buffering across chunks, a switch on event type, text accumulation for text_delta, string concatenation (not parsing) for input_json_delta until the block closes, explicit handling for error and unexpected stream termination, and graceful skipping of ping. Get those five things right and the rest of your streaming UI logic — rendering, retry, tool dispatch — sits cleanly on top.
questions
Why does my JSON.parse fail on tool_use delta chunks? Because each input_json_delta only contains a fragment of the JSON string, not a complete object. Concatenate all fragments for a block and parse once, after content_block_stop.
Do I need to handle ping events specially? No, but your parser must not crash on them. They carry no payload and exist only to keep the connection alive — simply ignore any event type you don't explicitly handle.
How do I know when a stream ended successfully vs. dropped? Track whether a message_stop event was received. If the HTTP connection closes without it, or you see an error event instead, treat the response as incomplete and retry the request.