← Blog

Claude API Streaming: Chunked Response Handling Guide

2026-10-04 · 5 min read · SubToAPI Team

Why chunked response handling trips people up

When you set "stream": true on a Claude API request, you don't get back one JSON blob — you get a chunked HTTP response made up of multiple Server-Sent Events (SSE), each carrying a fragment of the final message. The chunked part matters because TCP packets and HTTP chunks don't align with event boundaries: a single data: line can arrive split across two reads, or two full events can arrive in one read. If you parse naively (e.g. JSON.parse() on every chunk you receive from the socket), you will eventually hit malformed JSON errors in production, usually under load or on slow networks.

This guide covers how Claude's streaming format is structured, how to buffer and parse it correctly, and the common failure modes (partial JSON, out-of-order tool-use deltas, dropped connections) along with working code.

The shape of a Claude streaming response

A streaming request returns Content-Type: text/event-stream with a sequence of named events, each as a event: line followed by a data: line containing JSON:

event: message_start
data: {"type":"message_start","message":{"id":"msg_01...","role":"assistant",...}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":12}}

event: message_stop
data: {"type":"message_stop"}

Each event is separated by a blank line. That blank line is your real delimiter — not the chunk boundary from the transport layer. This distinction is the root of almost every "my stream parser crashes randomly" bug.

The correct buffering strategy

Treat the incoming bytes as a continuous stream and accumulate them in a string buffer. Split on double newlines (\n\n) to extract complete events, and keep any trailing partial event in the buffer for the next read.

async function readStream(response, onEvent) {
  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = "";

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    buffer += decoder.decode(value, { stream: true });
    const parts = buffer.split("\n\n");
    buffer = parts.pop(); // keep incomplete tail

    for (const part of parts) {
      const lines = part.split("\n");
      const dataLine = lines.find((l) => l.startsWith("data:"));
      if (!dataLine) continue;
      const json = dataLine.slice(5).trim();
      if (json === "[DONE]") continue;
      try {
        onEvent(JSON.parse(json));
      } catch (err) {
        console.error("Failed to parse event chunk:", json);
      }
    }
  }
}

Three details matter here:

Reassembling the final message

content_block_delta events only give you fragments. To reconstruct the full text, accumulate deltas by index:

let blocks = {};

function onEvent(evt) {
  if (evt.type === "content_block_start") {
    blocks[evt.index] = evt.content_block.text ?? "";
  }
  if (evt.type === "content_block_delta" && evt.delta.type === "text_delta") {
    blocks[evt.index] += evt.delta.text;
  }
}

For tool use, deltas arrive as input_json_delta fragments that need to be concatenated into a single JSON string before parsing — don't try to parse each fragment individually, since a tool call's arguments are usually streamed as partial, syntactically invalid JSON until the final chunk.

Handling dropped or stalled connections

Chunked streams can stall mid-response — a proxy timeout, a mobile network drop, a load balancer idle-connection kill. Build in:

  1. A read timeout per chunk, not just per request. If no bytes arrive for N seconds, abort and retry rather than hanging indefinitely.
  2. Idempotent retry logic that re-sends the same request rather than resuming mid-stream — Claude's API does not support resuming a stream from a given offset.
  3. Graceful partial-output handling on your UI side: if a stream dies after content_block_start but before message_stop, you should still render what you have and mark the response as incomplete rather than discarding it.

Where this gets harder in production

This buffering logic is straightforward for a single request in a demo, but it gets noticeably more complex once you're running it across many concurrent requests, multiple API keys, and client types (browser fetch, server-to-server, mobile). You end up re-implementing the same SSE buffering, reconnection, and usage-tracking logic in every service that talks to Claude.

This is one of the reasons teams put a layer like SubToAPI between their apps and their Claude access: streaming is already handled correctly end to end, so your client code just consumes clean SSE events without worrying about chunk boundaries, UTF-8 splitting, or stalled-connection retries. See the streaming docs for the exact event format, or the quickstart to get an API key running in a few minutes. Pricing starts at €9/month on the Solo plan, with team seats on the pricing page.

A minimal correct consumer, end to end

curl -N https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 512,
    "stream": true,
    "messages": [{"role": "user", "content": "Explain SSE in two sentences."}]
  }'

curl -N disables output buffering so you see events as they arrive — useful for debugging whether your issue is in the server response or your client-side parsing. Full request/response field details are in the messages docs.

questions

Why do I get "Unexpected end of JSON input" errors when parsing Claude's stream? You're almost certainly parsing raw chunk data instead of buffering until a full event (ending in \n\n) has arrived. Split on the double-newline delimiter first, keep incomplete tails in a buffer, and only parse complete data: lines.

Can I resume a Claude stream after a connection drop? No — there's no offset-based resume mechanism. The standard approach is to detect the stall with a read timeout, discard the partial response, and retry the full request.

How do I stream tool-use arguments correctly? Tool inputs arrive as input_json_delta fragments tied to a content block index. Concatenate all fragments for that index into one string and only call JSON.parse once the block's content_block_stop event arrives — intermediate fragments are not valid JSON on their own.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →