Claude API Streaming: Server-Sent Events Guide
What SSE streaming means for the Claude API
When you call the Claude Messages API with "stream": true, Anthropic doesn't send back one big JSON blob after the model finishes thinking. Instead it opens an HTTP connection and pushes the response token-by-token as a sequence of Server-Sent Events (SSE) — a simple, long-lived text protocol where each chunk is prefixed with event: and data: lines. Your client reads the connection as it arrives, instead of waiting for the full completion.
This matters for anything user-facing: chat UIs, coding assistants, voice agents. Without streaming, a 2,000-token response can mean 10-20 seconds of silence before anything renders. With SSE, the first tokens show up in a few hundred milliseconds and the rest trickles in as the model generates it. This guide covers the actual event types Claude sends, how to parse them correctly in raw curl/JavaScript, and where it's easy to get streaming wrong.
The anatomy of a Claude SSE stream
A streamed Claude response is not just a stream of text deltas — it's a sequence of distinct event types that together describe the full lifecycle of the message:
message_start— the response object is created, includesid,model, and initial (empty) usage.content_block_start— a new content block begins (text, or a tool_use block).content_block_delta— incremental content. For text this istext_delta; for tool calls it'sinput_json_deltawith partial JSON fragments.content_block_stop— the current content block is complete.message_delta— top-level changes, most importantlystop_reasonand updatedusage.message_stop— the stream is finished.ping— keep-alive events sent periodically; safe to ignore.error— something went wrong mid-stream (rate limit, overload, etc.).
Each of these arrives as its own SSE frame. A raw frame looks like this:
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
Note the blank line after data: — that's part of the SSE spec and marks the end of the event.
Streaming with curl
The quickest way to see this in action is curl with -N to disable buffering:
curl -N https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 512,
"stream": true,
"messages": [{"role": "user", "content": "Write a haiku about rivers"}]
}'
You'll see the full sequence of message_start, several content_block_delta events carrying one or a few tokens each, then content_block_stop, message_delta, and message_stop scroll past in real time.
Parsing SSE in JavaScript
The browser and Node both support reading a ReadableStream body line by line. Here's a minimal parser that reconstructs the full text as deltas arrive:
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 512,
stream: true,
messages: [{ role: "user", content: "Explain SSE in two sentences" }],
}),
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let fullText = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop(); // keep incomplete line for next chunk
for (const line of lines) {
if (!line.startsWith("data:")) continue;
const payload = line.slice(5).trim();
if (!payload) continue;
const event = JSON.parse(payload);
if (event.type === "content_block_delta" && event.delta?.type === "text_delta") {
fullText += event.delta.text;
process.stdout.write(event.delta.text);
}
}
}
The important detail most people get wrong: network chunks don't align with SSE event boundaries. A single read() call might deliver half an event, or three events at once. You have to buffer partial lines and only parse once you've got a complete data: line — the example above does this with the buffer.split("\n") / lines.pop() pattern.
Handling errors and reconnects
Mid-stream errors show up as an event: error frame with a body like {"type":"error","error":{"type":"overloaded_error","message":"..."}}. Because the connection is already open and may have delivered partial content, you need to decide per use case whether to discard the partial output or keep it and retry just the remainder. There's no built-in resume token — if the stream drops, you restart the request from scratch.
Also watch for:
- Idle timeouts on proxies/load balancers in front of your app — SSE connections that sit quiet too long between tokens can get killed by infrastructure that doesn't know about
pingevents. - Buffering by intermediaries — some CDNs and reverse proxies buffer responses by default, which defeats the purpose of streaming. Disable buffering explicitly (
X-Accel-Buffering: noon nginx, for example) if you're proxying streamed responses yourself.
Streaming through SubToAPI
If you're exposing Claude to your own app or team without wiring in the full Anthropic SDK and billing setup, SubToAPI gives you the same SSE streaming behavior behind a sub_live_... application key, plus usage metadata per request and per-seat visibility into who's consuming tokens. The endpoint and event format match what's described above — stream: true on a request to /v1/messages streams content_block_delta events exactly the same way. See /docs/streaming for setup and /docs/quickstart if you're starting from zero. Plans start at €9/month on the Solo tier with a free trial at /signup, and /pricing has the full breakdown for teams.
Questions
Does streaming cost more than non-streaming requests? No. Token usage and pricing are identical — streaming only changes how the response is delivered, not how many tokens are billed.
Can I stream tool use (function calling) responses? Yes. Tool calls arrive as content_block_start with a tool_use block followed by input_json_delta events that build up the arguments as partial JSON fragments; you concatenate them before parsing. See /docs/tools for the full tool-use flow.
Why does my stream stop producing events but the connection stays open? That's usually a ping event or a long generation pause — check you're not filtering those out as errors. If the connection truly hangs with no ping or message_stop, check for buffering proxies between your server and the client.