Claude API SSE Streaming Implementation Guide
Streaming responses from the Claude API means handling Server-Sent Events (SSE): a persistent HTTP connection that pushes response chunks as they're generated instead of waiting for the full completion. This guide shows exactly how to set up streaming, parse the event types Claude sends, and handle the edge cases that break naive implementations in production.
If you're here because your streaming client hangs, drops partial tokens, or fails silently on errors, the fixes below address the most common causes: incomplete SSE line buffering, missing event-type handling, and not accounting for content_block_delta vs message_delta events separately.
Why SSE instead of polling or WebSockets
SSE is a one-way protocol over plain HTTP. The server keeps the connection open and sends text-formatted events as data: {...}\n\n lines. For LLM output, this is the right fit because:
- You only need server-to-client push, not bidirectional messaging.
- It works through standard HTTP infrastructure (proxies, load balancers) without special upgrade handshakes like WebSockets require.
- Browsers and most HTTP clients have built-in or trivial support for parsing it.
Claude's streaming API uses this format: when you set "stream": true in your request, the response body becomes a sequence of named events rather than a single JSON object.
Setting up the request
A streaming request looks almost identical to a standard one, with two differences: the stream flag and how you read the response body.
curl -N https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"stream": true,
"messages": [{"role": "user", "content": "Explain SSE in two sentences."}]
}'
The -N flag disables curl's output buffering so you see events as they arrive instead of all at once at the end. This is the single most common mistake when testing streaming manually: without it, curl looks like it's not streaming even though the server is sending data correctly.
The event types you'll receive
Claude's SSE stream sends several distinct event types, each with its own event: line and JSON payload:
message_start— fires once, contains the initial message object (id, model, empty content).content_block_start— a new content block (text or tool use) is beginning.content_block_delta— the actual incremental content. For text blocks this is{"type": "text_delta", "text": "..."}.content_block_stop— the current block is complete.message_delta— top-level message metadata updates, includingstop_reasonand usage totals.message_stop— the stream is finished.
A client that only listens for content_block_delta and ignores the rest will miss stop reasons and final token counts — which matters if you're billing or logging usage per request.
Parsing SSE correctly in JavaScript
The most reliable way to parse SSE manually (without a library) is to buffer raw text and split on double newlines, since a single data: payload can arrive split across multiple TCP packets:
async function streamMessage(body) {
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({ ...body, stream: true }),
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const events = buffer.split("\n\n");
buffer = events.pop(); // keep incomplete trailing chunk
for (const chunk of events) {
const lines = chunk.split("\n");
const dataLine = lines.find((l) => l.startsWith("data:"));
if (!dataLine) continue;
const payload = JSON.parse(dataLine.slice(5).trim());
if (payload.type === "content_block_delta") {
process.stdout.write(payload.delta.text ?? "");
}
}
}
}
The key detail: never assume a read() call gives you a complete event. Always split on \n\n and hold back the last, possibly-incomplete fragment for the next iteration. This is where most hand-rolled SSE parsers silently corrupt output under load.
Handling errors mid-stream
Unlike a standard request, a streaming connection can fail after you've already received partial content. You need to handle:
errorevents sent inline in the stream (overloaded model, rate limits triggered mid-response).- Connection drops from proxies or load balancers with aggressive idle timeouts — keep-alive comments (
: ping\n\n) help here but aren't always enough. - Partial tool-use blocks — if a
content_block_stopnever arrives for a tool call, don't execute it as if it's complete.
Wrap your reader loop in a try/catch, and track whether you've seen message_stop before treating the stream as successfully finished. If the connection closes without it, treat the response as incomplete and retry with the same parameters.
Where SubToAPI fits in
If you're building this for an internal tool or product and don't want to maintain SSE parsing, reconnection logic, and per-key usage tracking yourself, SubToAPI wraps your existing Claude access in a standard HTTPS API with streaming support already implemented correctly — including proper event framing and usage metadata per request. You get a sub_live_... application key instead of managing raw model credentials, with the same request shape shown above. See the streaming docs and quickstart for details, or check pricing if you're evaluating it for a team.
Testing your implementation
Before shipping, verify your client against these cases:
- A normal short response — confirm text arrives incrementally, not all at once.
- A long response (1000+ tokens) — confirm no chunks are dropped or duplicated.
- A tool-use response — confirm
content_block_start/stoppairs are tracked per block index. - A forced error (invalid model name) — confirm your error handling fires instead of hanging.
- A mid-stream network interruption (kill the connection manually) — confirm you detect the missing
message_stop.
Questions
Does SSE guarantee delivery if my connection drops? No. SSE has no built-in resume mechanism for LLM responses — if the connection drops, you need to detect the missing message_stop event and retry the full request.
Can I use WebSockets instead of SSE for Claude streaming? The Claude API only exposes SSE for streaming responses, not WebSockets. SSE is sufficient since the data flow is one-directional.
Why does my curl command not show streaming output? Curl buffers output by default. Add the -N flag to disable buffering and see events arrive incrementally instead of all at once.