Claude API Streaming Token by Token: How It Works
When people search for "claude api streaming token by token," they're usually trying to answer one of two questions: how does the streaming protocol actually deliver individual tokens, and how do I parse that stream correctly in my own code without dropping characters or breaking on partial JSON. This article answers both.
Claude's API streams responses using Server-Sent Events (SSE), not raw token-by-token bytes. Each event contains a small JSON payload, and the actual text tokens arrive inside content_block_delta events as incremental text fragments. Understanding the event structure is the difference between a streaming integration that feels instant and one that stutters, double-renders text, or crashes on malformed JSON.
Why Stream at the Token Level
A non-streaming request waits for the full response before returning anything. For a 2,000-token answer, that can mean 10-30 seconds of silence before the user sees a single character. Token-level streaming instead pushes text as it's generated, so the first word can appear in under a second. This matters for:
- Chat interfaces — perceived latency drops dramatically even if total generation time is unchanged.
- CLI tools and agents — you can start processing partial output (e.g., detecting a tool call) before generation finishes.
- Long-form generation — summaries, code, or documents render progressively instead of appearing all at once.
The SSE Event Sequence
A streamed Claude response emits a defined sequence of event types, in order:
message_start— the message object is created, with empty content and initial usage stats.content_block_start— a new content block begins (text or tool_use).content_block_delta— repeated many times, each carrying a small piece of text (text_delta) or a partial JSON fragment (input_json_deltafor tool calls).content_block_stop— the current block is complete.message_delta— carries updates likestop_reasonand final usage counts.message_stop— the stream is finished.
The tokens themselves live inside the text_delta field of content_block_delta events. A single API "token" from the model's perspective often doesn't map 1:1 to a delta chunk — deltas are batched for efficiency — but the effect from the client's side is the same: text arrives incrementally.
Raw SSE Example
event: message_start
data: {"type":"message_start","message":{"id":"msg_01...","role":"assistant","content":[],"usage":{"input_tokens":15}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":", world"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":8}}
event: message_stop
data: {"type":"message_stop"}
Note that each data: line is a complete, self-contained JSON object. You never need to concatenate JSON fragments across lines — only the text fields need concatenating to reconstruct the full message.
Parsing the Stream Correctly
The most common mistake is trying to parse the stream as one giant JSON blob. Instead, process it line by line, buffer partial lines until you have a full SSE event, then parse just the data: payload.
curl example
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-d '{
"model": "claude-sonnet-4-20250514",
"max_tokens": 512,
"stream": true,
"messages": [{"role": "user", "content": "Explain SSE streaming in two sentences."}]
}'
Node.js parsing example
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
},
body: JSON.stringify({
model: "claude-sonnet-4-20250514",
max_tokens: 512,
stream: true,
messages: [{ role: "user", content: "Stream this word by word." }],
}),
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let fullText = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop(); // keep incomplete line for next chunk
for (const line of lines) {
if (!line.startsWith("data:")) continue;
const payload = line.slice(5).trim();
if (!payload) continue;
const event = JSON.parse(payload);
if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
fullText += event.delta.text;
process.stdout.write(event.delta.text); // token-by-token output
}
}
}
This handles the common edge cases: chunks arriving mid-line, empty keep-alive lines, and events you don't care about (like ping, which Anthropic sends periodically to keep the connection alive).
Streaming Through a Managed API Key
If you're running Claude access through SubToAPI, the streaming contract is the same SSE format, so this parsing code works unchanged — you just point requests at https://api.subtoapi.app/v1/messages with Authorization: Bearer $SUBTOAPI_KEY instead of an x-api-key header. That's useful if your team is already using SubToAPI for issuing scoped sub_live_ keys per application or environment rather than sharing a single raw key. See /docs/streaming for the exact request shape and /docs/messages for the full message schema.
Common Pitfalls
- Not handling
pingevents — they have nodatapayload structure you need, just skip anything that isn'tcontent_block_delta,message_start,message_delta, ormessage_stop. - Assuming one delta equals one token — deltas are text fragments, not guaranteed single tokens. Don't build logic that depends on delta boundaries matching word or token boundaries.
- Ignoring
input_json_delta— if the response includes tool use, JSON arguments also stream incrementally and must be concatenated before parsing as JSON. - Buffering the whole response before displaying anything — defeats the purpose of streaming. Flush text to the UI or terminal as each delta arrives.
FAQs
Does Claude stream literal model tokens or just text chunks? The API streams text fragments via text_delta events, not raw model tokens. Fragment size varies and isn't guaranteed to align with tokenizer boundaries, but the effect — text appearing progressively — is the same for UI purposes.
Can I get usage/cost data while streaming? Input token usage arrives in message_start; final output token counts arrive in the message_delta event near the end of the stream, so you get accurate cost data without waiting for a separate call.
Is streaming slower or more expensive than non-streaming requests? No — total generation time and token costs are the same either way. Streaming only changes when the client receives data, not how much data is generated or billed.