← Blog

Claude API Streaming Token by Token: How It Works

2026-09-28 · 4 min read · SubToAPI Team

When people search for "claude api streaming token by token," they're usually trying to answer one of two questions: how does the streaming protocol actually deliver individual tokens, and how do I parse that stream correctly in my own code without dropping characters or breaking on partial JSON. This article answers both.

Claude's API streams responses using Server-Sent Events (SSE), not raw token-by-token bytes. Each event contains a small JSON payload, and the actual text tokens arrive inside content_block_delta events as incremental text fragments. Understanding the event structure is the difference between a streaming integration that feels instant and one that stutters, double-renders text, or crashes on malformed JSON.

Why Stream at the Token Level

A non-streaming request waits for the full response before returning anything. For a 2,000-token answer, that can mean 10-30 seconds of silence before the user sees a single character. Token-level streaming instead pushes text as it's generated, so the first word can appear in under a second. This matters for:

The SSE Event Sequence

A streamed Claude response emits a defined sequence of event types, in order:

  1. message_start — the message object is created, with empty content and initial usage stats.
  2. content_block_start — a new content block begins (text or tool_use).
  3. content_block_delta — repeated many times, each carrying a small piece of text (text_delta) or a partial JSON fragment (input_json_delta for tool calls).
  4. content_block_stop — the current block is complete.
  5. message_delta — carries updates like stop_reason and final usage counts.
  6. message_stop — the stream is finished.

The tokens themselves live inside the text_delta field of content_block_delta events. A single API "token" from the model's perspective often doesn't map 1:1 to a delta chunk — deltas are batched for efficiency — but the effect from the client's side is the same: text arrives incrementally.

Raw SSE Example

event: message_start
data: {"type":"message_start","message":{"id":"msg_01...","role":"assistant","content":[],"usage":{"input_tokens":15}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":", world"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":8}}

event: message_stop
data: {"type":"message_stop"}

Note that each data: line is a complete, self-contained JSON object. You never need to concatenate JSON fragments across lines — only the text fields need concatenating to reconstruct the full message.

Parsing the Stream Correctly

The most common mistake is trying to parse the stream as one giant JSON blob. Instead, process it line by line, buffer partial lines until you have a full SSE event, then parse just the data: payload.

curl example

curl https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-sonnet-4-20250514",
    "max_tokens": 512,
    "stream": true,
    "messages": [{"role": "user", "content": "Explain SSE streaming in two sentences."}]
  }'

Node.js parsing example

const response = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "content-type": "application/json",
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-20250514",
    max_tokens: 512,
    stream: true,
    messages: [{ role: "user", content: "Stream this word by word." }],
  }),
});

const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let fullText = "";

while (true) {
  const { done, value } = await reader.read();
  if (done) break;
  buffer += decoder.decode(value, { stream: true });

  const lines = buffer.split("\n");
  buffer = lines.pop(); // keep incomplete line for next chunk

  for (const line of lines) {
    if (!line.startsWith("data:")) continue;
    const payload = line.slice(5).trim();
    if (!payload) continue;

    const event = JSON.parse(payload);
    if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
      fullText += event.delta.text;
      process.stdout.write(event.delta.text); // token-by-token output
    }
  }
}

This handles the common edge cases: chunks arriving mid-line, empty keep-alive lines, and events you don't care about (like ping, which Anthropic sends periodically to keep the connection alive).

Streaming Through a Managed API Key

If you're running Claude access through SubToAPI, the streaming contract is the same SSE format, so this parsing code works unchanged — you just point requests at https://api.subtoapi.app/v1/messages with Authorization: Bearer $SUBTOAPI_KEY instead of an x-api-key header. That's useful if your team is already using SubToAPI for issuing scoped sub_live_ keys per application or environment rather than sharing a single raw key. See /docs/streaming for the exact request shape and /docs/messages for the full message schema.

Common Pitfalls

FAQs

Does Claude stream literal model tokens or just text chunks? The API streams text fragments via text_delta events, not raw model tokens. Fragment size varies and isn't guaranteed to align with tokenizer boundaries, but the effect — text appearing progressively — is the same for UI purposes.

Can I get usage/cost data while streaming? Input token usage arrives in message_start; final output token counts arrive in the message_delta event near the end of the stream, so you get accurate cost data without waiting for a separate call.

Is streaming slower or more expensive than non-streaming requests? No — total generation time and token costs are the same either way. Streaming only changes when the client receives data, not how much data is generated or billed.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →