Claude API Streaming Chat UI in React: Full Guide
Building a chat interface that streams Claude's response token-by-token — instead of waiting for the full reply — is what makes an AI product feel responsive. This guide shows how to wire a React frontend to a streaming Claude API endpoint, parse server-sent events as they arrive, and update the UI incrementally without janky re-renders.
The short answer: Claude's API supports streaming via server-sent events (SSE) when you set "stream": true in the request. On the frontend, you read the response body as a stream, parse each data: chunk, extract the text delta, and append it to your message state. Below is a complete, working pattern for React, including the backend proxy you need because Claude API keys should never be called directly from the browser.
Why you need a backend for streaming
You can't call the Claude API straight from client-side React — your API key would be exposed in the browser's network tab. You need a thin backend (Node.js, Next.js API route, or a hosted proxy) that holds the key and forwards the stream to the client.
If you don't want to run and maintain that proxy yourself, SubToAPI gives you an HTTPS endpoint with application keys (sub_live_...) that you can call from a server you control, with streaming, usage metadata, and rate limiting already handled. The examples below work the same way against api.subtoapi.app or against Anthropic's API directly — swap the base URL and headers.
Backend: a minimal streaming proxy
// server.js (Express)
app.post("/api/chat", async (req, res) => {
const upstream = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
stream: true,
messages: req.body.messages,
}),
});
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
const reader = upstream.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
res.write(decoder.decode(value));
}
res.end();
});
This just pipes the upstream SSE stream through to your own /api/chat endpoint, which the React app calls. See /docs/streaming for the full event format and /docs/quickstart for setup.
Frontend: consuming the stream in React
The core challenge in React is parsing SSE chunks (which can arrive split across multiple read() calls) and updating message state without triggering a re-render per character.
import { useState, useRef } from "react";
function useChatStream() {
const [messages, setMessages] = useState([]);
const [isStreaming, setIsStreaming] = useState(false);
const bufferRef = useRef("");
async function sendMessage(text) {
const userMsg = { role: "user", content: text };
setMessages((prev) => [...prev, userMsg, { role: "assistant", content: "" }]);
setIsStreaming(true);
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: [...messages, userMsg] }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
bufferRef.current = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
bufferRef.current += decoder.decode(value, { stream: true });
const lines = bufferRef.current.split("\n\n");
bufferRef.current = lines.pop(); // keep incomplete chunk for next read
for (const line of lines) {
if (!line.startsWith("data: ")) continue;
const payload = line.slice(6);
if (payload === "[DONE]") continue;
try {
const event = JSON.parse(payload);
if (event.type === "content_block_delta") {
const delta = event.delta?.text ?? "";
setMessages((prev) => {
const next = [...prev];
next[next.length - 1] = {
...next[next.length - 1],
content: next[next.length - 1].content + delta,
};
return next;
});
}
} catch {
// partial JSON, ignore until buffer completes
}
}
}
setIsStreaming(false);
}
return { messages, sendMessage, isStreaming };
}
Key details that matter in practice:
- Buffer incomplete lines. SSE chunks don't align with event boundaries, so you must accumulate text and only process complete
data:lines, keeping the remainder for the next read. - Update only the last message. Mutating the last array item avoids re-creating the whole message list on every token, which keeps rendering cheap even at high token rates.
{ stream: true }indecoder.decodeis required to correctly handle multi-byte UTF-8 characters split across chunk boundaries — otherwise emoji or non-ASCII text can render as garbled characters mid-stream.
Rendering the chat UI
function ChatWindow({ messages, isStreaming }) {
return (
<div className="chat-window">
{messages.map((m, i) => (
<div key={i} className={`message ${m.role}`}>
<span>{m.content}</span>
{isStreaming && i === messages.length - 1 && m.role === "assistant" && (
<span className="cursor">▍</span>
)}
</div>
))}
</div>
);
}
A blinking cursor on the last assistant message is enough visual feedback — you don't need a spinner since text is already appearing.
Handling stop and errors
Two things every streaming chat UI needs beyond happy-path rendering:
- A stop/cancel button. Store the
AbortControllerused in thefetchcall and call.abort()on click, then mark the message as interrupted in state. - Stream-level errors. If the upstream connection drops mid-response, catch the error in your read loop and append an error message rather than leaving the UI stuck showing a cursor forever.
const controller = new AbortController();
fetch("/api/chat", { signal: controller.signal, /* ... */ });
// later: controller.abort();
Scaling beyond a single user
Once you have multiple users hitting your chat UI, you'll want per-user API keys, usage tracking, and rate limits rather than sharing one Anthropic key across your whole app. SubToAPI issues separate sub_live_... keys per application or team seat, tracks token usage per key, and exposes it in a dashboard — useful if you're billing customers or just want to know which feature is burning tokens. Plans start at €9/month for Solo, with Team and Scale tiers for multi-seat setups — see /pricing. You can get a key and test the streaming endpoint from a free trial at /signup.
FAQ
Does Claude's API support streaming responses? Yes. Set "stream": true in your Messages API request and the response arrives as server-sent events with content_block_delta chunks containing incremental text, terminated by a message_stop event.
Can I stream Claude responses directly from the browser without a backend? No — doing so exposes your API key. You need a server-side proxy (Node.js, Next.js API route, or a hosted gateway like SubToAPI) that holds the key and forwards the SSE stream to the client.
Why does my React UI lag or freeze during streaming? Usually because you're re-rendering the entire message list on every token. Update only the last message object by index, and avoid re-mounting the message list component on each delta — use stable key props and memoize static messages.