Claude API Chat Widget: Embed It on Your Website
If you want to drop a chat widget onto your website that's powered by Claude, the short answer is: you never call Anthropic's API directly from the browser. You build a small backend endpoint that holds the API key, and your frontend widget talks to that endpoint instead. This guide walks through exactly how to wire that up, from the HTML snippet to the streaming response handling.
The reason this matters: Claude's API key is a secret. If you embed it in client-side JavaScript, anyone who opens dev tools can copy it and run up your bill. So every working "Claude chat widget" setup has the same shape — a lightweight server proxy sitting between your visitors and the model.
The architecture in one picture
Website visitor → Chat widget (JS) → Your backend endpoint → Claude API → streamed response back
Your widget is just HTML/CSS/JS that renders a chat bubble, an input box, and a message list. It posts user messages to your own /api/chat route. That route does the actual API call and streams tokens back down to the browser.
Step 1: Build the backend endpoint
A minimal Node/Express handler that forwards to Claude:
app.post("/api/chat", async (req, res) => {
const { message, history = [] } = req.body;
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [...history, { role: "user", content: message }],
}),
});
const data = await response.json();
res.json({ reply: data.content[0].text });
});
This works, but every reply waits for the full response before anything appears — not great for a widget where users expect a typing effect.
Step 2: Add streaming for a real chat feel
Streaming is what makes a widget feel alive instead of frozen while Claude thinks. You set stream: true and pipe server-sent events straight through to the browser:
app.post("/api/chat/stream", async (req, res) => {
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
stream: true,
messages: req.body.messages,
}),
});
response.body.pipe(res);
});
On the frontend, read the stream with EventSource or a fetch reader and append each text delta to the message bubble as it arrives.
Step 3: The widget markup and script
A barebones embeddable widget can be a single <script> tag that injects its own container, so clients can add it to any page with one line:
<script src="https://yourdomain.com/widget.js" data-position="bottom-right"></script>
Inside widget.js, create the DOM elements, attach a submit listener, and call your streaming endpoint:
async function sendMessage(text) {
const res = await fetch("/api/chat/stream", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: conversationHistory }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let botText = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
botText += decoder.decode(value);
updateBotBubble(botText);
}
}
Keep conversationHistory in memory (or localStorage for persistence across page loads) so Claude has context for follow-up questions.
Step 4: Handling the backend without maintaining infrastructure
Running your own proxy means you're responsible for key rotation, rate limiting, retry logic, and usage tracking across every site you embed the widget on. If you'd rather skip that layer, SubToAPI turns your existing Claude access into a standard HTTPS API with its own sub_live_... keys — you point your widget's backend at https://api.subtoapi.app/v1/messages instead of Anthropic directly, and you get per-application keys, streaming, and usage metadata in one dashboard without standing up your own proxy server. It's a drop-in swap for the fetch call above; the request and response shapes follow the same Messages format. Check the quickstart to see the full request cycle.
Step 5: Keep the widget fast and contained
A few practical things that make the difference between a demo and something you'd actually ship:
- Cap
max_tokenson widget replies (300–500 is usually enough for a support-style chat) so responses don't run long and the UI doesn't feel sluggish. - Trim history — send only the last 6–10 turns to the API, not the entire conversation, to control latency and cost.
- Add a system prompt scoped to your site's purpose (e.g., "You are a support assistant for Acme's docs") so the widget doesn't wander into unrelated topics.
- Rate-limit per visitor on your backend endpoint, independent of whatever limits your API key already has, to stop one bad actor from hammering the widget.
- Use streaming (docs/streaming) for anything the user will wait more than a second or two for — it changes the perceived speed dramatically even if total generation time is the same.
Should the widget call tools?
If your widget needs to look up order status, search docs, or hit an internal API, use tool use instead of trying to parse free text. Define a tool schema, let Claude decide when to call it, execute the function server-side, and feed the result back in a follow-up turn. This keeps the widget's answers grounded in real data instead of guesses.
questions
Can I embed Claude directly in client-side JavaScript without a backend? No. The API key must stay server-side. Every production chat widget uses a backend endpoint as a proxy between the browser and the model.
Do I need streaming for a chat widget, or is a single response enough? Streaming isn't required but strongly recommended — without it, users stare at a blank bubble while the full reply generates, which feels slow even on fast models.
What's the easiest way to add API keys and usage tracking per client site? Issue a separate key per deployment. SubToAPI gives each application its own sub_live_... key with usage metadata, so you can see which embedded widget is generating traffic without building that tracking yourself.