← Blog

Claude API Next.js Chatbot Integration Guide

2026-10-02 · 5 min read · SubToAPI Team

Building a chatbot with the Claude API in Next.js comes down to three pieces: a server-side route handler that calls Claude and keeps your API key off the client, a streaming response so the UI feels responsive instead of waiting for a full reply, and a small amount of client state to render the conversation. This guide walks through all three using the App Router, with a working example you can copy into a project today.

If you've searched for this topic, you're probably past the "what is Claude" stage and want the actual wiring: where the fetch call lives, how to stream tokens into React, and how to handle conversation history across turns. That's what's below.

Project setup

You need a Next.js 14+ project with the App Router. Create a route handler for the chat endpoint:

app/
  api/
    chat/
      route.ts
  page.tsx

Install nothing extra — the native fetch API and ReadableStream handle everything, no SDK required.

The server route handler

The route handler is the only place your API key should ever appear. It receives messages from the client, forwards them to Claude, and streams the response back.

// app/api/chat/route.ts
export const runtime = "edge";

export async function POST(req) {
  const { messages } = await req.json();

  const response = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      stream: true,
      messages,
    }),
  });

  return new Response(response.body, {
    headers: { "Content-Type": "text/event-stream" },
  });
}

Running on the Edge runtime keeps latency low since the route just proxies a stream — it doesn't buffer the full response in memory.

If you'd rather not manage raw Anthropic auth headers, streaming event parsing, and key rotation yourself, SubToAPI exposes the same /v1/messages shape behind a sub_live_... key, so this same route handler works unchanged — just swap the host and header. See /docs/quickstart for the setup and /docs/streaming for streaming-specific details.

Parsing the stream on the client

Claude's streaming format sends server-sent events with content_block_delta chunks. On the client, read the stream manually and append text as it arrives:

// app/page.tsx
"use client";
import { useState } from "react";

export default function Chat() {
  const [messages, setMessages] = useState([]);
  const [input, setInput] = useState("");

  async function sendMessage() {
    const next = [...messages, { role: "user", content: input }];
    setMessages(next);
    setInput("");

    const res = await fetch("/api/chat", {
      method: "POST",
      body: JSON.stringify({ messages: next }),
    });

    const reader = res.body.getReader();
    const decoder = new TextDecoder();
    let assistantText = "";
    setMessages([...next, { role: "assistant", content: "" }]);

    while (true) {
      const { done, value } = await reader.read();
      if (done) break;
      const chunk = decoder.decode(value);

      for (const line of chunk.split("\n")) {
        if (!line.startsWith("data:")) continue;
        try {
          const event = JSON.parse(line.slice(5));
          if (event.type === "content_block_delta") {
            assistantText += event.delta.text;
            setMessages((prev) => [
              ...prev.slice(0, -1),
              { role: "assistant", content: assistantText },
            ]);
          }
        } catch {}
      }
    }
  }

  return (
    <div>
      {messages.map((m, i) => (
        <p key={i}><strong>{m.role}:</strong> {m.content}</p>
      ))}
      <input value={input} onChange={(e) => setInput(e.target.value)} />
      <button onClick={sendMessage}>Send</button>
    </div>
  );
}

This is the minimum viable chatbot: user types, message is pushed to state, the server streams Claude's reply, and the UI updates token by token.

Handling conversation history

Claude is stateless between requests — every call to /v1/messages needs the full message array, not just the latest turn. The next array in the example above already does this by appending to messages before sending. For longer conversations, you'll eventually want to trim history to stay under the model's context window, either by truncating old turns or summarizing them before the next request.

Tool use and structured responses

If your chatbot needs to call functions — looking up order status, querying a database, hitting an internal API — Claude's tool use feature lets the model request a function call instead of returning plain text. Your route handler checks for a tool_use content block, executes the function server-side, and sends the result back as a tool_result in the next message. This fits naturally into the same route handler shown above since it's just another branch in how you process the response before streaming it to the client. SubToAPI supports the same tool-calling flow if you're using it as your backend — see /docs/tools for the request format.

Error handling and rate limits

Production chatbots need to handle three failure modes: network errors mid-stream, rate limit responses (HTTP 429), and malformed JSON in SSE chunks (wrap the JSON.parse call in try/catch, as shown above, since partial chunks can arrive mid-object). For rate limits specifically, return a clear message to the user rather than letting the stream silently die — check response.status before piping the body.

Why some teams front Claude with SubToAPI

Two recurring pain points when shipping a Next.js chatbot to production: managing Anthropic API keys per environment (dev, staging, prod, and per-developer), and getting usage visibility across a team without building a dashboard yourself. SubToAPI wraps your existing Claude access in a standard HTTPS API with sub_live_... application keys, so each environment or team member gets its own key, usage is visible in one dashboard, and the request/response shape matches /v1/messages closely enough that the route handler above needs only a host and header swap. Plans start at Solo €9, with Team (€19/seat) and Scale (€49/seat) tiers for larger teams, and there's a free trial at /signup. Full request details are in /docs/messages and /pricing has the plan breakdown.

Questions

Do I need a Node.js backend, or can I call Claude directly from the browser? Call it from a server. Browser requests would expose your API key in client-side JavaScript, which anyone can read from devtools. A Next.js route handler keeps the key server-side while still giving you a streaming response.

Why does streaming matter for a chatbot UI? Without streaming, the browser waits for the entire response before showing anything, which feels slow for longer answers. Streaming renders tokens as they arrive, matching the perceived responsiveness users expect from chat interfaces.

Can I use the Edge runtime for the chat route? Yes, and it's recommended. Edge functions have lower cold-start latency and handle streaming responses well since they don't need to buffer the full body before forwarding it, which keeps time-to-first-token low.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →