← Blog

Claude API Discord Bot: Full Build Guide

2026-10-05 · 5 min read · SubToAPI Team

Building a Discord bot that uses Claude means wiring three pieces together: a Discord gateway client that listens for messages, an HTTP call to Claude's messages endpoint, and some logic to keep conversation context per channel or user. This guide walks through all three, using Node.js and discord.js, and covers the parts that trip people up: rate limits, streaming into Discord's edit API, and managing per-user memory without a database migration.

If you just want the shortest path from "nothing" to "bot replies with Claude," skip to the Quick Start below. If you need production details — retries, concurrency, cost control — read the whole thing.

What you need before starting

Quick start: minimal bot

Install dependencies:

npm install discord.js node-fetch

Core bot file:

import { Client, GatewayIntentBits } from "discord.js";
import fetch from "node-fetch";

const client = new Client({
  intents: [
    GatewayIntentBits.Guilds,
    GatewayIntentBits.GuildMessages,
    GatewayIntentBits.MessageContent,
  ],
});

client.on("messageCreate", async (message) => {
  if (message.author.bot) return;
  if (!message.mentions.has(client.user)) return;

  const prompt = message.content.replace(/<@!?\d+>/, "").trim();
  const reply = await askClaude(prompt);
  await message.reply(reply.slice(0, 2000)); // Discord's 2000-char limit
});

async function askClaude(userMessage) {
  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      messages: [{ role: "user", content: userMessage }],
    }),
  });
  const data = await res.json();
  return data.content[0].text;
}

client.login(process.env.DISCORD_TOKEN);

That's a working bot: mention it in a channel, it responds with Claude's answer. Everything after this is making it production-grade.

Handling Discord's 2000-character limit

Claude responses often exceed Discord's message cap. Split long replies into chunks instead of truncating:

function chunkMessage(text, limit = 2000) {
  const chunks = [];
  for (let i = 0; i < text.length; i += limit) {
    chunks.push(text.slice(i, i + limit));
  }
  return chunks;
}

for (const chunk of chunkMessage(reply)) {
  await message.reply(chunk);
}

For a nicer UX, send the first chunk as a reply and the rest as follow-up messages in the same channel, not as replies to avoid spamming the reply chain.

Streaming responses into Discord

Discord has no native token-streaming UI, but you can simulate it by editing a message as chunks arrive. This gives users a "typing" feel instead of a long silent wait:

const placeholder = await message.reply("Thinking...");
let buffer = "";

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4-5",
    max_tokens: 1024,
    stream: true,
    messages: [{ role: "user", content: userMessage }],
  }),
});

let lastEdit = Date.now();
for await (const chunk of res.body) {
  const lines = chunk.toString().split("\n").filter((l) => l.startsWith("data:"));
  for (const line of lines) {
    const payload = JSON.parse(line.replace("data: ", ""));
    if (payload.delta?.text) buffer += payload.delta.text;
  }
  if (Date.now() - lastEdit > 800) {
    await placeholder.edit(buffer.slice(0, 2000));
    lastEdit = Date.now();
  }
}
await placeholder.edit(buffer.slice(0, 2000));

Throttle edits to roughly once per second — Discord rate-limits message edits per channel, and editing on every token will get your bot throttled fast. See /docs/streaming for the full event format if you're parsing server-sent events directly.

Managing conversation context

Discord bots usually need to remember recent messages per channel or thread. A simple in-memory map works for low-traffic bots:

const histories = new Map(); // channelId -> messages[]

function getHistory(channelId) {
  if (!histories.has(channelId)) histories.set(channelId, []);
  return histories.get(channelId);
}

function pushMessage(channelId, role, content) {
  const history = getHistory(channelId);
  history.push({ role, content });
  if (history.length > 20) history.shift(); // cap context size
}

For anything that needs to survive restarts or scale across bot shards, move this to Redis or a small Postgres table keyed by channel ID, storing just role/content pairs.

Rate limits and concurrency

Discord bots can receive bursts of messages across many channels simultaneously. Each one triggering a Claude call means you need to think about concurrency:

If you're running the bot for a community or team and want visibility into who's driving usage without building your own logging, routing calls through SubToAPI gives you per-key usage metadata and a dashboard, which is useful once more than one person maintains the bot. See /docs/quickstart for key setup and /pricing for plan details — Solo works fine for a single bot, Team adds seats if multiple people manage it.

Error handling basics

Wrap the Claude call in error handling so the bot doesn't go silent on failures:

try {
  const reply = await askClaude(prompt);
  await message.reply(reply.slice(0, 2000));
} catch (err) {
  console.error(err);
  await message.reply("Something went wrong talking to Claude. Try again shortly.");
}

Log the full error server-side, but keep the user-facing message generic — Discord users don't need to see raw status codes.

Deploying the bot

Run it as a long-lived Node process (PM2, a Docker container, or a small VM/Fly.io app). Discord bots use a persistent WebSocket connection, so serverless functions with cold starts aren't a good fit unless you split the gateway listener from the Claude-calling logic via a queue.

questions

Do I need a paid Claude plan to run a Discord bot? You need API access with billing attached, either directly from Anthropic or through a service like SubToAPI that turns your Claude access into an API key with a free trial at /signup.

Can the bot handle multiple servers at once? Yes — discord.js handles multiple guilds in one process by default. Just make sure your context storage is keyed by channel or guild ID, not global, so conversations don't bleed between servers.

How do I stop the bot from responding to every message? Gate replies behind a mention, a slash command, or a specific channel using message.mentions.has(client.user) or Discord's slash command interactions instead of listening to all messageCreate events.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →