Claude API Discord Bot: Full Build Guide
Building a Discord bot that uses Claude means wiring three pieces together: a Discord gateway client that listens for messages, an HTTP call to Claude's messages endpoint, and some logic to keep conversation context per channel or user. This guide walks through all three, using Node.js and discord.js, and covers the parts that trip people up: rate limits, streaming into Discord's edit API, and managing per-user memory without a database migration.
If you just want the shortest path from "nothing" to "bot replies with Claude," skip to the Quick Start below. If you need production details — retries, concurrency, cost control — read the whole thing.
What you need before starting
- A Discord application and bot token from the Discord Developer Portal
- A Claude API key (directly from Anthropic, or a proxied key like a
sub_live_...key from a service such as SubToAPI if you want unified billing and a dashboard across your team's bots) - Node.js 18+ and
discord.jsv14 MESSAGE CONTENTprivileged intent enabled in the Discord portal
Quick start: minimal bot
Install dependencies:
npm install discord.js node-fetch
Core bot file:
import { Client, GatewayIntentBits } from "discord.js";
import fetch from "node-fetch";
const client = new Client({
intents: [
GatewayIntentBits.Guilds,
GatewayIntentBits.GuildMessages,
GatewayIntentBits.MessageContent,
],
});
client.on("messageCreate", async (message) => {
if (message.author.bot) return;
if (!message.mentions.has(client.user)) return;
const prompt = message.content.replace(/<@!?\d+>/, "").trim();
const reply = await askClaude(prompt);
await message.reply(reply.slice(0, 2000)); // Discord's 2000-char limit
});
async function askClaude(userMessage) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{ role: "user", content: userMessage }],
}),
});
const data = await res.json();
return data.content[0].text;
}
client.login(process.env.DISCORD_TOKEN);
That's a working bot: mention it in a channel, it responds with Claude's answer. Everything after this is making it production-grade.
Handling Discord's 2000-character limit
Claude responses often exceed Discord's message cap. Split long replies into chunks instead of truncating:
function chunkMessage(text, limit = 2000) {
const chunks = [];
for (let i = 0; i < text.length; i += limit) {
chunks.push(text.slice(i, i + limit));
}
return chunks;
}
for (const chunk of chunkMessage(reply)) {
await message.reply(chunk);
}
For a nicer UX, send the first chunk as a reply and the rest as follow-up messages in the same channel, not as replies to avoid spamming the reply chain.
Streaming responses into Discord
Discord has no native token-streaming UI, but you can simulate it by editing a message as chunks arrive. This gives users a "typing" feel instead of a long silent wait:
const placeholder = await message.reply("Thinking...");
let buffer = "";
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
stream: true,
messages: [{ role: "user", content: userMessage }],
}),
});
let lastEdit = Date.now();
for await (const chunk of res.body) {
const lines = chunk.toString().split("\n").filter((l) => l.startsWith("data:"));
for (const line of lines) {
const payload = JSON.parse(line.replace("data: ", ""));
if (payload.delta?.text) buffer += payload.delta.text;
}
if (Date.now() - lastEdit > 800) {
await placeholder.edit(buffer.slice(0, 2000));
lastEdit = Date.now();
}
}
await placeholder.edit(buffer.slice(0, 2000));
Throttle edits to roughly once per second — Discord rate-limits message edits per channel, and editing on every token will get your bot throttled fast. See /docs/streaming for the full event format if you're parsing server-sent events directly.
Managing conversation context
Discord bots usually need to remember recent messages per channel or thread. A simple in-memory map works for low-traffic bots:
const histories = new Map(); // channelId -> messages[]
function getHistory(channelId) {
if (!histories.has(channelId)) histories.set(channelId, []);
return histories.get(channelId);
}
function pushMessage(channelId, role, content) {
const history = getHistory(channelId);
history.push({ role, content });
if (history.length > 20) history.shift(); // cap context size
}
For anything that needs to survive restarts or scale across bot shards, move this to Redis or a small Postgres table keyed by channel ID, storing just role/content pairs.
Rate limits and concurrency
Discord bots can receive bursts of messages across many channels simultaneously. Each one triggering a Claude call means you need to think about concurrency:
- Queue requests per channel so a single channel's bursty activity doesn't flood the API
- Set a reasonable
max_tokens(512–1024 for chat-style replies) to keep latency and cost predictable - Use a single API key across your bot rather than per-user keys — track usage centrally instead
If you're running the bot for a community or team and want visibility into who's driving usage without building your own logging, routing calls through SubToAPI gives you per-key usage metadata and a dashboard, which is useful once more than one person maintains the bot. See /docs/quickstart for key setup and /pricing for plan details — Solo works fine for a single bot, Team adds seats if multiple people manage it.
Error handling basics
Wrap the Claude call in error handling so the bot doesn't go silent on failures:
try {
const reply = await askClaude(prompt);
await message.reply(reply.slice(0, 2000));
} catch (err) {
console.error(err);
await message.reply("Something went wrong talking to Claude. Try again shortly.");
}
Log the full error server-side, but keep the user-facing message generic — Discord users don't need to see raw status codes.
Deploying the bot
Run it as a long-lived Node process (PM2, a Docker container, or a small VM/Fly.io app). Discord bots use a persistent WebSocket connection, so serverless functions with cold starts aren't a good fit unless you split the gateway listener from the Claude-calling logic via a queue.
questions
Do I need a paid Claude plan to run a Discord bot? You need API access with billing attached, either directly from Anthropic or through a service like SubToAPI that turns your Claude access into an API key with a free trial at /signup.
Can the bot handle multiple servers at once? Yes — discord.js handles multiple guilds in one process by default. Just make sure your context storage is keyed by channel or guild ID, not global, so conversations don't bleed between servers.
How do I stop the bot from responding to every message? Gate replies behind a mention, a slash command, or a specific channel using message.mentions.has(client.user) or Discord's slash command interactions instead of listening to all messageCreate events.