← Blog

Claude API Discord Bot Development Tutorial

2026-10-03 · 5 min read · SubToAPI Team

Building a Discord bot that talks like Claude means wiring three pieces together: a Discord client that listens for messages, an HTTP call to Claude's Messages API, and some logic to turn a chat channel into a usable conversation. None of this is exotic — it's a standard webhook-style integration — but there are a handful of details (rate limits, message length caps, streaming edits, conversation memory) that trip people up the first time.

This tutorial walks through the full setup using Node.js and discord.js, from registering the bot application to deploying a working /ask command that calls Claude and replies in-channel. The same pattern works whether you call Claude directly through Anthropic's API or through a managed gateway like SubToAPI — we'll use the latter in the code examples since it simplifies key management and gives you usage metadata for free, but the Discord-side code is identical either way.

What You Need Before Starting

Step 1: Register the Discord Application

  1. Go to the Discord Developer Portal and create a new application.
  2. Under Bot, click "Add Bot" and copy the token — treat it like a password.
  3. Under OAuth2 → URL Generator, select the bot scope and permissions: Send Messages, Read Message History, Use Slash Commands.
  4. Open the generated URL and invite the bot to your test server.

Step 2: Set Up the Project

mkdir claude-discord-bot && cd claude-discord-bot
npm init -y
npm install discord.js dotenv

Create a .env file:

DISCORD_TOKEN=your-discord-bot-token
SUBTOAPI_KEY=sub_live_xxxxxxxxxxxx

Step 3: Register a Slash Command

Discord bots should use slash commands rather than parsing every message in a channel — it's cleaner and avoids accidentally responding to unrelated chatter.

// register-commands.js
import { REST, Routes, SlashCommandBuilder } from 'discord.js';
import 'dotenv/config';

const command = new SlashCommandBuilder()
  .setName('ask')
  .setDescription('Ask Claude a question')
  .addStringOption(opt =>
    opt.setName('prompt').setDescription('Your question').setRequired(true));

const rest = new REST({ version: '10' }).setToken(process.env.DISCORD_TOKEN);
await rest.put(Routes.applicationCommands('YOUR_APP_ID'), {
  body: [command.toJSON()],
});
console.log('Command registered');

Run this once with node register-commands.js.

Step 4: Call Claude from the Bot

This is the core integration. When the slash command fires, send the prompt to Claude's Messages endpoint and reply with the result.

// bot.js
import { Client, GatewayIntentBits, Events } from 'discord.js';
import 'dotenv/config';

const client = new Client({ intents: [GatewayIntentBits.Guilds] });

client.on(Events.InteractionCreate, async (interaction) => {
  if (!interaction.isChatInputCommand() || interaction.commandName !== 'ask') return;

  const prompt = interaction.options.getString('prompt');
  await interaction.deferReply(); // Discord gives you ~15 min, Claude is usually fast enough without this

  try {
    const res = await fetch('https://api.subtoapi.app/v1/messages', {
      method: 'POST',
      headers: {
        'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
        'Content-Type': 'application/json',
      },
      body: JSON.stringify({
        model: 'claude-sonnet-4-5',
        max_tokens: 1024,
        messages: [{ role: 'user', content: prompt }],
      }),
    });

    const data = await res.json();
    const text = data.content?.[0]?.text ?? 'No response generated.';

    // Discord caps messages at 2000 characters
    await interaction.editReply(text.slice(0, 1990));
  } catch (err) {
    console.error(err);
    await interaction.editReply('Something went wrong talking to Claude.');
  }
});

client.login(process.env.DISCORD_TOKEN);

Run it with node bot.js, then type /ask in your server. The full request/response shape for the Messages endpoint is documented at /docs/messages if you want to add system prompts, multi-turn history, or adjust temperature.

Step 5: Handle Long Responses with Streaming

Discord's 2000-character limit means long Claude responses get truncated unless you either split them into multiple messages or stream and edit a single message as tokens arrive. Streaming also makes the bot feel responsive instead of waiting 5–10 seconds in silence.

The pattern: open a streaming connection, buffer tokens, and edit the Discord message every few hundred milliseconds instead of on every token (Discord rate-limits message edits aggressively — roughly 5 edits per message per 5 seconds). See /docs/streaming for the event format; the client-side logic is the same server-sent-events parsing you'd use in any JS app, just feeding a setInterval-based edit instead of a DOM update.

Step 6: Add Conversation Memory

A single-turn bot is fine for quick lookups, but most Discord use cases (support bots, roleplay bots, coding assistants) need context across messages. Keep it simple at first: store the last N messages per channel or thread in memory (a Map keyed by channel ID) and include them in the messages array on each call. For anything beyond a toy project, move that history to Redis or a database so it survives bot restarts.

Step 7: Keep an Eye on Keys and Rate Limits

A Discord bot that gets popular can generate a lot of concurrent Claude calls fast — every /ask in every server is a separate request. Two things matter here:

Step 8: Deploy It

For a hobby bot, a small VPS or a free-tier container service running node bot.js with a process manager (pm2 or systemd) is enough. For production bots with thousands of servers, you'll want horizontal scaling behind a sharding manager — discord.js has built-in support for this once you cross Discord's sharding threshold (around 2,500 guilds).

questions

Do I need an Anthropic API key, or can I use a third-party service? Either works — the Discord-side code is identical. A gateway like SubToAPI gives you an application-style key (sub_live_...), usage metadata, and team seats on top of the same Claude models; see /docs/quickstart for setup.

Why does my bot's response get cut off in Discord? Discord messages are capped at 2000 characters. Truncate, split into multiple messages, or use streaming with periodic edits so you can send a follow-up message when the limit is reached.

How do I prevent users from spamming my bot and running up API costs? Add per-user or per-channel cooldowns (a simple timestamp Map works), cap max_tokens per request, and monitor usage through your API provider's dashboard so you catch spikes early.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →