Claude API Slack Bot Integration: A Developer Guide
Building a Slack bot powered by Claude means connecting Slack's Events API to a backend that calls Claude, then posting the response back into the channel or thread. The architecture is straightforward: Slack sends you an event webhook when someone mentions your bot or DMs it, your server extracts the message text and thread history, sends it to Claude, and replies via Slack's chat.postMessage API. The hard parts are authentication, thread context, rate limits, and keeping the bot responsive inside Slack's tight timeout windows.
This guide covers the practical pieces: setting up the Slack app, handling events correctly, managing conversation context across a thread, and dealing with streaming and latency so your bot doesn't feel sluggish.
Setting up the Slack app
You need a Slack app with:
- Event Subscriptions enabled, subscribed to
app_mentionandmessage.im(for DMs) - A bot token (
xoxb-...) withchat:write,app_mentions:read, andim:historyscopes - A public HTTPS endpoint Slack can POST events to (ngrok or a real server)
Slack requires your endpoint to respond within 3 seconds, or it retries the event — which, if your handler isn't idempotent, means duplicate Claude calls and duplicate replies. Acknowledge immediately and process asynchronously:
app.post("/slack/events", (req, res) => {
if (req.body.type === "url_verification") {
return res.send(req.body.challenge);
}
res.sendStatus(200); // ack immediately
handleEvent(req.body.event).catch(console.error);
});
Do the actual Claude call and Slack reply inside handleEvent, outside the request/response cycle.
Calling Claude from the event handler
Once you have the message text, call Claude the same way you would from any backend service. If you're using SubToAPI, the request looks like a standard HTTPS call with your sub_live_... key:
async function handleEvent(event) {
if (event.bot_id) return; // ignore bot's own messages
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
messages: [{ role: "user", content: event.text }]
})
});
const data = await response.json();
const reply = data.content[0].text;
await slackClient.chat.postMessage({
channel: event.channel,
thread_ts: event.thread_ts || event.ts,
text: reply
});
}
Always reply in-thread (thread_ts) rather than posting a new top-level message — it keeps channels readable and lets you reconstruct conversation history later.
Handling thread context
Slack conversations aren't single-turn. If a user replies inside a thread, Claude needs the prior messages to respond coherently. Use conversations.replies to pull the thread, then map each Slack message into Claude's messages array:
async function getThreadMessages(channel, threadTs) {
const result = await slackClient.conversations.replies({
channel,
ts: threadTs
});
return result.messages.map(m => ({
role: m.bot_id ? "assistant" : "user",
content: m.text
}));
}
Strip the bot's own @mention prefix from user messages before sending them to Claude — otherwise every turn includes <@U12345> noise. A simple regex replace on event.text handles this:
const cleanText = event.text.replace(/<@[A-Z0-9]+>/g, "").trim();
Cap thread length before sending to Claude. Long-running threads can exceed reasonable token budgets fast; truncate to the last 15–20 messages or summarize older turns if the thread is a persistent support channel.
Streaming into Slack
Slack doesn't support token-by-token streaming the way a chat UI does, but you can simulate progressive updates by editing the message as tokens arrive, which makes long responses feel faster instead of waiting 10+ seconds for a full answer:
const msg = await slackClient.chat.postMessage({
channel: event.channel,
thread_ts: event.ts,
text: "Thinking..."
});
let buffer = "";
const stream = await getStreamingResponse(cleanText);
for await (const chunk of stream) {
buffer += chunk;
if (buffer.length % 200 < 20) {
await slackClient.chat.update({
channel: event.channel,
ts: msg.ts,
text: buffer
});
}
}
Throttle updates — Slack rate-limits chat.update calls per channel, so updating on every token will get you rate limited quickly. Batching every ~200 characters is a reasonable middle ground. See /docs/streaming for details on consuming streamed responses from the API.
Multi-channel and multi-workspace considerations
If your bot runs across several Slack workspaces or your team has multiple bots (support, internal docs, code review), it's worth separating API keys per bot so you can see usage and cost per integration independently rather than one shared key with no breakdown. This also limits blast radius if one bot's key leaks or starts misbehaving — you can revoke a single key without taking down every integration.
SubToAPI issues separate sub_live_... application keys under one account, with per-key usage visible in the dashboard, which maps cleanly onto a "one key per bot" pattern. Signup includes a free trial, and setup is a single API call rather than provisioning separate Anthropic console access for each integration — see /docs/quickstart.
Error handling and fallbacks
Slack bots fail visibly — a stuck "Thinking..." message or a thrown error in a public channel looks bad. Always wrap the Claude call in a try/catch and post a fallback message on failure:
try {
// Claude call + Slack reply
} catch (err) {
await slackClient.chat.postMessage({
channel: event.channel,
thread_ts: event.ts,
text: "Something went wrong processing that — try again in a moment."
});
}
Log the actual error server-side, not in Slack. Also set a reasonable timeout on your Claude call (10–15 seconds) so a slow response doesn't hang the handler indefinitely.
questions
Do I need a separate Anthropic account to build a Slack bot with Claude? No — you can use SubToAPI to turn existing Claude access into an HTTPS API with an application key, which works the same way for a Slack bot's backend calls as any direct integration. See /docs/quickstart.
How do I avoid duplicate replies when Slack retries an event? Acknowledge the event immediately with a 200 response before processing, and track event IDs you've already handled (even a short-lived in-memory cache works) to skip retries.
Can the bot maintain context across a whole Slack thread? Yes — pull the thread with conversations.replies, map each message to a Claude role, and send the full (or truncated) history as the messages array on each request. See /docs/messages for the message format.