Claude API Usage Alerts: Slack Integration Guide
Anthropic's console shows you usage graphs, but it won't ping your team when spend spikes overnight or a deploy starts throwing 429s. If you want Claude API usage alerts in Slack, you need to build a small pipeline yourself: pull usage data on a schedule, check it against thresholds, and post to a Slack webhook when something crosses the line.
This is a common gap for teams running Claude in production. Nobody notices a runaway loop calling the API 10x more than expected until the invoice arrives, or worse, until a customer-facing feature starts failing because a key got rate-limited. The fix isn't complicated — it's a scheduled job plus a webhook — but there are a few decisions worth getting right: what to alert on, how often to check, and where the usage numbers come from in the first place.
What to alert on
Not every usage metric deserves a Slack ping. Pick a small set of signals that actually require action:
- Daily/monthly spend threshold — e.g., "alert at 80% of monthly budget"
- Request rate approaching your plan's limit — catches noisy loops before they trigger throttling
- Error rate spike — a jump in 4xx/5xx responses usually means a bad deploy or an expired key
- Sudden drop to zero requests — often more telling than a spike; it can mean an integration silently broke
Alerting on raw token counts per request is usually noise. Alert on trends and thresholds, not individual data points.
Step 1: Get a Slack incoming webhook
Create a Slack app (or use an existing one) and add an Incoming Webhook for the channel you want alerts in. Slack gives you a URL like:
https://hooks.slack.com/services/T000/B000/XXXXXXXXXXXX
Posting to it is a plain HTTP POST with a JSON body:
curl -X POST "$SLACK_WEBHOOK_URL" \
-H "Content-Type: application/json" \
-d '{"text": "⚠️ Claude API spend hit 80% of monthly budget ($720/$900)."}'
Keep the webhook URL in a secret manager or environment variable, not in source control.
Step 2: Get usage data on a schedule
You need a source of truth for usage. If you're calling the Anthropic API directly, you're responsible for logging every request's tokens and cost yourself, since the raw API doesn't expose a usage-query endpoint you can poll. If you're routing traffic through SubToAPI, usage metadata — tokens, cost, per-key breakdown — is already tracked per application key, so you can pull it straight from your account rather than building a logging layer from scratch.
A simple cron job (or scheduled Lambda/Cloud Function) checks usage every 15–30 minutes:
import fetch from "node-fetch";
const BUDGET_USD = 900;
const WARN_THRESHOLD = 0.8;
async function checkUsage() {
// pull today's/this month's aggregated usage from your logging
// or dashboard/export, then compute spend so far
const spendSoFar = await getMonthToDateSpend();
if (spendSoFar / BUDGET_USD >= WARN_THRESHOLD) {
await postToSlack(
`⚠️ Claude API spend at ${(spendSoFar / BUDGET_USD * 100).toFixed(0)}% ` +
`of monthly budget ($${spendSoFar.toFixed(2)} / $${BUDGET_USD}).`
);
}
}
async function postToSlack(text) {
await fetch(process.env.SLACK_WEBHOOK_URL, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ text }),
});
}
checkUsage();
Run this with cron, a GitHub Actions scheduled workflow, or your platform's job scheduler. The key requirement is that getMonthToDateSpend() returns a real number — how you implement it depends on where your usage data lives.
Step 3: Add debounce logic
Without debouncing, you'll get a Slack message every 15 minutes once you cross a threshold, which trains people to ignore the channel. Track the last alert state (in a database row, a small JSON file, or a cache key) and only post once per threshold crossed per day:
const alertedToday = await getAlertFlag("budget_80pct");
if (spendSoFar / BUDGET_USD >= WARN_THRESHOLD && !alertedToday) {
await postToSlack(/* ... */);
await setAlertFlag("budget_80pct", true, { ttl: "24h" });
}
Reset the flag daily so each new day gets a fresh check.
Step 4: Alert on rate limits and errors, not just spend
Cost alerts get the most attention, but rate-limit and error alerts prevent outages. If you're logging response status codes from your Claude calls, add a second check:
async function checkErrorRate() {
const { total, errors } = await getLastHourStats();
const errorRate = errors / total;
if (errorRate > 0.05 && total > 20) {
await postToSlack(
`🚨 Claude API error rate at ${(errorRate * 100).toFixed(1)}% ` +
`over the last hour (${errors}/${total} requests failed).`
);
}
}
The total > 20 guard avoids false alarms from small sample sizes during low-traffic periods.
Where SubToAPI fits
If you're managing this across multiple applications or team members, the bookkeeping part — knowing which key made which request, at what cost, with what latency — is usually the hardest piece to build reliably. SubToAPI exposes that as structured usage metadata per application key, with per-seat visibility across your team, so your Slack alert script has a clean data source to query instead of parsing raw logs. It doesn't post to Slack natively, but pulling numbers into the kind of script above takes a few lines once the underlying usage tracking is already solved. Check the docs for the request/response shape, or start with the quickstart if you're setting up API access for the first time. Plans and seat pricing are on the pricing page.
Keep it simple
Resist the urge to build a full alerting platform for this. A scheduled script, a Slack webhook, and two or three well-chosen thresholds will catch the incidents that actually matter — runaway costs, approaching rate limits, and elevated error rates — without turning your Slack channel into noise.
FAQs
Does Anthropic's console support Slack alerts natively? No. The console shows usage dashboards but has no built-in alerting or Slack integration. You need to poll usage data yourself and push it to a webhook.
What's the easiest way to post to Slack from a script? Create a Slack incoming webhook URL for your channel, then send a JSON POST request with a text field. No SDK is required for basic alerts.
How often should I check usage for alerts? Every 15–30 minutes is enough for cost and rate-limit alerts. Error-rate alerts benefit from tighter windows (5–10 minutes) since they signal active incidents.