← Blog

Monitor Claude API Spend in Real Time: A Practical Guide

2026-09-30 · 5 min read · SubToAPI Team

If you're asking how to monitor Claude API spend in real time, the short answer is: you need per-request usage metadata (input/output tokens, cost), a place to aggregate it instantly, and a way to see it broken down by key, project, or team member — not a monthly invoice you check after the damage is done. Anthropic's console shows usage, but it's not built for live, per-key tracking across a team, and it doesn't push alerts when a key starts burning through budget.

This matters because Claude spend is unpredictable by nature. A single runaway loop, an unbounded max_tokens, or a teammate testing a long-context prompt in production can turn a €50 day into a €500 day before anyone notices. Real-time visibility is the difference between catching that in minutes versus finding out at the end of the billing cycle.

Why Claude spend is hard to track by default

The raw Claude API gives you token counts in each response, but turning that into "who spent what, right now" requires building infrastructure:

Most teams start by logging to a database and writing a dashboard. That works, but it's ongoing maintenance — pricing changes, new models get added, and someone has to keep the cost math correct.

What real-time monitoring actually needs

1. Per-key or per-user API keys

You can't monitor spend "in real time" if every request uses the same shared API key. The first step is issuing separate keys per application, environment, or team member so usage can be attributed. If you're on the raw Anthropic API, this means managing your own key-issuing layer, since Anthropic doesn't give you unlimited scoped sub-keys out of the box.

2. Usage metadata on every response

Every Claude response includes token usage. You need to capture input_tokens and output_tokens (and cache-read/cache-write tokens if you're using prompt caching) on every call, then multiply by current per-model pricing to get a cost figure.

{
  "usage": {
    "input_tokens": 512,
    "output_tokens": 128,
    "cache_read_input_tokens": 0
  }
}

3. A live aggregation layer

Raw logs aren't monitoring. You need something that sums spend by key/day/hour as requests come in, so a dashboard or alert can react within seconds, not after a batch job runs overnight.

4. Thresholds and alerts

Real-time monitoring is only useful if it triggers action. Set a daily or monthly cap per key, and get notified (or have requests blocked) when it's hit.

Building it yourself vs. using a dashboard that already tracks it

If you're calling the Claude API directly, you can build this with a middleware layer that wraps every request, logs usage, and writes to a time-series-friendly store (Postgres with hourly rollups works fine at moderate volume). It's a reasonable project if you have the time and it's core to your product.

If it's not core to your product, it's usually faster to use a layer that already does this. SubToAPI sits between your app and Claude and gives every application its own sub_live_... key, with usage metadata returned on each response and tracked per key in a dashboard — so you can see spend accumulate in near real time without building the logging pipeline yourself.

A typical setup looks like this:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [
      {"role": "user", "content": "Summarize this changelog."}
    ]
  }'

Each response carries usage data you can log on your own side too, but because spend is already attributed per key in the dashboard, you get visibility across your whole team — Solo, Team, or Scale plans — without writing aggregation code. Full request/response shape is in the docs and messages reference.

Practical habits that reduce surprise spend

Monitoring in real time is most valuable when paired with a few defensive habits:

Setting this up in a few minutes

If you want live per-key spend tracking without building the pipeline yourself:

  1. Sign up and start a free trial.
  2. Create separate application keys for each project or environment from the dashboard.
  3. Swap your Anthropic base URL for the SubToAPI endpoint and point requests at your sub_live_... key — see the quickstart.
  4. Watch spend accumulate per key in the dashboard as requests come in.
  5. Compare plans on pricing as your team grows — Solo for individuals, Team and Scale for per-seat usage across multiple builders.

FAQ

Does the Claude API report cost directly, or just tokens?

Anthropic's API returns token counts (input_tokens, output_tokens, and cache-related fields), not a dollar figure. You calculate cost by multiplying token counts by the current per-model rate, which means your monitoring layer needs to stay updated when pricing changes.

Can I monitor spend per team member, not just per project?

Yes, if each team member or application has its own API key. Shared keys make per-person attribution impossible — the fix is issuing distinct keys and aggregating usage by key, which is what a per-key dashboard like SubToAPI's is built for.

What's the fastest way to catch a runaway cost spike?

Set a hard daily cap per key and alert (or block requests) when it's hit, rather than relying on end-of-month invoice review. Combined with explicit max_tokens limits on every request, this catches most spikes within hours instead of weeks.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →