← Blog

Claude API Token Usage Tracking Script (How-To)

2026-10-10 · 5 min read · SubToAPI Team

If you're calling the Claude API directly, you're responsible for tracking how many tokens each request burns through — Anthropic doesn't give you a usage dashboard by default, just a usage object on every response. A token usage tracking script is simply code that reads that object, logs it somewhere durable (a file, a database, a logging service), and aggregates it by whatever dimension you care about: user, endpoint, model, or day.

This matters for three practical reasons: you need to catch cost spikes before the invoice arrives, you need per-user or per-feature attribution if you're billing customers or allocating budget internally, and you need historical data to decide when to switch models or optimize prompts. Below is a working approach you can adapt, plus a note on when writing and maintaining this yourself stops being worth it.

What Claude's API gives you to work with

Every response from the Messages API includes a usage field:

{
  "id": "msg_01...",
  "model": "claude-opus-4",
  "usage": {
    "input_tokens": 512,
    "output_tokens": 189,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

That's the entire source of truth. There's no separate usage endpoint to poll — you have to capture this object at the moment of the call and persist it yourself. If you're streaming responses, the final message_stop event in the SSE stream carries the same usage data, so your tracking logic needs to handle both response shapes.

A minimal tracking script

The simplest version wraps your API call, extracts usage, and appends a row to a log file or table. Here's a Node.js example that tracks usage per call and tags it with a user ID:

import fs from "fs";

async function callClaudeTracked(userId, messages, model = "claude-opus-4") {
  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
      "content-type": "application/json",
    },
    body: JSON.stringify({ model, max_tokens: 1024, messages }),
  });

  const data = await res.json();
  const { input_tokens, output_tokens } = data.usage;

  const record = {
    timestamp: new Date().toISOString(),
    userId,
    model,
    input_tokens,
    output_tokens,
    total_tokens: input_tokens + output_tokens,
  };

  fs.appendFileSync("usage_log.jsonl", JSON.stringify(record) + "\n");
  return data;
}

This is fine for a prototype. For anything running in production you'll want to swap the fs.appendFileSync call for a write to Postgres, SQLite, or a time-series store, since a flat JSONL file gets painful to query once you have more than a few thousand rows.

Aggregating usage by user or day

Once you're logging consistently, aggregation is just a query. If you're using SQLite or Postgres:

SELECT
  user_id,
  DATE(timestamp) AS day,
  SUM(input_tokens) AS total_input,
  SUM(output_tokens) AS total_output,
  SUM(input_tokens + output_tokens) AS total_tokens
FROM usage_log
GROUP BY user_id, day
ORDER BY day DESC;

If you want to estimate cost, multiply by the published per-token rate for your model and add input/output separately, since they're usually priced differently. Keep the raw token counts and the rate as separate fields in your schema — rates change, and you don't want to have baked a stale price into historical rows.

Handling streaming responses

If you stream, don't try to count tokens from the text chunks yourself — token counts from character length are unreliable. Instead, listen for the final event in the stream and pull usage from there:

for await (const event of stream) {
  if (event.type === "message_delta" && event.usage) {
    // cumulative usage so far
  }
  if (event.type === "message_stop") {
    // final usage is on the preceding message_delta or message event
  }
}

The exact event names depend on the SDK version you're using, so check the response payload shape in your logs before assuming the field names above are current.

Where this gets harder

A hand-rolled script works well for a single API key and a single service. It gets messy fast once you have multiple apps, multiple team members with their own keys, or a need to see usage broken down by application rather than by raw API key — Anthropic's dashboard doesn't give you that breakdown natively, so you end up building your own tagging convention and hoping everyone on the team follows it.

This is the gap SubToAPI fills. Instead of writing and maintaining a logging pipeline for every service that calls Claude, you issue a separate sub_live_... key per application or team member from one dashboard, and usage metadata — tokens in, tokens out, per-key breakdowns — is already tracked and visible without writing a script at all. If you're already deep into building your own tracker, it's worth checking /pricing to see if the per-seat cost is lower than the engineering time of maintaining it. The /docs/messages page shows the exact response shape, including usage fields, if you want to compare it against what you're parsing today.

Keeping the script maintainable

A few practices make a difference once this script is running in production:

If you're building this inside a larger integration, pairing it with retry and backoff logic is also worth doing early — the /docs/quickstart guide covers the request lifecycle if you're setting this up from scratch.

questions

Does the Claude API have a built-in usage tracking endpoint? No. Usage data is returned inline in each response's usage field — there's no separate endpoint to query historical totals, so you need to capture and store it yourself at call time.

Can I track token usage without modifying my application code? Not with the raw Anthropic API, since usage only appears in the response payload. Using a proxy layer like SubToAPI, which surfaces usage metadata per key in a dashboard, avoids adding logging code to every service.

What's the difference between input and output token tracking? Input tokens cover your prompt and context (often including cached content), output tokens cover the generated response. They're usually priced differently, so track and sum them separately rather than combining them into one "tokens used" number.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →