Claude API Cost Per Conversation Tracking Guide
Tracking Claude API cost per conversation means attributing the token usage (and therefore dollar cost) of every request to the specific conversation, user, or session that generated it — not just looking at a monthly total. The Anthropic Messages API returns token counts on every response, but it doesn't group them by conversation for you. You have to capture that data yourself and map it to a conversation ID at the point of logging.
This matters because a single aggregate invoice tells you almost nothing useful. If your app serves 500 users and your Claude bill jumps 40% in a week, you need to know which conversations, features, or customers caused it — a runaway agent loop, a user pasting huge documents, or a new feature that calls the model five times per request. Per-conversation tracking turns an opaque bill into a debuggable, attributable cost model.
What the API actually gives you
Every Claude Messages API response includes a usage object:
{
"id": "msg_01XYZ",
"model": "claude-sonnet-4-5",
"usage": {
"input_tokens": 812,
"output_tokens": 347,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}
That's the raw material for cost tracking, but it's per-request, not per-conversation. A single conversation might involve several requests (each turn, each tool call round trip), so you need to sum usage across every request that belongs to the same conversation ID before you can report a meaningful per-conversation cost.
Build a minimal cost ledger
The simplest reliable approach is: generate a conversation ID in your app, attach it to every API call as metadata, and log token usage against that ID after each response.
import crypto from "crypto";
const conversationId = crypto.randomUUID();
async function sendTurn(conversationId, messages) {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages,
}),
});
const data = await res.json();
await logUsage({
conversationId,
model: data.model,
inputTokens: data.usage.input_tokens,
outputTokens: data.usage.output_tokens,
timestamp: Date.now(),
});
return data;
}
logUsage writes a row to a database (Postgres, SQLite, whatever you already use). Once you have rows keyed by conversationId, computing cost per conversation is a straightforward aggregation query:
SELECT conversation_id,
SUM(input_tokens) AS total_input,
SUM(output_tokens) AS total_output,
SUM(input_tokens) * 0.000003 +
SUM(output_tokens) * 0.000015 AS estimated_cost_usd
FROM api_usage
GROUP BY conversation_id
ORDER BY estimated_cost_usd DESC;
(Replace the per-token rates with whatever your current model pricing is — rates differ by model and change over time, so don't hardcode them permanently.)
Where this breaks down at scale
The manual ledger approach works fine for a single app with one API key. It gets harder once you have multiple services calling Claude, multiple environments (staging vs production), or a team where several engineers each hold their own API key. At that point you're reconciling usage across N different logging pipelines, and conversation-level attribution silently falls apart the moment someone forgets to pass the conversationId through.
This is the problem SubToAPI is built around. Instead of calling Anthropic directly with a shared key, each application or environment gets its own sub_live_... key, and every request made with that key is logged centrally with full usage metadata — input tokens, output tokens, model, and timestamp — visible from one dashboard. You still generate your own conversation IDs and pass them through your own logs for fine-grained grouping, but you no longer have to build separate pipelines per key or per team member to answer "what did this conversation cost" at the account level. See the quickstart for key setup and the Messages docs for the exact response shape, which matches the Anthropic API so your existing logging code doesn't need to change.
Practical patterns for per-conversation attribution
Tag at the source, not after the fact. Attach a conversation_id (and ideally a user_id) to your logging call the moment you make the request, not after parsing the response. If the request fails partway through a stream, you still want a partial usage record.
Include tool-use rounds in the total. If you're using tool use, a single conversation turn can trigger multiple round trips to the model (initial call, tool result, follow-up call). Each round trip has its own usage object — sum all of them under the same conversation ID or your per-conversation cost will be understated.
Account for streaming separately. With streaming responses, the usage data arrives in the final message_delta event, not upfront. Make sure your event handler captures it before closing the connection, otherwise you'll log zero-cost conversations that actually consumed tokens.
Separate estimated vs. billed cost. Token-based math gives you a close estimate, but your actual invoice is the source of truth. Use conversation-level tracking for debugging and alerting, not for finance-grade reconciliation.
Roll up by feature, not just by conversation. Once you have conversation-level rows, tag each conversation with a feature or endpoint label (chatbot, summarizer, agent). This lets you answer "which feature is driving cost" in addition to "which conversation."
A simple dashboard query
Once logging is in place, a weekly report is usually enough to catch anomalies:
SELECT feature,
COUNT(DISTINCT conversation_id) AS conversations,
SUM(input_tokens + output_tokens) AS total_tokens,
AVG(input_tokens + output_tokens) AS avg_tokens_per_conversation
FROM api_usage
WHERE timestamp > now() - interval '7 days'
GROUP BY feature
ORDER BY total_tokens DESC;
If avg_tokens_per_conversation suddenly spikes for one feature, that's your signal to look at prompt length, context window growth, or an agent loop that isn't terminating cleanly.
questions
Does the Claude API report cost directly, or only tokens? Only tokens. The usage field in every response gives input_tokens and output_tokens; you multiply by the current per-token rate for your model to get an estimated dollar cost. There's no cost field in the API response itself.
What's the easiest way to group usage by conversation? Generate a conversation ID in your own application, pass it alongside every request in your logging layer, and sum the usage.input_tokens and usage.output_tokens from each response under that ID. Tool-use round trips and streamed responses each need their own usage capture.
How do teams track cost per user or per project across multiple API keys? Give each project or team member a separate key so usage is naturally segmented at the source, then aggregate centrally. SubToAPI does this by issuing scoped sub_live_... keys per app with usage metadata visible in one dashboard, so you're not stitching together logs from several separate Anthropic accounts.