Claude API Logging and Monitoring Setup Guide
A production Claude integration without logging is a black box. When a user complains that a response was slow, wrong, or never arrived, you need to be able to answer: what prompt was sent, what model responded, how long it took, how many tokens were used, and whether an error occurred. This article walks through a practical logging and monitoring setup for the Claude API — what to capture, how to structure it, and how to turn raw logs into alerts you actually act on.
The short version: log every request and response at the metadata level (not necessarily full content), track latency and token usage per call, capture errors with status codes and retry counts, and feed all of it into a dashboard or alerting system so you find out about problems before your users tell you.
What to log on every request
At minimum, capture these fields for every call to the Claude API:
- Request ID — generate your own UUID per call, independent of any provider request ID, so you can correlate logs across services
- Model — which model was used (this matters once you start routing between models for cost or capability reasons)
- Timestamp — when the request started and when it completed
- Latency — total round-trip time, and time-to-first-token if you're streaming
- Token counts — input tokens, output tokens, and cache read/write tokens if you use prompt caching
- Status — success, client error, server error, timeout
- Retry count — how many times this request was retried before succeeding or failing
- User or tenant ID — whichever identifier lets you trace usage back to an account
Whether you log full prompt and response text is a policy decision, not a technical one. Many teams log a hash or truncated snippet in production and only log full content in a staging environment or behind a feature flag, to avoid storing sensitive data in logs that have wider access than your primary database.
A minimal structured logging wrapper
Wrap every API call in a function that emits a structured log line regardless of outcome:
async function callClaude(payload, context) {
const requestId = crypto.randomUUID();
const start = Date.now();
let status = "success";
let usage = null;
let errorMessage = null;
try {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
if (!res.ok) {
status = res.status >= 500 ? "server_error" : "client_error";
}
const data = await res.json();
usage = data.usage;
return data;
} catch (err) {
status = "exception";
errorMessage = err.message;
throw err;
} finally {
console.log(JSON.stringify({
requestId,
model: payload.model,
userId: context.userId,
status,
latencyMs: Date.now() - start,
inputTokens: usage?.input_tokens,
outputTokens: usage?.output_tokens,
error: errorMessage,
}));
}
}
This gives you a JSON line per call that any log aggregator (CloudWatch, Datadog, Loki, Logtail) can parse without custom regex. Ship stdout to your aggregator rather than writing your own file-based logging — it's less to maintain and scales with your deployment.
Monitoring streaming requests
Streaming responses need an extra metric: time-to-first-token. A request that takes 8 seconds to fully complete but starts streaming after 200ms feels fast to a user; one that starts streaming after 6 seconds feels broken, even if total latency is similar. Capture the timestamp of the first chunk separately from the timestamp of completion. If you're implementing streaming from scratch, see /docs/streaming for the event format and how chunks map to content blocks.
For tool use requests, also log which tools were invoked and how many tool-call round trips a single user turn required — a conversation that triggers five tool calls costs and behaves very differently from one that triggers zero. /docs/tools covers the tool-call event structure if you need the exact fields to parse.
Turning logs into alerts
Logs you never look at are not monitoring. At minimum, set up alerts for:
- Error rate — alert if the 5xx or client-error rate over a 5-minute window exceeds a threshold (start with something strict, like 2%, and tune from there)
- P95 latency — alert if the 95th percentile latency for a given model doubles compared to its 24-hour baseline
- Token usage spikes — a sudden jump in output tokens per request often indicates a prompt regression or a loop
- Rate limit hits — track 429 responses separately; frequent rate limiting is a capacity problem, not a bug
Most teams wire these into whatever alerting stack they already use — PagerDuty, Opsgenie, a Slack webhook — rather than building something bespoke. The key is making sure token usage and latency, not just raw error counts, are part of the alert surface, since a Claude integration can degrade in cost or speed without ever throwing an error.
Where SubToAPI fits in
If you're running Claude through SubToAPI, a chunk of this is already handled for you. Every application key (sub_live_...) reports usage metadata — tokens, latency, and request status — in the dashboard, broken down per key and per team seat, so you get per-application monitoring without wiring up a separate aggregation pipeline for basic usage visibility. You'd still want your own application-level logging for correlating requests with your own user IDs and business logic, but the raw usage and cost tracking comes built in. See /docs/quickstart to get an API key, and /docs/messages for the request format if you're integrating against it directly. Plans start at Solo €9/month, with Team and Scale tiers adding per-seat usage visibility — details on /pricing.
questions
Do I need to log full prompt and response text? Not necessarily. Log metadata (tokens, latency, status) in production by default, and only log full content behind a flag or in a restricted environment, since prompts and responses can contain sensitive user data.
What's the single most useful metric to monitor first? Error rate by status code, broken down by model. It's cheap to compute, surfaces outages and bad requests immediately, and gives you a baseline before you add latency or token-based alerts.
How do I monitor streaming latency specifically? Capture two timestamps per request: when the first chunk arrives (time-to-first-token) and when the stream closes (total latency). Track them separately — they measure different user experiences. See /docs/streaming for the event format.