Claude API Request Logging and Monitoring Guide
If you're running Claude in production, you need to know what's being sent, what's coming back, how long it takes, and what it costs — without digging through application code every time something looks off. Claude API request logging and monitoring means capturing structured data about every call (prompts, responses, latency, token usage, errors) and surfacing it somewhere you can actually query and alert on.
This matters for three practical reasons: debugging why a specific response was wrong or truncated, tracking spend before it surprises you at the end of the month, and catching reliability problems (rate limits, timeouts, malformed tool calls) before users report them. Below is a concrete setup you can apply whether you're calling the Anthropic API directly or through a proxy.
What to log on every request
At minimum, capture these fields per call:
- Request ID — a UUID you generate, used to correlate logs across services
- Timestamp — request start and end time
- Model — which Claude model handled the call
- Input tokens / output tokens — from the response usage object
- Latency — time to first byte (for streaming) and total duration
- HTTP status code — 200, 429, 529, etc.
- Error body — if the call failed, the full error payload
- Caller identity — which user, team, or application triggered the call
- Prompt and response — full text, or at least a hash/truncated version if you have privacy constraints
Don't log raw API keys. If you're storing prompts and responses for debugging, make sure you have a retention policy and, if you handle regulated data, a redaction step before anything hits disk.
A minimal logging wrapper
async function callClaude(messages, { model = "claude-sonnet-4-5" } = {}) {
const requestId = crypto.randomUUID();
const start = Date.now();
try {
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify({ model, max_tokens: 1024, messages }),
});
const data = await res.json();
const latencyMs = Date.now() - start;
logEvent({
requestId,
model,
status: res.status,
inputTokens: data.usage?.input_tokens,
outputTokens: data.usage?.output_tokens,
latencyMs,
ok: res.ok,
});
return data;
} catch (err) {
logEvent({ requestId, model, error: String(err), latencyMs: Date.now() - start });
throw err;
}
}
logEvent can write to stdout (for a log aggregator to pick up), push to a queue, or insert directly into a database table. The important part is that every call goes through one place, so you never have untracked requests scattered across your codebase.
Where to send the logs
Three common patterns, roughly in order of setup effort:
- Structured stdout + log aggregator — print JSON lines and let Datadog, CloudWatch, or Loki ingest them. Fast to set up, good for ephemeral debugging, weaker for long-term analytics.
- A dedicated table — insert each event into Postgres or a similar store. Lets you run SQL for cost reports, slow-query investigations, and per-user usage breakdowns.
- A metrics/tracing backend — emit latency and token counts as metrics (Prometheus, OpenTelemetry) so you get dashboards and alerting without querying raw logs.
Most teams end up combining all three: metrics for alerting, a database table for reporting, and full request/response logs retained for a shorter window (7–30 days) for debugging.
Monitoring: what to alert on
Logging tells you what happened. Monitoring tells you when something's wrong right now. Set alerts on:
- Error rate — spike in non-200 responses over a 5-minute window
- 429 / 529 responses — rate limiting or overload, usually needs backoff logic, not just a page
- P95 latency — sudden increases often mean a model or region issue, not your code
- Token usage anomalies — a prompt template change that doubles input tokens will double your bill before anyone notices in a dashboard
- Silent failures — streaming responses that end mid-tool-call or with an unexpected stop reason
A simple daily job that aggregates yesterday's logs into total requests, total tokens, error rate, and average latency per model is often more useful than a live dashboard nobody checks.
If you don't want to build this yourself
Building and maintaining a logging pipeline — ingestion, storage, retention, dashboards — is a real project, not a one-afternoon script. If your priority is having clean per-key usage metadata without standing up your own infrastructure, SubToAPI gives you that out of the box: every request made with a sub_live_... key is already tracked with token counts and usage data visible in the dashboard, so you get the monitoring layer without building it. It sits in front of your existing Claude access and exposes a standard HTTPS API — see the quickstart and messages docs for the request format, or streaming if you need token-by-token output with the same visibility. Plans start at €9/month with a free trial at signup; full details on pricing.
Keeping it maintainable
A few habits keep a logging setup useful instead of becoming its own maintenance burden:
- Version your logging schema — add fields, don't repurpose existing ones
- Set retention explicitly (e.g., 30 days for full payloads, 1 year for aggregated metrics)
- Tag requests with an environment field (dev/staging/prod) from day one
- Review the error log weekly even if nothing paged — slow degradation rarely trips alerts
Questions
Does Anthropic provide built-in request logging? The API itself doesn't give you a dashboard of historical requests — you need to capture usage data from each response (the usage object) and store it yourself, or use a layer that does this for you.
What's the minimum I should log to track costs accurately? Input tokens, output tokens, model name, and timestamp per request. That's enough to compute spend per day, per model, or per user if you tag requests with a caller ID.
Should I store full prompts and responses? Only if you have a clear reason (debugging, quality review) and a retention/redaction policy. For cost and performance monitoring alone, token counts and latency are usually sufficient.