← Blog

Claude API Request Logging and Monitoring Guide

2026-10-02 · 5 min read · SubToAPI Team

If you're running Claude in production, you need to know what's being sent, what's coming back, how long it takes, and what it costs — without digging through application code every time something looks off. Claude API request logging and monitoring means capturing structured data about every call (prompts, responses, latency, token usage, errors) and surfacing it somewhere you can actually query and alert on.

This matters for three practical reasons: debugging why a specific response was wrong or truncated, tracking spend before it surprises you at the end of the month, and catching reliability problems (rate limits, timeouts, malformed tool calls) before users report them. Below is a concrete setup you can apply whether you're calling the Anthropic API directly or through a proxy.

What to log on every request

At minimum, capture these fields per call:

Don't log raw API keys. If you're storing prompts and responses for debugging, make sure you have a retention policy and, if you handle regulated data, a redaction step before anything hits disk.

A minimal logging wrapper

async function callClaude(messages, { model = "claude-sonnet-4-5" } = {}) {
  const requestId = crypto.randomUUID();
  const start = Date.now();

  try {
    const res = await fetch("https://api.anthropic.com/v1/messages", {
      method: "POST",
      headers: {
        "x-api-key": process.env.ANTHROPIC_API_KEY,
        "anthropic-version": "2023-06-01",
        "content-type": "application/json",
      },
      body: JSON.stringify({ model, max_tokens: 1024, messages }),
    });

    const data = await res.json();
    const latencyMs = Date.now() - start;

    logEvent({
      requestId,
      model,
      status: res.status,
      inputTokens: data.usage?.input_tokens,
      outputTokens: data.usage?.output_tokens,
      latencyMs,
      ok: res.ok,
    });

    return data;
  } catch (err) {
    logEvent({ requestId, model, error: String(err), latencyMs: Date.now() - start });
    throw err;
  }
}

logEvent can write to stdout (for a log aggregator to pick up), push to a queue, or insert directly into a database table. The important part is that every call goes through one place, so you never have untracked requests scattered across your codebase.

Where to send the logs

Three common patterns, roughly in order of setup effort:

  1. Structured stdout + log aggregator — print JSON lines and let Datadog, CloudWatch, or Loki ingest them. Fast to set up, good for ephemeral debugging, weaker for long-term analytics.
  2. A dedicated table — insert each event into Postgres or a similar store. Lets you run SQL for cost reports, slow-query investigations, and per-user usage breakdowns.
  3. A metrics/tracing backend — emit latency and token counts as metrics (Prometheus, OpenTelemetry) so you get dashboards and alerting without querying raw logs.

Most teams end up combining all three: metrics for alerting, a database table for reporting, and full request/response logs retained for a shorter window (7–30 days) for debugging.

Monitoring: what to alert on

Logging tells you what happened. Monitoring tells you when something's wrong right now. Set alerts on:

A simple daily job that aggregates yesterday's logs into total requests, total tokens, error rate, and average latency per model is often more useful than a live dashboard nobody checks.

If you don't want to build this yourself

Building and maintaining a logging pipeline — ingestion, storage, retention, dashboards — is a real project, not a one-afternoon script. If your priority is having clean per-key usage metadata without standing up your own infrastructure, SubToAPI gives you that out of the box: every request made with a sub_live_... key is already tracked with token counts and usage data visible in the dashboard, so you get the monitoring layer without building it. It sits in front of your existing Claude access and exposes a standard HTTPS API — see the quickstart and messages docs for the request format, or streaming if you need token-by-token output with the same visibility. Plans start at €9/month with a free trial at signup; full details on pricing.

Keeping it maintainable

A few habits keep a logging setup useful instead of becoming its own maintenance burden:

Questions

Does Anthropic provide built-in request logging? The API itself doesn't give you a dashboard of historical requests — you need to capture usage data from each response (the usage object) and store it yourself, or use a layer that does this for you.

What's the minimum I should log to track costs accurately? Input tokens, output tokens, model name, and timestamp per request. That's enough to compute spend per day, per model, or per user if you tag requests with a caller ID.

Should I store full prompts and responses? Only if you have a clear reason (debugging, quality review) and a retention/redaction policy. For cost and performance monitoring alone, token counts and latency are usually sufficient.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →