Claude API Audit Logging for Compliance: A Setup Guide
If you're integrating Claude into a product that touches customer data, finance, healthcare, or internal tooling with access controls, you'll eventually get the question: "can you show me who called the model, with what data, and when?" That's audit logging, and the Claude API does not give it to you out of the box. Anthropic's API returns a response and some usage tokens — it doesn't persist a queryable log of requests, doesn't tag them by internal user, and doesn't give you a compliance-ready export.
This article covers what a proper audit trail for Claude API usage needs to contain, how to build it yourself if you're calling the API directly, and how to get most of it without building anything if you're using a gateway layer like SubToAPI.
What "audit logging for compliance" actually means here
Compliance frameworks (SOC 2, ISO 27001, HIPAA, internal security reviews) generally want answers to five questions for any system that processes data:
- Who initiated the action (which internal user, service account, or API key)
- What was sent and received (or at minimum, metadata about it — some frameworks require content, others forbid logging raw PII)
- When it happened (timestamp, ideally with timezone and latency)
- Where it came from (IP, environment, application/service name)
- What happened as a result (success, error, tokens consumed, cost)
A single console.log of your Claude calls is not an audit log. An audit log needs to be immutable (or at least tamper-evident), queryable by user and time range, retained for a defined period, and exportable for an auditor.
Building audit logging around raw Claude API calls
If you're calling api.anthropic.com directly, you own 100% of this. The minimum viable version:
async function callClaudeWithAudit(payload, context) {
const startedAt = Date.now();
const requestId = crypto.randomUUID();
const auditEntry = {
request_id: requestId,
actor_id: context.userId,
actor_role: context.role,
app: context.appName,
ip: context.ip,
model: payload.model,
started_at: new Date(startedAt).toISOString(),
};
try {
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01",
"content-type": "application/json",
},
body: JSON.stringify(payload),
});
const json = await response.json();
await writeAuditLog({
...auditEntry,
status: response.status,
input_tokens: json.usage?.input_tokens,
output_tokens: json.usage?.output_tokens,
duration_ms: Date.now() - startedAt,
// Decide deliberately whether to store prompt/response text
prompt_hash: hashIfNeeded(payload.messages),
});
return json;
} catch (err) {
await writeAuditLog({ ...auditEntry, status: "error", error: String(err) });
throw err;
}
}
A few decisions you need to make explicitly, not by default:
- Do you store raw prompt/response content, or just a hash? Storing full content gives you better forensic ability but creates a second place where sensitive data lives, which itself needs its own access controls and retention policy. Many teams hash the content and only store full text for a short, separately-encrypted window.
- Where does the log go? A database table works, but append-only storage (a dedicated audit table with no
UPDATE/DELETEgrants, or a write-once log shipping pipeline to something like S3 with object lock) is what auditors actually want to see. - Who can read the audit log? If every engineer with database access can also edit the audit log, it's not really an audit log. Separate the write path from the admin read path.
- Retention. SOC 2 typically expects at least 90 days to a year depending on control type; some regulated industries require years. Decide this before storage costs force your hand.
This is maybe a day of engineering work for a basic version, and then ongoing maintenance every time you add a new endpoint, a new internal service, or a new failure mode that needs to be captured.
Logging by API key instead of by code path
One mistake teams make is logging at the application-code level only — a try/catch around one function — and missing all the other places Claude gets called from (cron jobs, admin scripts, a second service that reuses the same key). If multiple services share one Anthropic API key, you lose the ability to attribute usage cleanly, which defeats the "who" requirement before you've even started.
The fix is to issue a distinct key per service or per team, so attribution happens at the credential level, not just in application logs. This is one of the reasons per-key usage separation matters even outside of compliance — it also makes cost allocation and incident response faster.
Doing less manual work: gateway-level logging
If you don't want to build and maintain this logging layer yourself, routing Claude traffic through a gateway that already separates keys and tracks usage per credential removes a chunk of the work. SubToAPI issues scoped application keys (sub_live_...) per team member or service, so every request against https://api.subtoapi.app/v1/messages is already attributable to a specific key without you having to thread a context.userId through every call site:
curl https://api.subtoapi.app/v1/messages \
-H "Authorization: Bearer $SUBTOAPI_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4.5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Summarize this ticket."}]
}'
Because each key maps to a team seat in the dashboard, usage metadata (tokens, model, timestamp, success/failure) is already segmented by who made the call, which covers the "who," "when," and "what happened" parts of the audit requirement without custom infrastructure. You still own the decision about whether to log full prompt content elsewhere for your own compliance needs — SubToAPI doesn't replace a data retention policy, but it removes the key-management and per-call attribution plumbing. See the quickstart and messages docs for the request format, and pricing for plan details if you're evaluating it for a team.
A minimal compliance checklist
Before an audit, confirm you can answer these without manual digging:
- Can you list every Claude API call made by a specific user in the last 90 days?
- Can you prove the audit log itself hasn't been edited (append-only storage, checksums, or a managed platform)?
- Do you know exactly what's stored in each log entry — full content, hashes, or metadata only — and is that documented?
- Is API key access reviewed and rotated, with old keys revoked, not just rotated silently?
- Is there a documented retention and deletion policy for the logs themselves?
If any of these is "we'd have to check," that's the gap to close first — it's usually cheaper to fix before an auditor asks than during the review.
FAQ
Does Claude's API provide built-in audit logs? No. The API returns usage tokens in each response but does not store or expose a historical, queryable log of requests. You need to capture and store that data yourself or use a layer that does it for you.
What's the difference between application logs and audit logs? Application logs are for debugging and are often mutable, rotated, or incomplete. Audit logs need to be attributable to a specific actor, tamper-evident, retained for a defined period, and reviewable by someone other than the engineer who wrote the code.
Should I log full prompt and response content? Only if you have a clear policy for it. Logging full content gives better forensic detail but increases your sensitive-data footprint. Many teams log hashes or truncated metadata by default and store full content separately with stricter access controls, if at all.