← Blog

Building a Claude API Billing API for Customers

2026-10-06 · 5 min read · SubToAPI Team

What "billing API for customers" actually means here

If you're searching for a Claude API billing API for customers, you're probably not asking about Anthropic's own invoice or payment system. You're building a product on top of Claude — a SaaS, an internal tool, an agency reselling AI features — and you need a way to meter usage per customer so you can bill them, cap them, or show them a usage dashboard. Anthropic's console gives you organization-wide usage, not a per-customer billing layer.

The short answer: you need three things working together — isolated credentials per customer (or at least a reliable way to tag requests), usage metadata on every call (input tokens, output tokens, model used), and a place to aggregate that data into something you can invoice from or expose to customers directly. Below is how to build that, and where a layer like SubToAPI removes most of the manual work.

Why Anthropic's native billing doesn't solve this

Anthropic bills you, the account owner, for total token consumption across your org. There's no built-in concept of "customer A used 40,000 tokens, customer B used 12,000 tokens" unless you build that mapping yourself. If you're reselling Claude access or embedding it in a multi-tenant product, you end up needing:

Most teams solve this by writing a thin metering layer around the raw Claude API. It works, but it's a maintenance burden you didn't sign up for.

The DIY approach: metering at the request layer

The baseline pattern looks like this: wrap every Claude call, capture the usage object from the response, and write it to your own database keyed by customer ID.

async function callClaudeForCustomer(customerId, messages) {
  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
      "content-type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      messages,
    }),
  });

  const data = await res.json();

  await db.usageLedger.insert({
    customerId,
    inputTokens: data.usage.input_tokens,
    outputTokens: data.usage.output_tokens,
    model: data.model,
    createdAt: new Date(),
  });

  return data;
}

That's fine for a prototype. In production it falls apart quickly: streaming responses need usage extracted from the final event, failed requests still need partial accounting, and you now own a billing-adjacent database that has to be correct, because it's directly tied to money.

A cleaner path: per-customer API keys with usage baked in

Instead of tagging requests yourself, give each customer (or each internal project) its own application API key and let usage metadata come attached to every call. This is the model SubToAPI is built around: you connect your Claude access once, then issue scoped sub_live_... keys — one per customer, per project, or per environment — and every request against that key returns usage data you can read programmatically.

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Summarize this ticket."}]
  }'

Because each customer gets a distinct key, "billing per customer" becomes "read usage per key" instead of building your own attribution logic. You still own the invoicing step (Stripe, manual, whatever fits your product), but you're no longer responsible for the metering plumbing or for keeping token counts accurate across streaming and tool-use calls. See /docs/quickstart for setup and /docs/messages for the request/response shape, including streaming at /docs/streaming and tool calls at /docs/tools.

What this gets you that a raw wrapper doesn't

Designing the billing logic on top

Once usage data is reliable, the billing API you build for your customers is just an aggregation and pricing layer:

  1. Decide your unit: per-token, per-request, or a flat markup over Claude's cost
  2. Pull usage per key on a schedule (daily rollups are usually enough)
  3. Apply your margin and generate a line item per customer
  4. Expose a read endpoint so customers can see their own usage in near real time, not just at invoice time
async function getCustomerUsage(subtoapiKey) {
  // Pseudocode: pull your own stored usage records for this key,
  // or your aggregation if you're logging usage from responses.
  const usage = await db.usageLedger.sumByKey(subtoapiKey);
  return {
    inputTokens: usage.inputTokens,
    outputTokens: usage.outputTokens,
    estimatedCost: usage.outputTokens * YOUR_RATE_PER_TOKEN,
  };
}

This is also where plan tiers matter for your own cost base. SubToAPI's plans — Solo at €9, Team at €19/seat, Scale at €49/seat — determine how many seats and keys you're working with, which affects how you structure markup for your customers. Check current details at /pricing before deciding how to price your own product on top.

Keep it simple at first

Don't build a full metering system before you have customers to meter. Start with one key per customer, log usage from every response, and generate invoices manually or through Stripe metered billing. Automate the parts that get repetitive — usage pulls, invoice generation — once the pattern is proven. If you want to skip the key-management and metadata plumbing entirely, /signup gets you a working setup in minutes, and /docs covers the rest.

FAQ

Does Anthropic provide per-customer billing natively? No. Anthropic's console shows organization-wide usage and billing. Per-customer attribution has to be built by you, either through request tagging or by issuing separate keys per customer.

What's the easiest way to bill customers for Claude usage without building a metering system? Issue each customer a distinct API key so usage is naturally isolated per key, then read usage metadata from responses instead of calculating token counts yourself. This is the approach SubToAPI's application keys are designed for.

Should I pass my raw Claude API key to customers? No. Giving out a shared or raw key means no isolation, no per-customer revocation, and no reliable way to meter individual usage. Use scoped application keys per customer instead.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →