Claude API Multi-Tenant SaaS Architecture Guide
Building a multi-tenant SaaS on top of the Claude API means solving four problems at once: keeping each customer's data and usage isolated, controlling cost per tenant, giving each tenant their own credentials without exposing your master API key, and tracking usage accurately enough to bill for it. None of these are solved by calling the Anthropic API directly from your backend — you need a layer in between.
The short answer to "how do I architect this" is: put a proxy or gateway between your application and the Claude API, give every tenant (or every tenant's application) a scoped key that maps back to your one Anthropic account, and log every request against that key for rate limiting, quotas, and billing. Everything below breaks down how to implement that in practice.
Tenant isolation models
There are two common ways to isolate tenants when you're building on a single upstream Claude API account:
Logical isolation (most SaaS products use this). All tenants share the same underlying Anthropic account and rate limits, but every tenant gets a distinct key issued by your platform. Requests are tagged, logged, and billed per key. This is cheaper to operate and scales well because you're not managing N separate Anthropic accounts.
Physical isolation. Each tenant (or each enterprise customer) gets a dedicated Anthropic account or API key pool, usually reserved for compliance-heavy customers who need contractual guarantees that their prompts and data never share infrastructure with anyone else's traffic. This is more expensive to run and only worth it for a handful of large accounts.
Most products start with logical isolation and offer physical isolation as an enterprise tier later. Don't over-engineer this on day one.
Give every tenant their own scoped key
Whatever isolation model you pick, don't let tenant applications call Claude directly with your organization's raw API key. If one tenant's key leaks, you want to revoke that one key — not rotate your entire account and break every other tenant at once.
The pattern is:
Tenant App → Your Gateway (scoped key) → Claude API (your master credentials)
Your gateway is the only thing that holds the real Anthropic credentials. Tenants get keys scoped to their tenant ID, with their own rate limits and spend caps. This is exactly the model SubToAPI implements if you don't want to build the gateway yourself: it sits in front of your Claude access and issues sub_live_... keys per application, so each tenant (or each internal service) gets its own key, its own usage metadata, and its own revoke button, without you managing token exchange or key rotation logic. See the quickstart for how key issuance works.
If you're building the gateway yourself, a minimal key-issuance table looks like this:
create table api_keys (
id uuid primary key,
tenant_id uuid not null references tenants(id),
key_hash text not null,
rate_limit_rpm int not null default 60,
monthly_token_cap bigint,
revoked_at timestamptz
);
Hash the key at rest, look it up on every request, and reject anything revoked or over cap before you ever touch the Claude API.
Rate limiting and quotas per tenant
A single noisy tenant can burn through your Anthropic rate limit and degrade the experience for everyone else. Enforce limits at two levels:
- Per-tenant request rate (requests per minute) to stop bursts from one customer starving others.
- Per-tenant token budget (daily or monthly) tied to their subscription tier, checked before you forward the request upstream.
A simple middleware check before the Claude call:
async function enforceQuota(tenantId) {
const usage = await getMonthlyTokenUsage(tenantId);
const cap = await getTenantCap(tenantId);
if (usage >= cap) {
throw new Error("quota_exceeded");
}
}
Keep this check fast — an in-memory or Redis counter updated asynchronously from your logging pipeline works better than querying a usage table on every request.
Usage tracking that supports billing
If tenants pay you based on usage (per-seat, per-token, or tiered), you need reliable, per-tenant usage metadata, not just aggregate logs. Every response from the Claude API includes token counts in the usage field — capture input_tokens and output_tokens per request and attribute them to the tenant and, ideally, the specific application or feature that made the call.
const response = await client.messages.create({
model: "claude-sonnet-4-5",
max_tokens: 1024,
messages: [{ role: "user", content: prompt }],
});
await logUsage({
tenantId,
inputTokens: response.usage.input_tokens,
outputTokens: response.usage.output_tokens,
});
This is the same data you'd expose to tenants in a usage dashboard, and it's what a platform like SubToAPI already surfaces per key out of the box, along with request logs and streaming support — see /docs/messages for the response shape and /docs/streaming if your tenants need token-by-token output.
Where the gateway should live
For most SaaS products, the gateway is a thin service sitting between your application layer and Claude: it authenticates the tenant key, checks quota, forwards the request, streams the response back, and writes a usage record. Keep it stateless and horizontally scalable — the state (quotas, key status, usage totals) lives in your database or cache, not in the gateway process.
If tenants need tool use (function calling) against their own data sources, route tool definitions and execution through your gateway too, so each tenant's tools stay scoped to their own tenant context and can't leak into another tenant's session. The tool use docs cover the request/response format if you're implementing this from scratch.
Teams that don't want to build and maintain this gateway — key issuance, quota enforcement, streaming, usage logs — can point their existing Claude access through SubToAPI instead and get application keys, per-seat team access, and usage metadata immediately; see /pricing for the plan breakdown and /signup to start a trial.
Security checklist
- Master Anthropic credentials never leave your gateway process.
- Tenant keys are hashed at rest and revocable individually.
- Every request is logged with tenant ID, token counts, and timestamp before being discarded.
- Quota checks happen before the upstream call, not after.
- Tool execution is scoped per tenant session, never shared.
questions
Do I need a separate Anthropic account per tenant? No. Most multi-tenant SaaS products use one Anthropic account behind a gateway, with per-tenant scoped keys for isolation and billing. Dedicated accounts are usually reserved for enterprise customers with specific compliance requirements.
How do I bill tenants based on Claude usage? Capture input_tokens and output_tokens from every API response and attribute them to the tenant that triggered the request. Aggregate this into daily or monthly totals and map it to your pricing tiers.
What's the fastest way to add multi-tenant key management without building it myself? A hosted proxy that issues scoped keys per application, tracks usage per key, and handles streaming and tool use — for example SubToAPI — lets you skip building the gateway layer and start with per-tenant keys on day one.