Claude API Latency Benchmark Comparison Test Guide
Anyone searching for a "Claude API latency benchmark comparison test" usually wants one of two things: a number to sanity-check against ("is 2.5s for a 500-token response normal?"), or a repeatable method to run their own test because published benchmarks are always stale by the time you read them. This article gives you both — a methodology you can run in five minutes, and the variables that actually explain the differences you'll see between runs, models, and providers.
The short answer: Claude API latency is dominated by output token count, not network overhead. A well-formed request to claude-3-5-haiku typically returns a first token in 300–600ms and streams at roughly 60–120 tokens/second depending on load. Larger models like claude-opus-4 have higher time-to-first-token (often 800ms–1.5s) but comparable or better per-token throughput once streaming starts. Any benchmark that doesn't separate time-to-first-token (TTFT) from total completion time is not measuring anything useful.
What "latency" actually means for an LLM API
Before comparing numbers, agree on what you're measuring. There are three distinct metrics that get lumped together as "latency":
- Time to first token (TTFT) — how long from request sent to the first byte of the response. This is what matters for interactive UIs where you're streaming to a chat window.
- Total completion time — how long the full response takes, dominated by output length. A 2,000-token answer will always take longer than a 50-token answer, regardless of provider.
- Tokens per second (throughput) — completion time minus TTFT, divided by output tokens. This is the number that's actually comparable across requests of different lengths.
If you're benchmarking for a chat product, optimize for TTFT. If you're benchmarking for batch summarization or data extraction, optimize for throughput and total time. Conflating the two produces misleading comparisons.
A minimal benchmark script
Here's a Node script that measures TTFT and throughput against a streaming endpoint. It works against the Anthropic API directly or against any compatible gateway — just swap the base URL and headers.
async function benchmark(prompt, model) {
const start = performance.now();
let firstTokenAt = null;
let tokenCount = 0;
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01"
},
body: JSON.stringify({
model,
max_tokens: 500,
stream: true,
messages: [{ role: "user", content: prompt }]
})
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
if (firstTokenAt === null) firstTokenAt = performance.now();
const chunk = decoder.decode(value);
tokenCount += (chunk.match(/"text":/g) || []).length;
}
const end = performance.now();
return {
ttft_ms: Math.round(firstTokenAt - start),
total_ms: Math.round(end - start),
approx_tokens: tokenCount,
tokens_per_sec: (tokenCount / ((end - firstTokenAt) / 1000)).toFixed(1)
};
}
Run this 20–30 times per model with the same prompt, throw out the first two runs (cold connection), and take the median — not the average. Latency distributions are right-skewed; a single slow outlier will distort a mean but barely moves a median.
Variables that skew your results
Most "Claude API is slow" complaints trace back to one of these, not the model itself:
- Prompt caching state. If you're using prompt caching and hit a cache miss, you'll see materially higher TTFT than a cache hit. Always note cache status in your results.
- Output length. Comparing a 50-token response against a 1,500-token response and calling one "faster" is comparing apples to oranges. Fix
max_tokensor measure throughput instead of total time. - Region and network path. If your client is far from the API's edge, you're measuring network latency, not model latency. Run benchmarks from the same region as your production deployment.
- Time of day / load. Latency varies with overall API load. A single benchmark run tells you almost nothing — run it across different hours and days before drawing conclusions.
- Tool use and extended thinking. Requests that invoke tools or use extended thinking have fundamentally different latency profiles because they involve multiple internal steps. Don't mix these into a "plain completion" benchmark.
- Streaming vs non-streaming. Non-streaming requests report latency as one blocking call, which inflates perceived TTFT to the full completion time. If you're benchmarking for UX, always use streaming.
Comparing across gateways and proxies
If you're testing Claude access through a proxy, gateway, or an API wrapper instead of calling Anthropic directly, add one more measurement: the overhead the gateway itself introduces. This is usually small (10–50ms) if the gateway is just forwarding requests and streaming responses through, but it can be large if the gateway buffers the full response before returning it — which defeats the purpose of streaming entirely.
This matters if you're evaluating a service like SubToAPI, which turns your existing Claude access into an HTTPS API with sub_live_... keys. SubToAPI streams responses through without buffering, so TTFT overhead versus calling Claude directly is minimal — you're mainly paying the cost of one extra network hop. When benchmarking any gateway, run the same script above against both the gateway and the direct API with identical prompts and max_tokens, then diff the medians. See the streaming docs for the exact event format, and the quickstart if you want to reproduce these numbers on your own account — a free trial is available at signup.
A reasonable benchmark checklist
- Separate TTFT from throughput from total time
- Use streaming, not blocking requests
- Fix
max_tokensand prompt length across comparisons - Run 20+ samples, report the median, discard the first two
- Note prompt cache hit/miss state explicitly
- Test from the same region as your production traffic
- Run at multiple times across at least 2–3 days before publishing a number
FAQ
Is Claude API latency consistent across models? No. Smaller models like Haiku have lower TTFT and higher throughput; larger models like Opus have higher TTFT but strong per-token throughput once streaming begins. Always benchmark the specific model you plan to use in production.
Why does my benchmark show wildly different numbers each run? Latency distributions are right-skewed — occasional slow outliers are normal. Run at least 20 samples and report the median, not the average, and rule out cache misses and network variance before concluding anything changed.
Does routing through a gateway or proxy add significant latency? It depends on whether the gateway streams responses through in real time or buffers them first. A pass-through streaming gateway typically adds under 50ms; a buffering one can add the full completion time to your perceived TTFT. Test both configurations directly against your own traffic pattern.