← Blog

Claude API Latency Benchmark Comparison Test Guide

2026-09-29 · 5 min read · SubToAPI Team

Anyone searching for a "Claude API latency benchmark comparison test" usually wants one of two things: a number to sanity-check against ("is 2.5s for a 500-token response normal?"), or a repeatable method to run their own test because published benchmarks are always stale by the time you read them. This article gives you both — a methodology you can run in five minutes, and the variables that actually explain the differences you'll see between runs, models, and providers.

The short answer: Claude API latency is dominated by output token count, not network overhead. A well-formed request to claude-3-5-haiku typically returns a first token in 300–600ms and streams at roughly 60–120 tokens/second depending on load. Larger models like claude-opus-4 have higher time-to-first-token (often 800ms–1.5s) but comparable or better per-token throughput once streaming starts. Any benchmark that doesn't separate time-to-first-token (TTFT) from total completion time is not measuring anything useful.

What "latency" actually means for an LLM API

Before comparing numbers, agree on what you're measuring. There are three distinct metrics that get lumped together as "latency":

If you're benchmarking for a chat product, optimize for TTFT. If you're benchmarking for batch summarization or data extraction, optimize for throughput and total time. Conflating the two produces misleading comparisons.

A minimal benchmark script

Here's a Node script that measures TTFT and throughput against a streaming endpoint. It works against the Anthropic API directly or against any compatible gateway — just swap the base URL and headers.

async function benchmark(prompt, model) {
  const start = performance.now();
  let firstTokenAt = null;
  let tokenCount = 0;

  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "content-type": "application/json",
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01"
    },
    body: JSON.stringify({
      model,
      max_tokens: 500,
      stream: true,
      messages: [{ role: "user", content: prompt }]
    })
  });

  const reader = res.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    if (firstTokenAt === null) firstTokenAt = performance.now();
    const chunk = decoder.decode(value);
    tokenCount += (chunk.match(/"text":/g) || []).length;
  }

  const end = performance.now();
  return {
    ttft_ms: Math.round(firstTokenAt - start),
    total_ms: Math.round(end - start),
    approx_tokens: tokenCount,
    tokens_per_sec: (tokenCount / ((end - firstTokenAt) / 1000)).toFixed(1)
  };
}

Run this 20–30 times per model with the same prompt, throw out the first two runs (cold connection), and take the median — not the average. Latency distributions are right-skewed; a single slow outlier will distort a mean but barely moves a median.

Variables that skew your results

Most "Claude API is slow" complaints trace back to one of these, not the model itself:

Comparing across gateways and proxies

If you're testing Claude access through a proxy, gateway, or an API wrapper instead of calling Anthropic directly, add one more measurement: the overhead the gateway itself introduces. This is usually small (10–50ms) if the gateway is just forwarding requests and streaming responses through, but it can be large if the gateway buffers the full response before returning it — which defeats the purpose of streaming entirely.

This matters if you're evaluating a service like SubToAPI, which turns your existing Claude access into an HTTPS API with sub_live_... keys. SubToAPI streams responses through without buffering, so TTFT overhead versus calling Claude directly is minimal — you're mainly paying the cost of one extra network hop. When benchmarking any gateway, run the same script above against both the gateway and the direct API with identical prompts and max_tokens, then diff the medians. See the streaming docs for the exact event format, and the quickstart if you want to reproduce these numbers on your own account — a free trial is available at signup.

A reasonable benchmark checklist

FAQ

Is Claude API latency consistent across models? No. Smaller models like Haiku have lower TTFT and higher throughput; larger models like Opus have higher TTFT but strong per-token throughput once streaming begins. Always benchmark the specific model you plan to use in production.

Why does my benchmark show wildly different numbers each run? Latency distributions are right-skewed — occasional slow outliers are normal. Run at least 20 samples and report the median, not the average, and rule out cache misses and network variance before concluding anything changed.

Does routing through a gateway or proxy add significant latency? It depends on whether the gateway streams responses through in real time or buffers them first. A pass-through streaming gateway typically adds under 50ms; a buffering one can add the full completion time to your perceived TTFT. Test both configurations directly against your own traffic pattern.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →