← Blog

Claude API Multi-Agent Orchestration Example

2026-10-03 · 5 min read · SubToAPI Team

What multi-agent orchestration means with the Claude API

Multi-agent orchestration is a pattern where instead of sending one prompt to Claude and getting one answer, you split a task across several Claude calls that play different roles — a planner, one or more workers, and a reviewer — and you write code that routes messages between them. The Claude API doesn't have a built-in "agent" object; orchestration is something you build on top of it using the Messages API, tool use, and your own control flow.

This article walks through a concrete example: a research-and-summarize pipeline with three agent roles, the code that wires them together, and the practical issues (retries, context limits, cost) you'll hit when you move this from a demo to production.

The architecture: planner, workers, reviewer

A common and simple orchestration pattern looks like this:

  1. Planner agent — takes the user's request and breaks it into subtasks.
  2. Worker agents — each subtask is handled by a separate Claude call, often with tool access (web search, a database, a calculator).
  3. Reviewer agent — takes the worker outputs and produces a single, checked final answer.

Each "agent" here is just a Claude API call with a specific system prompt and a specific slice of context. The orchestration logic — deciding who runs when, what gets passed forward, what happens on failure — lives in your application code, not inside the model.

Step 1: the planner call

The planner's job is to output a structured list of subtasks. Forcing structured output with tool use makes this reliable:

const planResponse = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
    "content-type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-opus-4-20250514",
    max_tokens: 500,
    tools: [{
      name: "create_plan",
      description: "Return a list of subtasks for the research request.",
      input_schema: {
        type: "object",
        properties: {
          subtasks: { type: "array", items: { type: "string" } },
        },
        required: ["subtasks"],
      },
    }],
    tool_choice: { type: "tool", name: "create_plan" },
    messages: [
      { role: "user", content: "Research the competitive landscape for low-code API gateways and summarize pricing models." },
    ],
  }),
});

const plan = await planResponse.json();
const subtasks = plan.content[0].input.subtasks;

Forcing tool_choice to a specific tool guarantees you get parsed JSON back instead of free text you need to regex out of a sentence. This is the backbone of any orchestration layer — if the planner's output format is unreliable, everything downstream breaks.

Step 2: fan out to worker agents

Each subtask becomes its own Claude call, run in parallel. Workers should get a narrow system prompt so they don't try to solve the whole problem themselves:

async function runWorker(subtask) {
  const res = await fetch("https://api.anthropic.com/v1/messages", {
    method: "POST",
    headers: {
      "x-api-key": process.env.ANTHROPIC_API_KEY,
      "anthropic-version": "2023-06-01",
      "content-type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-20250514",
      max_tokens: 800,
      system: "You are a research worker. Answer only the specific subtask given. Be factual and concise. Do not speculate beyond the task.",
      messages: [{ role: "user", content: subtask }],
    }),
  });
  const data = await res.json();
  return data.content[0].text;
}

const results = await Promise.all(subtasks.map(runWorker));

Running workers with Promise.all is the main latency win of this pattern — three subtasks that would take 9 seconds sequentially finish in roughly 3 seconds in parallel, assuming you're not rate-limited.

Step 3: reviewer agent merges and checks

The reviewer gets all worker outputs as context and is told explicitly to cross-check them against each other, not just summarize blindly:

const reviewResponse = await fetch("https://api.anthropic.com/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": process.env.ANTHROPIC_API_KEY,
    "anthropic-version": "2023-06-01",
    "content-type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-opus-4-20250514",
    max_tokens: 1200,
    system: "You are a reviewer. You will receive multiple research findings. Merge them into one coherent report. Flag any contradictions between sources instead of silently resolving them.",
    messages: [{
      role: "user",
      content: results.map((r, i) => `Finding ${i + 1}:\n${r}`).join("\n\n"),
    }],
  }),
});

Using a stronger model for the planner and reviewer (where mistakes compound) and a cheaper/faster model for workers is a common cost optimization in this pattern — not every agent needs the same model.

Failure handling you need in production

A demo script can skip error handling; an orchestration pipeline cannot, because a single failed worker call can silently produce a corrupted final report.

Where billing and keys get complicated

A multi-agent pipeline makes 3-10x more API calls per user request than a single-prompt app, which means usage tracking, rate limits, and per-team cost visibility matter more here than in a simple chatbot wrapper. If you're building this on top of a Claude Pro or Team subscription rather than metered API billing, you also need a way to expose that access as a normal HTTPS API with real keys.

That's the specific problem SubToAPI solves: it turns your existing Claude access into an API with sub_live_... keys, so each agent role in your pipeline can call https://api.subtoapi.app/v1/messages the same way it would call Anthropic's endpoint directly, with streaming and tool use supported and usage visible per key in one dashboard. For a multi-agent setup specifically, this matters because you can issue a separate key per agent role or per environment (dev planner key, prod worker key, etc.) and see usage broken down that way instead of one undifferentiated total. Start with the quickstart, check tool use for the planner/reviewer pattern above, and compare plans if you're running this across a team.

questions

Do I need a special "agent" API or framework to orchestrate multiple Claude calls? No. The Messages API and tool use are enough — orchestration is control flow in your own code (sequencing, parallelizing, passing outputs between calls), not a separate Claude feature.

Should every agent in the pipeline use the same Claude model? Not necessarily. Many teams use a stronger model for planning and review steps, where errors compound, and a faster/cheaper model for narrow worker tasks, to cut latency and cost.

What's the biggest failure mode in multi-agent pipelines? Context bloat and silent partial failures — passing too much history into every agent, and not handling a single failed worker call before it reaches the reviewer, which produces confidently wrong final output.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →