← Blog

Claude API Multi-Agent Orchestration Setup Guide

2026-09-30 · 5 min read · SubToAPI Team

Setting up multi-agent orchestration with the Claude API means running several Claude instances, each with a distinct role and system prompt, coordinated by a controller that routes messages, manages shared state, and merges results. There's no special "multi-agent" endpoint — you build orchestration as an application layer on top of standard Claude API calls, using patterns like a manager agent that delegates to worker agents, or a pipeline where each agent transforms the previous agent's output.

This guide walks through a concrete setup: how to structure roles, route requests, handle shared context, and avoid the failure modes that make multi-agent systems expensive and hard to debug.

What multi-agent orchestration actually means here

A single Claude call is one request, one system prompt, one conversation. Multi-agent orchestration means running multiple independent (or semi-independent) Claude conversations that each specialize in a task, then combining their outputs programmatically. Common patterns:

All of these are orchestration logic in your own code (Node.js, Python, whatever), not something Claude does for you automatically.

Step 1: Define agent roles with separate system prompts

Each agent should have a narrow, well-defined system prompt. Vague, overlapping roles are the number one cause of multi-agent systems producing redundant or contradictory work.

const AGENTS = {
  planner: {
    system: "You break a user task into 3-5 concrete subtasks. Output only a numbered list, no commentary.",
    model: "claude-sonnet-4"
  },
  coder: {
    system: "You write code for a single subtask. Output only code and a one-line explanation.",
    model: "claude-sonnet-4"
  },
  reviewer: {
    system: "You review code for bugs and security issues. Output a list of concrete issues or 'LGTM'.",
    model: "claude-sonnet-4"
  }
};

Keep each system prompt focused on one job. If an agent's prompt needs "and also handle edge case X," that's usually a sign you need a fourth agent, not a longer prompt.

Step 2: Build the orchestrator as a state machine

The orchestrator is just code that decides which agent runs next and what context it receives. A simple sequential pipeline:

async function runPipeline(task) {
  const plan = await callAgent("planner", task);
  const subtasks = parsePlan(plan);

  const results = [];
  for (const subtask of subtasks) {
    const code = await callAgent("coder", subtask);
    const review = await callAgent("reviewer", code);
    results.push({ subtask, code, review });
  }
  return results;
}

async function callAgent(role, input) {
  const agent = AGENTS[role];
  const res = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: agent.model,
      system: agent.system,
      max_tokens: 1024,
      messages: [{ role: "user", content: input }]
    })
  });
  const data = await res.json();
  return data.content[0].text;
}

For non-linear orchestration (manager deciding which worker to call next), swap the loop for a router function that inspects the manager's output and dispatches accordingly. Keep the routing logic deterministic where possible — parse structured output (JSON or numbered lists) rather than trying to interpret free text, since free-text parsing is where orchestration silently breaks.

Step 3: Manage shared state and context passing

Agents don't automatically share context. Decide explicitly what each agent sees:

A practical middle ground is a shared state object your orchestrator updates after each agent call, then serializes into the next agent's prompt:

let state = { task, decisions: [], issues: [] };

function buildContext(state) {
  return `Current state:\n${JSON.stringify(state, null, 2)}`;
}

This keeps token usage bounded and gives you a clear audit trail of what each agent knew at each step.

Step 4: Handle failures and loops

Multi-agent systems fail in specific, predictable ways:

If you're running this in production, put retries and timeouts around every agent call individually, not around the whole pipeline — a single agent timing out shouldn't force you to restart the entire orchestration from scratch.

Where SubToAPI fits

Multi-agent setups typically mean many concurrent API calls across agents, which makes usage visibility important — you want to know which agent role is burning tokens, not just a total bill. SubToAPI turns your Claude access into a standard HTTPS API with sub_live_... application keys, so each agent (or each team member building one) can have its own key with separate usage metadata, while streaming and tool use work the same way they do with a direct integration. See the quickstart to get a key, messages docs for the request format used in the examples above, and tools docs if your agents need function calling. Plans start at €9/month with a free trial — check pricing or sign up directly.

FAQ

Do I need a special API or SDK for multi-agent orchestration with Claude? No. Orchestration is application logic you write yourself — a controller that makes multiple standard Claude API calls and routes results between them. There's no dedicated multi-agent endpoint.

How many agents should a typical setup use? Start with 2-3 (e.g., planner, worker, reviewer). Each additional agent adds latency, cost, and failure surface, so only add one when a single agent's role is genuinely overloaded.

How do I keep token costs under control with multiple agents? Pass summarized state instead of full conversation history where possible, set max iteration limits per subtask, and log token usage per agent so you can spot the expensive one early.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →