Claude API Multi-Agent System Architecture Guide
A multi-agent system built on the Claude API is a set of independent Claude calls, each with a narrow role, that pass structured messages to each other under the coordination of an orchestrator. Instead of one giant prompt trying to plan, research, write, and verify, you split those jobs into separate agents that each get a focused system prompt, a limited tool set, and a clear input/output contract.
This matters because single-prompt agents degrade fast on long, multi-step tasks — context gets crowded, tool calls get tangled, and errors compound silently. A well-architected multi-agent system isolates failure, makes each step debuggable, and lets you scale different agents independently (a cheap classifier agent can run far more often than an expensive research agent). Below is a practical breakdown of the architecture patterns that actually work with Claude, plus the design decisions you'll need to make early.
Core Architecture Patterns
Most production multi-agent systems on Claude fall into one of four shapes:
- Orchestrator-worker: A single "manager" agent decomposes a task, dispatches subtasks to specialized worker agents, and merges results. This is the most common pattern and the easiest to reason about.
- Pipeline: Agents run in a fixed sequence — researcher → writer → editor — where each stage's output is the next stage's input. Simple to build, but rigid if a step needs to loop back.
- Debate/consensus: Two or more agents critique each other's output before a final answer is accepted. Useful for high-stakes reasoning tasks where you want an adversarial check.
- Hierarchical: A top-level orchestrator delegates to mid-level orchestrators, which delegate to leaf workers. Needed once you have more than 5–6 distinct roles and a flat orchestrator becomes a bottleneck.
Start with orchestrator-worker unless you have a specific reason not to. It maps cleanly onto Claude's tool-use model: the orchestrator can literally call each worker as a "tool."
Designing the Orchestrator
The orchestrator's job is decomposition and routing, not doing the work itself. Keep its system prompt tight:
You are a task router. Given a user request, decide which
specialist agents are needed and in what order. Do not answer
the request yourself. Output a JSON plan with steps, each
naming an agent and the input it should receive.
Treat each worker as a callable tool using Claude's tool-use format (see /docs/tools for the request shape). This gives you two benefits: the orchestrator's decisions are structured and parseable, and you get a clean audit trail of which agent was invoked with what input.
{
"name": "research_agent",
"description": "Gathers factual information on a topic before writing.",
"input_schema": {
"type": "object",
"properties": {
"topic": { "type": "string" },
"depth": { "type": "string", "enum": ["quick", "deep"] }
},
"required": ["topic"]
}
}
When the orchestrator calls research_agent, your application code intercepts that tool call, runs it as a separate Claude request with its own system prompt and message history, and feeds the result back into the orchestrator's conversation as a tool result.
State and Message Passing
Each agent should have its own message history — don't let the orchestrator's full transcript leak into worker context by default. Pass only what the worker needs:
- Task input: the specific subtask, not the whole conversation.
- Relevant prior outputs: summarized, not raw, if they're long.
- Constraints: format requirements, tone, length limits.
Keep a shared "task state" object in your application layer (not inside any single agent's context) that tracks step status, outputs, and retries. This is where you catch loops — if an agent produces the same output twice or a worker gets called more than N times, halt and escalate to a human or a fallback path.
A Minimal Working Example
async function callAgent(role, systemPrompt, userInput) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
system: systemPrompt,
messages: [{ role: "user", content: userInput }],
max_tokens: 1024
})
});
return res.json();
}
async function runPipeline(topic) {
const research = await callAgent(
"researcher",
"You gather concise, factual notes on the given topic.",
topic
);
const draft = await callAgent(
"writer",
"You write a clear draft using the provided research notes.",
research.content[0].text
);
return callAgent(
"editor",
"You tighten prose and check factual consistency against the notes.",
draft.content[0].text
);
}
This is a pipeline, not a full orchestrator, but it shows the core idea: each agent is a separate, stateless call with a narrow prompt. Because every agent authenticates with its own application key, you can track usage per agent role instead of per user — useful when one agent (say, a research step with deep tool use) is far more expensive than the others. SubToAPI issues these as sub_live_... keys so each agent or service in your architecture gets its own key and usage line in the dashboard; see /docs/quickstart to generate one.
Streaming and Concurrency
Worker agents that run independently (not in sequence) should run concurrently — a fan-out research step across three topics doesn't need to be serial. Use Promise.all for independent calls, but cap concurrency deliberately; unbounded fan-out against a shared API key can hit rate limits and makes cost unpredictable.
For the final agent in a chain that's user-facing (the writer or the summarizer step), stream the response so users see output as it's generated rather than waiting for the whole pipeline to finish. See /docs/streaming for the event format — internal agent-to-agent calls generally don't need streaming since nothing is watching them in real time. Full request/response shapes for both streaming and non-streaming calls are documented at /docs/messages.
Common Pitfalls
- Over-decomposition: five agents for a task one well-prompted call could do adds latency and failure surface for no benefit.
- Unbounded loops: debate/consensus patterns need a hard iteration cap, or two agents will happily disagree forever.
- Context bleed: passing full transcripts between agents instead of summaries inflates token cost and confuses each agent's focused role.
- No cost visibility per agent: without per-key or per-agent usage tracking, you can't tell which role is driving your bill. If you're running this across a team, /pricing shows how seat-based plans handle multiple agents or services sharing a workspace.
FAQ
Do I need separate API keys for each agent?
Not strictly, but it's the easiest way to get per-agent cost and usage visibility. One key per role or service, tracked separately in your dashboard, makes debugging expensive agents much faster than digging through a single combined log.
How many agents is too many?
There's no fixed number, but each added agent introduces latency and a new failure point. If your task can be reliably done by one well-scoped prompt with tool use, don't split it. Reach for multi-agent when steps genuinely need different context, tools, or expertise.
Should agents share conversation history?
Generally no. Pass summarized inputs and outputs between agents rather than full transcripts. This keeps each agent's context focused, reduces token cost, and avoids one agent's mistakes or verbosity contaminating the next one's reasoning.