Claude API Multi-Agent System Design Guide
Designing a multi-agent system with the Claude API means splitting a complex task across multiple Claude calls, each playing a specialized role, and coordinating them with code rather than asking one giant prompt to do everything. This matters because single-prompt agents hit quality ceilings fast: they lose track of long instructions, mix up roles, and can't parallelize independent subtasks like research, drafting, and validation.
The practical answer is to treat agents as functions with a defined input/output contract, orchestrate them with a controller process (not with the model itself trying to "remember" the whole pipeline), and use tool calling to let agents fetch data or trigger actions in a structured way. Below is a breakdown of the patterns that actually work in production, plus the pitfalls that trip people up.
Core architecture patterns
There are three patterns worth knowing before you write any code.
Orchestrator-worker. A single "manager" agent breaks a task into subtasks and dispatches them to specialized worker agents (a researcher, a writer, a critic). The orchestrator doesn't do the work itself — it routes and aggregates. This is the most common pattern because it's easy to debug: each worker has a narrow system prompt and a small context window to reason over.
Pipeline. Agents run in a fixed sequence, each consuming the previous agent's output. Good for workflows like: extract → summarize → classify → format. No dynamic routing needed, which makes it cheaper and more predictable, but less flexible if the task shape varies.
Debate/critic loop. Two or more agents review each other's output — one drafts, another critiques against a rubric, and the loop repeats until the critic approves or a max-iteration limit hits. Useful for tasks where correctness matters more than speed (code review, compliance checks).
Most real systems combine these: an orchestrator-worker setup where one of the workers is itself a critic loop.
Designing the agent contract
Each agent should have:
- A single responsibility — "summarize this document," not "summarize, then decide what to do next, then format."
- A strict input schema — what the orchestrator passes in.
- A strict output schema — ideally enforced via tool use so you get structured JSON back instead of parsing prose.
- No shared memory with other agents unless you explicitly pass it. Agents should not assume context from a different agent's conversation.
Here's a minimal orchestrator pattern using sequential calls:
async function callAgent(systemPrompt, userMessage) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
system: systemPrompt,
messages: [{ role: "user", content: userMessage }]
})
});
const data = await res.json();
return data.content[0].text;
}
async function runPipeline(task) {
const research = await callAgent(
"You are a research agent. Extract key facts only, no commentary.",
task
);
const draft = await callAgent(
"You are a writing agent. Turn facts into a clear paragraph.",
research
);
const critique = await callAgent(
"You are a critic. List factual errors or unclear phrasing. Reply 'OK' if none.",
draft
);
return { draft, critique };
}
This is deliberately boring — plain async/await, no framework. Multi-agent systems fail more often from unclear responsibilities than from missing orchestration libraries.
Using tool use for structured handoffs
When agents need to pass data to each other or to external systems, don't rely on free-text parsing. Define tools and let Claude call them with validated arguments. This gives you a typed contract between agents instead of regex-scraping a response.
{
"name": "submit_research",
"description": "Submit structured research findings",
"input_schema": {
"type": "object",
"properties": {
"facts": { "type": "array", "items": { "type": "string" } },
"sources": { "type": "array", "items": { "type": "string" } }
},
"required": ["facts"]
}
}
When the research agent calls this tool, the orchestrator reads tool_use blocks directly instead of hoping the model formatted JSON correctly inside prose. See /docs/tools for the request/response shape.
State, retries, and cost control
Three things break multi-agent systems in production that don't show up in a demo:
- Context bloat. Passing full conversation history between every agent call multiplies token usage fast. Pass only the specific output each downstream agent needs, not the whole chain.
- Cascading failures. If one agent in a pipeline returns malformed output, the next agent either crashes or quietly hallucinates around it. Validate each agent's output against its schema before passing it forward, and fail loudly with a retry rather than letting garbage propagate.
- Cost explosion. Five agents each making a call per task means 5x the API calls of a single-prompt system. Track token usage per agent, not just per request, so you know which role is expensive and whether it actually needs a larger model.
If you're running this across a team — one person building the orchestrator, another tuning the critic agent — a shared API layer avoids everyone juggling separate credentials and losing visibility into who's spending what. SubToAPI turns your existing Claude access into application API keys (sub_live_...) with per-key usage metadata, so you can see token consumption broken down by agent role if you tag keys accordingly, and add teammates as seats instead of sharing one credential. Streaming and standard Messages API compatibility (see /docs/messages and /docs/streaming) mean the orchestration code above works unchanged — only the base URL and key differ. Check /pricing for Solo, Team, and Scale tiers, or start with a free trial at /signup.
Testing the system
Multi-agent systems need two layers of tests: unit tests per agent (does the research agent return facts in the right schema given a known input?) and integration tests on the full pipeline (does the end-to-end output meet quality bar X?). Run the unit tests against fixed inputs with assertions on schema, not on exact text — model outputs vary slightly between runs even at low temperature.
FAQs
Do I need a framework like LangGraph to build a multi-agent system with Claude? No. Frameworks help with complex branching logic, but a simple orchestrator-worker pattern is often just a few async functions calling the Messages API in sequence or in parallel with Promise.all.
How many agents is too many? There's no fixed number, but each additional agent adds latency and cost. Start with the minimum (often 2–3: do the task, check the task) and only split further when a single agent's responsibility is genuinely too broad to do well.
Should agents share conversation history? Generally no. Pass only the specific output one agent needs from another. Shared full history increases token cost and lets agents drift off their assigned role.