Claude API Tool Use Multi-Step Workflow Guide
Building a multi-step tool use workflow with the Claude API means handling the loop where Claude calls a tool, you execute it, you send the result back, and Claude decides whether to call another tool or respond to the user. This is different from a single tool call — the model has to maintain context across several rounds, chain outputs from one tool into inputs for the next, and know when to stop. Most of the complexity isn't in the Claude API call itself, it's in the orchestration loop you write around it.
This guide walks through the actual mechanics of a multi-step tool workflow: how the conversation state grows, how to detect when Claude wants to keep working versus when it's done, and the common failure modes that break long tool chains.
What makes a workflow "multi-step"
A single tool call is: user asks a question, Claude requests a tool, you run it, you return the result, Claude answers. A multi-step workflow repeats that cycle multiple times before producing a final answer — for example:
- Claude calls
search_ordersto find a customer's order ID - Claude calls
get_shipping_statususing that order ID - Claude calls
send_emailto notify the customer - Claude returns a final text response summarizing what happened
Each step depends on the output of the previous one. The model isn't just picking tools — it's reasoning about what it learned from each result before deciding the next action.
The core loop
Regardless of provider, the pattern is the same: you keep calling the Messages API in a loop, appending the assistant's tool use and your tool results to the conversation, until the response contains no more tool calls.
let messages = [
{ role: "user", content: "Find the shipping status for order #4821 and email the customer." }
];
while (true) {
const response = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
tools: toolDefinitions,
messages
})
}).then(r => r.json());
messages.push({ role: "assistant", content: response.content });
const toolCalls = response.content.filter(b => b.type === "tool_use");
if (toolCalls.length === 0) {
// Final answer reached
console.log(response.content.find(b => b.type === "text")?.text);
break;
}
const toolResults = [];
for (const call of toolCalls) {
const result = await runTool(call.name, call.input);
toolResults.push({
type: "tool_result",
tool_use_id: call.id,
content: JSON.stringify(result)
});
}
messages.push({ role: "user", content: toolResults });
}
The stop_reason field on each response tells you why the model stopped generating. When it's tool_use, there's work left to do. When it's end_turn (or similar), the model is giving its final answer and the loop should exit.
Tracking state across steps
As the workflow grows, the messages array becomes the entire memory of the task. There's no separate "state" object — everything Claude knows about what happened in step 1 comes from what you put back into the conversation in step 2. A few practical consequences:
- Keep tool results concise. If a tool returns a 50KB JSON blob, trim it before sending it back. Every step re-sends the full conversation history, so bloated results slow every subsequent call and burn tokens.
- Don't silently drop failures. If a tool call errors, return that error as the
tool_resultcontent instead of throwing. Claude can often recover — retry with different input, try another tool, or tell the user it failed — but only if it sees the failure. - Cap the number of steps. Add a hard iteration limit (5–10 is reasonable for most workflows) to prevent infinite loops where the model keeps requesting tools without converging. If you hit the limit, return a message asking Claude to summarize what it has so far.
Handling branching logic
Multi-step workflows often need Claude to choose different paths depending on earlier results — for example, only sending an email if the shipping status is "delayed." You don't need special logic for this on your side; it's handled by writing clear tool descriptions and letting the model's own reasoning branch naturally. The orchestration code stays identical regardless of which path Claude takes — it just keeps feeding tool results back until Claude stops calling tools.
Where you do need extra logic is in deciding which tools are available at each step. If a workflow has a destructive action (deleting a record, sending a payment), consider only exposing that tool after a prior confirmation step, rather than relying purely on the model's judgment in every case.
Debugging long tool chains
When a multi-step workflow misbehaves, the fastest way to debug it is to log the full messages array after each round and inspect exactly what Claude saw before making its next decision. Common issues:
- A tool result was returned in the wrong format (not matching the JSON shape described in the tool's schema), causing Claude to misinterpret it
- The
tool_use_idin your result didn't match the id in Claude's request, breaking the pairing - Token limits were hit mid-chain because results weren't trimmed, truncating earlier context
If you're running this in production, request logging and usage metadata per call help a lot here — SubToAPI's dashboard shows token usage and timing for every request, which makes it easier to spot which step in a chain is expensive or slow without adding your own instrumentation. See the tool use docs for the full request/response shape, and the quickstart if you're setting up API access for the first time.
Keeping it maintainable
As workflows grow past 3–4 steps, it helps to:
- Keep each tool narrowly scoped (one clear action) rather than building "do everything" tools
- Log
stop_reasonand tool names at each step for observability - Set a sensible
max_tokensper call — see our guide on the parameter if you're unsure — so partial responses don't get cut off mid-reasoning - Version your tool schemas, since changing an input shape mid-deployment can break in-flight conversations
None of this requires a framework. A while loop, a switch statement for tool execution, and careful logging cover the vast majority of real-world multi-step tool workflows.
Questions
Does Claude decide the whole workflow upfront, or one step at a time? One step at a time. Claude sees the conversation so far and decides the next single action — it doesn't plan all steps in advance, so your loop must call it again after each tool result.
How many tool calls can a single workflow chain together? There's no hard limit from the API, but practically you should cap iterations (5–10 is typical) to avoid runaway loops and control token usage, since the full history resends each round.
Can two workflows run in parallel for different users? Yes — each conversation is just a separate messages array and independent API call, so concurrent workflows don't interfere as long as you keep state scoped per request or per user session.