Claude Tool Use Parallel Execution: How It Works
When Claude decides a task needs more than one tool, it doesn't always call them one at a time. In a single assistant turn, Claude can return several tool_use blocks back to back — for example, checking weather in three cities or querying two different APIs — and expects you to execute them and send all the results back together. This is parallel tool execution, and it's a native behavior of the Claude API's tool use feature, not something you configure separately.
The short answer to "how do I get parallel tool calls with Claude": you don't explicitly request it. You give Claude a set of tool definitions, ask a question that reasonably requires multiple independent calls, and Claude's response may contain multiple tool_use content blocks in one message. Your job is to detect all of them, run them (ideally concurrently), and send back a matching set of tool_result blocks in a single user message before continuing the conversation.
How Claude Decides to Call Tools in Parallel
Claude evaluates whether the sub-tasks in your request are independent of each other. If a user asks "what's the weather in Paris, Tokyo, and New York," there's no reason to wait for one result before requesting the next, so Claude often emits three tool_use blocks in the same response. If the tasks are sequential — where one tool's output is needed as input to the next — Claude will call them one at a time across multiple turns instead.
You can't force parallel execution with a parameter, but you can encourage it:
- Give each tool a narrow, single-purpose definition rather than one do-everything tool.
- Phrase prompts so independence is clear ("for each of these three tickers, fetch the current price").
- Avoid system prompts that tell Claude to "do one thing at a time," which discourages batching.
If you specifically need to prevent parallel calls (for example, a tool that must never run concurrently due to rate limits), set tool_choice to force a single tool or add explicit instructions in your system prompt that Claude should call tools sequentially and wait for each result.
Structuring the Response for Multiple Tool Calls
A Claude response with parallel tool use looks like this — multiple tool_use blocks inside one content array:
{
"role": "assistant",
"content": [
{ "type": "text", "text": "I'll check all three cities." },
{ "type": "tool_use", "id": "toolu_01A", "name": "get_weather", "input": { "city": "Paris" } },
{ "type": "tool_use", "id": "toolu_01B", "name": "get_weather", "input": { "city": "Tokyo" } },
{ "type": "tool_use", "id": "toolu_01C", "name": "get_weather", "input": { "city": "New York" } }
]
}
Your code has to iterate the full content array, not just grab the first tool_use block. A common bug is handling only one tool call and silently dropping the rest, which causes Claude's next turn to hallucinate or error because it's waiting on results that never arrive.
Executing Tool Calls Concurrently
Once you've extracted all tool_use blocks, run them in parallel with Promise.all (or your language's equivalent) rather than looping with await on each one sequentially:
const toolCalls = response.content.filter(b => b.type === "tool_use");
const results = await Promise.all(
toolCalls.map(async (call) => {
const output = await executeTool(call.name, call.input);
return {
type: "tool_result",
tool_use_id: call.id,
content: JSON.stringify(output)
};
})
);
The tool_use_id in each tool_result must match the id from the corresponding tool_use block. Order within the array doesn't need to match the original sequence, but every tool_use block must get exactly one tool_result — Claude's API will reject the next turn if any are missing.
Sending Results Back in One Message
All tool_result blocks go into a single user message, appended after the assistant's message with the tool calls:
messages.push(response); // the assistant message with tool_use blocks
messages.push({ role: "user", content: results });
const next = await client.messages.create({
model: "claude-sonnet-4-5",
max_tokens: 1024,
tools,
messages
});
This is true whether you executed the tools in parallel or sequentially on your end — the API only cares that the message structure pairs each tool_use_id with a result before the conversation continues.
Error Handling Across Parallel Calls
If one tool call fails while others succeed, don't drop it — return a tool_result with is_error: true and a description of what went wrong:
{
"type": "tool_result",
"tool_use_id": "toolu_01B",
"content": "Rate limit exceeded for weather API",
"is_error": true
}
Claude will see the error alongside the successful results and can decide whether to retry, apologize to the user, or proceed with partial data. This matters more with parallel execution because a single failed request among several shouldn't block the whole response — handle each tool call's success and failure independently, ideally with Promise.allSettled if you want to avoid one rejection canceling the batch.
Building This on Top of an API Gateway
If you're routing Claude access through SubToAPI — turning your Claude subscription into an HTTPS API with sub_live_... keys — the same tool use mechanics apply, since the request and response format for tools matches Claude's Messages API. You define tools, send them with Authorization: Bearer $SUBTOAPI_KEY, and parse multiple tool_use blocks the same way. See the tool use docs for the exact request shape, and the streaming docs if you're combining parallel tool calls with streamed text. For teams running several agents or services against the same account, SubToAPI's usage metadata also helps you see which tool-heavy requests are driving cost, without changing how you build the tool execution logic itself.
questions
Does Claude always call tools in parallel when it could? No. Parallel tool use depends on the model judging the sub-tasks as independent. Ambiguous prompts or tools with overlapping purposes often get called sequentially across turns instead.
Can I limit Claude to only one tool call per turn? Yes, use tool_choice to force a specific tool, or instruct the model explicitly in the system prompt to call tools one at a time and wait for each result before proceeding.
What happens if I only handle one tool_use block in a multi-call response? The conversation breaks — Claude expects a tool_result for every tool_use_id it emitted. Missing results cause errors or malformed follow-up responses on the next turn.