Claude API Parallel Tool Use: A Working Example
What parallel tool use actually means
When you give Claude multiple tools in a single request, it doesn't always call them one at a time. If the model decides it needs two or more independent pieces of information to answer a prompt, it can return several tool_use blocks in a single response — all in one turn, instead of one tool call, a round trip, another tool call, another round trip. That's parallel tool use.
The practical benefit is latency and simplicity: instead of chaining multiple request/response cycles with Claude, you execute all the requested tools (often concurrently, since they're usually unrelated), send all the results back in one message, and get a final answer. Below is a complete, working example showing exactly how this looks in the Claude Messages API, from request to execution to final response.
The scenario
Suppose you're building an assistant that can check weather and look up stock prices. A user asks:
"What's the weather in Berlin and what's Apple's current stock price?"
These two sub-tasks have nothing to do with each other, so Claude should call both tools in the same turn rather than sequentially.
Defining the tools
[
{
"name": "get_weather",
"description": "Get the current weather for a city",
"input_schema": {
"type": "object",
"properties": {
"city": { "type": "string" }
},
"required": ["city"]
}
},
{
"name": "get_stock_price",
"description": "Get the current stock price for a ticker symbol",
"input_schema": {
"type": "object",
"properties": {
"ticker": { "type": "string" }
},
"required": ["ticker"]
}
}
]
The initial request
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"tools": [ /* tool definitions above */ ],
"messages": [
{
"role": "user",
"content": "What is the weather in Berlin and what is Apple'\''s current stock price?"
}
]
}'
What Claude returns
If Claude decides to use both tools, the response content array contains two tool_use blocks in the same assistant turn:
{
"role": "assistant",
"content": [
{ "type": "text", "text": "I'll check both of those for you." },
{
"type": "tool_use",
"id": "toolu_01A",
"name": "get_weather",
"input": { "city": "Berlin" }
},
{
"type": "tool_use",
"id": "toolu_01B",
"name": "get_stock_price",
"input": { "ticker": "AAPL" }
}
],
"stop_reason": "tool_use"
}
Notice there's no separate API call between the two tool invocations. Claude emitted both at once.
Executing both tools and sending results back
You run both tool functions — ideally in parallel since they're independent — and then send all results back in a single user message, matching each tool_result to its tool_use_id:
const toolUseBlocks = response.content.filter(b => b.type === "tool_use");
const results = await Promise.all(
toolUseBlocks.map(async (block) => {
if (block.name === "get_weather") {
return { tool_use_id: block.id, output: await getWeather(block.input.city) };
}
if (block.name === "get_stock_price") {
return { tool_use_id: block.id, output: await getStockPrice(block.input.ticker) };
}
})
);
const toolResultMessage = {
role: "user",
content: results.map(r => ({
type: "tool_result",
tool_use_id: r.tool_use_id,
content: JSON.stringify(r.output)
}))
};
You then append the assistant's message and this new user message to the conversation history and make one more request. Claude reads both results and produces a single, combined final answer:
"Berlin is currently 14°C with light rain. Apple (AAPL) is trading at $227.40."
That's the entire pattern: one request, one response with multiple tool_use blocks, parallel execution on your end, one tool_result message with multiple blocks, one final request.
Controlling parallel tool use
By default, Claude decides on its own whether to call one tool or several. If your application logic can't handle multiple simultaneous tool calls (for example, a stateful workflow where tool B depends on tool A's result), you can force sequential behavior with tool_choice:
"tool_choice": { "type": "auto", "disable_parallel_tool_use": true }
Setting disable_parallel_tool_use to true tells Claude to issue at most one tool_use block per turn, which is useful when tools have ordering dependencies you don't want the model guessing about.
Common mistakes with parallel tool use
- Returning results in separate messages. All
tool_resultblocks for a given turn must go in one user message, not multiple messages. Claude expects to see every pending tool call resolved before continuing. - Mismatched
tool_use_id. Each result must reference the exactidfrom its correspondingtool_useblock — not the tool name. - Assuming tools execute in order. If two tools are independent, don't assume Claude calls them left-to-right for a reason; treat the order as arbitrary and safe to parallelize.
- Forgetting partial failures. If one tool call fails, you still need to return a
tool_resultfor it (often with an error message as content) — you can't just omit it.
Where SubToAPI fits in
If you're exposing this kind of tool-calling flow to your own users — through an app, an internal tool, or a product built on top of Claude — you still need API keys, usage tracking per customer, and streaming support without building that infrastructure yourself. SubToAPI turns your Claude access into a standard HTTPS API with application-level keys (sub_live_...), full support for tool use and streaming, and per-key usage metadata out of the box. The request/response shapes for tools are the same ones shown above — see the tool use docs and the quickstart for details. Plans start at €9/month with a free trial at signup; full breakdown on pricing.
Questions
Does parallel tool use cost more tokens than sequential calls? Not meaningfully more per call, since the tool definitions and results are similar size either way. You actually save tokens overall because you avoid repeating the full conversation history across multiple round trips.
Can I force Claude to always call tools in parallel? No — Claude decides based on the prompt and available tools. You can only disable parallel calls (via disable_parallel_tool_use), not force them when the model determines only one tool is needed.
What happens if one of several parallel tool calls fails? You still return a tool_result block for the failed call, typically with an error message as its content, alongside the successful results from the other calls, all in the same message.