Claude API Tool Use Best Practices for Reliable Agents
Tool use (also called function calling) lets Claude decide when to invoke code you control — a database lookup, a search API, a calculator — instead of guessing an answer. Getting reliable behavior out of it isn't about picking the "right" model; it's about how you design the tool schema, structure the conversation loop, and handle the edge cases where Claude asks for something malformed or unnecessary.
This guide covers the practical decisions that separate a tool-use integration that works in a demo from one that survives production traffic: schema design, error handling, multi-tool orchestration, and cost control.
Design tool schemas like you're writing a function signature for a stranger
Claude only knows what your tool does from the name, description, and JSON schema you provide. Treat the description as documentation for a developer who has never seen your codebase.
Do this:
- Name tools with verbs:
search_orders, notorders. - Describe what the tool returns, not just what it accepts. "Returns the 5 most recent orders for a customer, sorted by date" is far more useful than "Gets orders."
- Mark required fields explicitly in the schema and keep optional fields genuinely optional — don't require parameters your backend can default.
- Use
enumfor constrained values (status codes, categories) instead of free-text strings. This alone eliminates a large class of malformed calls.
Avoid this:
- Vague descriptions like "Handles user requests." Claude will call it at the wrong times or with wrong arguments.
- Overloaded tools that do five things based on a
modeparameter. Split them — Claude picks the right tool far more reliably than it picks the right mode. - More than 10–15 tools active at once in a single request. If you have a large tool library, filter to the subset relevant to the current task before sending the request.
{
"name": "search_orders",
"description": "Returns up to 5 recent orders for a given customer_id, sorted by date descending. Use this when the user asks about order history, status, or tracking.",
"input_schema": {
"type": "object",
"properties": {
"customer_id": { "type": "string", "description": "The customer's UUID" },
"status": {
"type": "string",
"enum": ["pending", "shipped", "delivered", "cancelled"]
}
},
"required": ["customer_id"]
}
}
Handle the tool_use loop correctly
A tool call is not the end of a request — it's a round trip. Claude responds with a stop_reason of tool_use, you execute the tool, and you send the result back as a tool_result block in a new user message with the same tool_use_id. Skipping or mishandling any part of this breaks the conversation.
let response = await client.messages.create({
model: "claude-sonnet-4",
max_tokens: 1024,
tools: [searchOrdersTool],
messages: [{ role: "user", content: "Where's my last order?" }]
});
while (response.stop_reason === "tool_use") {
const toolUse = response.content.find(b => b.type === "tool_use");
const result = await runTool(toolUse.name, toolUse.input);
messages.push({ role: "assistant", content: response.content });
messages.push({
role: "user",
content: [{
type: "tool_result",
tool_use_id: toolUse.id,
content: JSON.stringify(result)
}]
});
response = await client.messages.create({
model: "claude-sonnet-4",
max_tokens: 1024,
tools: [searchOrdersTool],
messages
});
}
A few things that trip people up:
- Always echo the full assistant message back, including any text content alongside the
tool_useblock. Truncating it confuses the model on the next turn. - Set
is_error: trueon atool_resultwhen the tool call failed, rather than stuffing an error string into a normal result. This tells Claude to reason about the failure instead of treating it as valid data. - Cap the loop. Add a max iteration count (5–8 is usually plenty) so a confused model can't spiral into repeated tool calls and burn tokens indefinitely.
Let Claude parallelize when it makes sense
Claude can request multiple tool calls in a single turn if the tasks are independent — for example, checking inventory and pricing at the same time. Execute these concurrently and return all tool_result blocks together in one message. Don't force sequential round trips for tools that don't depend on each other; it adds latency for no benefit.
Validate everything, trust nothing
Claude generates plausible-looking arguments, not guaranteed-valid ones. Always validate tool_use.input against your own schema or type system before executing — especially for anything that touches a database, file system, or external API with side effects. If a required field is missing or malformed, return a tool_result with is_error: true and a clear message. Claude will usually self-correct and retry with better arguments on the next turn.
For destructive actions (deleting records, sending emails, charging cards), don't rely on the model deciding not to call the tool — gate the actual execution behind an explicit confirmation step in your application logic.
Keep system prompts and tool descriptions in sync
If your system prompt says "you can look up order status" but the tool is actually named get_order_details, Claude has to bridge that gap itself, which increases the chance of a wrong or missed call. Reference tools by their exact names in the system prompt when you want to nudge usage, and keep the two in the same document or generated from the same source so they don't drift.
Watch token and cost overhead
Tool definitions are sent as part of the prompt on every request, and tool_result payloads count against your context window. Large tool result blobs (full database rows, unfiltered API responses) inflate cost and can crowd out useful context. Trim results to what's actually relevant before sending them back.
If you're routing tool-use traffic through SubToAPI, usage metadata is tracked per key in the dashboard, so you can see which application or endpoint is driving token consumption from tool calls specifically — useful for spotting a tool that's returning oversized payloads. The tool use docs cover the request format SubToAPI expects, which mirrors the standard Messages API, and the streaming docs explain how tool_use events surface when you're streaming a response that includes a tool call.
questions
Do I need a different model for tool use? No — tool use works with any current Claude model via the Messages API. Model choice affects how well Claude selects the right tool and formats arguments, but it's not a separate feature or endpoint.
What happens if Claude calls a tool that doesn't exist? It shouldn't, if only tools you've defined are listed in the request. If you dynamically filter tools per request, make sure the filtered list matches what the system prompt references — otherwise Claude may hallucinate a call to a tool it saw earlier in the conversation history but that isn't currently available.
Should I always require confirmation before executing a tool? For read-only tools (lookups, searches), no — it adds friction for no safety benefit. For anything with a side effect (writes, payments, deletions), yes, gate execution behind explicit confirmation in your application, not just model judgment.