Build an AI Agent with Claude Tool Use: A Guide
What "building an AI agent" actually means with Claude
If you're searching for how to build an AI agent with Claude tool use, you're really asking one question: how do I let Claude decide when to call functions in my code, run them, and use the results to keep working toward a goal? That's the core of an agent — not a chatbot that answers one question, but a loop where the model can take actions, observe what happened, and take more actions until the task is done.
Claude's tool use (also called function calling) is the mechanism that makes this possible. You describe the tools available — a search function, a database query, a calculator, an API call — and Claude decides which one to invoke, with what arguments, based on the conversation. Your code executes the tool and feeds the result back. Repeat until Claude produces a final answer instead of another tool call. That loop is the entire architecture of a basic agent.
The three parts of an agentic tool-use setup
Every Claude-based agent has the same three moving pieces:
- Tool definitions — JSON schemas describing what each tool does and what arguments it takes.
- The agent loop — code that sends messages to Claude, checks if the response contains a tool call, executes it, and sends the result back.
- State/context — the growing conversation history that lets Claude remember what tools it already called and what they returned.
Here's a minimal tool definition for a weather lookup:
{
"name": "get_weather",
"description": "Get current weather for a city",
"input_schema": {
"type": "object",
"properties": {
"city": { "type": "string" }
},
"required": ["city"]
}
}
You pass an array of these tool definitions with every request. Claude reads them, decides if a tool is relevant, and if so, returns a tool_use content block instead of (or alongside) plain text.
Building the agent loop
The loop is the part people underestimate. It's not a single request-response — it's a while loop that keeps running until Claude stops asking for tools.
async function runAgent(userMessage, tools, executeTool) {
let messages = [{ role: "user", content: userMessage }];
while (true) {
const response = await callClaude(messages, tools);
const toolCalls = response.content.filter(b => b.type === "tool_use");
if (toolCalls.length === 0) {
// Claude gave a final answer, no more tools needed
return response.content.find(b => b.type === "text").text;
}
messages.push({ role: "assistant", content: response.content });
const toolResults = [];
for (const call of toolCalls) {
const result = await executeTool(call.name, call.input);
toolResults.push({
type: "tool_result",
tool_use_id: call.id,
content: JSON.stringify(result)
});
}
messages.push({ role: "user", content: toolResults });
}
}
A few things worth noting:
- Every
tool_useblock gets matched with atool_resultcarrying the sametool_use_id. Claude uses this to know which result belongs to which call. - The assistant's tool-call message must be pushed back into the conversation before the tool results, otherwise Claude loses track of what it asked for.
- The loop should have a hard iteration cap in production — an agent stuck in a call-tool-repeat cycle will burn tokens fast.
Designing tools Claude can actually use well
Agents fail less because of bad prompts and more because of vague tools. A few practical rules:
- One tool, one job. A tool called
manage_databasethat does five different things forces Claude to guess. Split it intoget_record,create_record,update_record. - Descriptions are prompts. The
descriptionfield is what Claude reads to decide relevance — write it like you're explaining the tool to a new engineer, not like a code comment. - Return structured, minimal results. Don't dump a raw API response into
tool_result. Strip it to what's relevant so Claude doesn't waste context re-parsing noise. - Make failures visible. If a tool call fails, return a
tool_resultdescribing the failure instead of throwing — Claude can often recover by trying a different approach.
Where infrastructure fits in
Once you're running an agent loop in production, you're not just calling Claude once — you're making several requests per user interaction, often with streaming output so the UI doesn't sit idle while a tool executes. This is where a lot of teams end up wiring together their own key rotation, usage tracking, and multi-project access instead of building the agent itself.
SubToAPI turns your existing Claude access into a standard HTTPS API with sub_live_ application keys, so each agent, service, or teammate gets its own key with its own usage visibility, without sharing raw credentials. Tool use, streaming, and message handling work the same way you'd expect from the underlying API — see the tool use docs, streaming docs, and messages reference for the exact request shapes. If you're prototyping an agent today, the quickstart gets you an API key in a few minutes, and plans start at Solo for solo builders and scale up to Team and Scale for multi-seat setups — check pricing or start with the free trial at signup.
A simple example end to end
Say you're building an agent that answers "what's the weather in the cheapest city to fly to from Berlin this weekend" — it needs a flights tool and a weather tool, called in sequence, with Claude deciding the order:
const tools = [flightsTool, weatherTool];
async function executeTool(name, input) {
if (name === "search_flights") return await searchFlights(input);
if (name === "get_weather") return await getWeather(input.city);
}
const answer = await runAgent(
"Cheapest weekend flight from Berlin and what's the weather there?",
tools,
executeTool
);
Claude will typically call search_flights first, read the cheapest destination from the result, then call get_weather with that city — no extra orchestration code required beyond the loop above. That's the practical payoff of tool use: multi-step reasoning without you hardcoding the steps.
Keep it simple before scaling up
Resist the urge to build a full multi-agent framework on day one. A single well-designed loop with two or three sharp tools solves most real problems. Add retries, timeouts, and logging around tool execution before you add more agents or more tools — that's usually where the actual reliability gains come from.
FAQs
Do I need LangChain or another framework to build an agent with Claude? No. Tool use is a native part of the Claude API, and the agent loop is roughly 30 lines of code as shown above. Frameworks add convenience for complex multi-agent systems, but a single-agent tool loop doesn't need one.
How many tools can an agent have? There's no hard cap in the API, but practically, keep it under 10–15 well-described tools per agent. Beyond that, Claude's tool selection accuracy tends to drop, and it's usually a sign you need multiple specialized agents instead.
What's the difference between tool use and an "agent"? Tool use is the mechanism (Claude requesting a function call). An agent is the surrounding loop and state management that lets that mechanism run repeatedly toward a goal instead of stopping after one call.