Building AI Agents with Claude API: A Practical Guide
Building AI agents with the Claude API means designing a loop where Claude decides which actions to take, calls tools to gather information or perform tasks, and continues reasoning until it reaches a final answer. Unlike a simple chat request, an agent is a system: it has memory across steps, access to external tools, and logic that decides when to stop.
This guide covers the core architecture of a Claude-based agent, how the tool-use loop actually works, where most implementations break in production, and practical patterns for keeping agents reliable and observable.
What Makes Something an "Agent"
A single API call to Claude that returns text is not an agent — it's a completion. An agent emerges when you add:
- A loop: the model's output can trigger another call rather than ending the conversation.
- Tools: functions the model can request to execute (search, database queries, code execution, API calls).
- State: conversation history, intermediate results, and task context persisted across steps.
- A stopping condition: logic that decides when the task is done versus when to keep iterating.
Claude supports this directly through its tool-use (function calling) feature. You define tools with JSON schemas, Claude decides when to invoke them, and your code executes them and returns results back into the conversation.
The Core Agent Loop
At a high level, every Claude-based agent follows the same pattern:
- Send the user's request plus available tool definitions to Claude.
- Claude responds either with a final answer or a
tool_userequest. - If it's a tool request, your code executes the tool and sends the result back as a
tool_result. - Repeat until Claude returns a final text response with no further tool calls.
Here's a minimal implementation:
async function runAgent(userMessage, tools, executeTool) {
let messages = [{ role: "user", content: userMessage }];
while (true) {
const response = await client.messages.create({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
tools,
messages,
});
messages.push({ role: "assistant", content: response.content });
const toolCalls = response.content.filter(b => b.type === "tool_use");
if (toolCalls.length === 0) {
return response.content.find(b => b.type === "text")?.text;
}
const toolResults = [];
for (const call of toolCalls) {
const result = await executeTool(call.name, call.input);
toolResults.push({
type: "tool_result",
tool_use_id: call.id,
content: JSON.stringify(result),
});
}
messages.push({ role: "user", content: toolResults });
}
}
This is the entire shape of most production agents. Everything else — retries, guardrails, logging, multi-agent coordination — is layered on top of this loop.
Designing Good Tools
The quality of your agent depends heavily on how well you define tools, not just how good the model is. A few practical rules:
- Narrow scope beats broad scope. A tool called
search_orderswith a clear schema works better than a genericdatabase_querytool that accepts raw SQL. - Return structured, bounded results. If a tool can return thousands of rows, truncate and summarize before sending it back to Claude — large tool outputs burn context and slow down reasoning.
- Make errors first-class. If a tool fails, return a clear error message as the tool result rather than throwing an exception that kills the loop. Claude can often recover by trying a different approach.
- Document edge cases in the description. Claude relies entirely on the tool's name, description, and parameter schema — there's no implicit knowledge of your system.
State and Memory
For short tasks, passing the full message history on every call is fine. For longer-running agents — multi-turn research tasks, background jobs, anything that spans minutes or more — you need an explicit state strategy:
- Persist the message array per session (database, Redis, or even a file for prototypes).
- Periodically summarize older turns to keep the context window manageable.
- Separate "working memory" (the current task's intermediate results) from "conversation memory" (what the user actually said).
A common mistake is treating the agent's full transcript as permanent state. In practice, you want a trimmed, relevant subset passed to each call — this keeps latency and cost predictable as the agent runs longer.
Stopping Conditions and Guardrails
Agents that loop indefinitely are a real production risk — both in cost and in correctness. Build in:
- A max iteration count. Hard-cap the loop (e.g., 10–15 tool calls) and return a fallback message if exceeded.
- Timeouts per tool call. A hanging external API shouldn't hang the whole agent.
- Confirmation steps for destructive actions. If a tool can delete data, send money, or send a message externally, add an explicit confirmation step before execution rather than trusting the model's judgment alone.
- Logging every step. Store the full sequence of tool calls, inputs, and outputs per run — this is how you debug agents that misbehave, and it's not optional once you have real users.
Running This in Production
The loop above assumes you already have Claude API access set up with proper authentication, streaming for long responses, and usage tracking per user or per feature. If you're building an agent that multiple team members or applications will call, you generally want:
- Scoped API keys per application rather than one shared key.
- Visibility into token usage per agent, since tool-heavy loops can consume more tokens than a single chat request.
- Streaming support so long-running agent steps don't block on a single large response.
This is where SubToAPI fits in if you're already paying for Claude access and want to turn it into a proper API layer: it gives you application-scoped keys (sub_live_...), streaming, tool-use support, and per-key usage metadata in one dashboard — useful when you're running several agents or giving teammates access without sharing a single credential. See the tool use docs and streaming docs for the specifics, or the quickstart if you're starting from zero.
A Minimal Checklist Before Shipping an Agent
- Tools have narrow, well-documented scopes with bounded outputs.
- The loop has a hard iteration cap and per-tool timeouts.
- Destructive actions require explicit confirmation.
- Every step is logged (input, tool call, output) for debugging.
- Token usage is tracked per agent run, not just per application.
Questions
Do I need a special framework to build agents with Claude? No. The tool-use loop is a few dozen lines of code, as shown above. Frameworks like LangChain or custom orchestration layers can help with complex multi-agent systems, but a single-agent tool loop is straightforward to build directly against the API.
How is an "agent" different from just using the chat API? A chat API call returns one response to one input. An agent runs a loop where the model can request tool execution, receive results, and continue reasoning across multiple steps before producing a final answer.
What's the biggest cause of unreliable agents? Poorly scoped tools and missing stopping conditions. Agents that loop without a max iteration count or that receive huge, unbounded tool outputs tend to drift, repeat actions, or burn through context unpredictably.