Claude API Autonomous Agent: Build Tutorial
What building an autonomous agent with the Claude API actually involves
An autonomous agent is a program that pursues a goal by repeatedly calling a language model, letting it decide which tool to use next, executing that tool, and feeding the result back in — without a human approving each step. Building one on the Claude API means wiring up four things: a tool definition schema, a loop that calls the model and executes whatever it asks for, a stopping condition so it doesn't run forever, and some form of state so it remembers what it already did.
This tutorial walks through that loop end to end with working code. It assumes you already have API access — either a direct Anthropic key or a wrapper service like SubToAPI that exposes the same /v1/messages interface over HTTPS with an application key. Either way, the agent logic below is identical; only the base URL and auth header change.
The agent loop, conceptually
Every autonomous agent built on Claude follows the same shape:
- Send the conversation history plus available tools to the model.
- If the model returns a
tool_useblock, execute that tool in your own code. - Append the tool result to the conversation as a
tool_result. - Send the updated conversation back to the model.
- Repeat until the model returns a plain text answer with no tool call, or you hit a stop condition.
That's it — there's no hidden "agent framework" magic underneath most agent libraries. They're all doing this loop with varying amounts of bookkeeping.
Step 1: Define your tools
Tools are just JSON Schema objects describing a function the model can call. Keep the first version narrow — two or three tools is enough to get a working agent.
{
"name": "search_docs",
"description": "Search internal documentation and return matching snippets.",
"input_schema": {
"type": "object",
"properties": {
"query": { "type": "string" },
"max_results": { "type": "integer", "default": 5 }
},
"required": ["query"]
}
}
If you're new to tool schemas, see /docs/tools for a working reference.
Step 2: Implement the loop in code
async function runAgent(userGoal, tools, executeTool) {
const messages = [{ role: "user", content: userGoal }];
const MAX_STEPS = 10;
for (let step = 0; step < MAX_STEPS; step++) {
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "claude-sonnet-4",
max_tokens: 1024,
tools,
messages
})
});
const data = await res.json();
messages.push({ role: "assistant", content: data.content });
const toolCall = data.content.find(b => b.type === "tool_use");
if (!toolCall) {
return data.content.find(b => b.type === "text")?.text ?? "Done.";
}
const result = await executeTool(toolCall.name, toolCall.input);
messages.push({
role: "user",
content: [{
type: "tool_result",
tool_use_id: toolCall.id,
content: JSON.stringify(result)
}]
});
}
return "Stopped: max steps reached.";
}
This is the entire skeleton of an autonomous agent. Everything else — planning quality, retry logic, memory — is built around this loop. If you're routing through SubToAPI, the request format matches Claude's native /v1/messages shape exactly, so this code works unchanged whether you're hitting Anthropic directly or a sub_live_... key. See /docs/messages for the full request/response reference.
Step 3: Add guardrails before you add intelligence
Autonomous agents fail loudly and expensively if you skip guardrails. At minimum, implement:
- A hard step limit. The
MAX_STEPSloop above prevents infinite tool-calling cycles. - Tool allowlisting. Never let the model call a tool name your
executeToolfunction doesn't explicitly recognize. - Timeouts per tool call. A hanging API call inside a tool shouldn't hang the whole agent.
- A cost ceiling. Track token usage per run and abort if it crosses a threshold — agents that loop on a flawed plan can burn tokens fast.
if (data.usage && data.usage.output_tokens > 50000) {
return "Stopped: token budget exceeded.";
}
Usage metadata is returned on every response, so tracking cumulative cost per agent run is a few lines of code, not a separate monitoring system.
Step 4: Give the agent memory across turns
A single messages array is fine for a single run, but real agents need to remember facts across sessions — user preferences, prior task outcomes, project context. Two practical approaches:
- Summarize and inject. After each run, ask Claude to compress the transcript into a short summary and store it. On the next run, prepend that summary instead of replaying the full history.
- External key-value store. Store structured facts (user_id → preferences) in your own database and fetch them before building the initial
messagesarray. This scales better than summarization once you have many agents running per user.
Either way, keep memory outside the model call itself — don't rely on the model to "remember" between HTTP requests, since each call is stateless.
Step 5: Make it observable
Before shipping an autonomous agent to production, log every step: which tool was called, with what input, what it returned, and how many tokens the step cost. When an agent misbehaves — and it will, eventually — this log is the only way to debug a multi-step decision chain after the fact. If you're using SubToAPI, per-key usage metadata is already visible in the dashboard, which covers token counts and request volume without extra instrumentation on your end.
Deploying the loop as an API
Once the loop works locally, wrap it in an HTTP endpoint (Express, FastAPI, whatever you already use) so other services can trigger agent runs. If you're building this as a product feature rather than a one-off script, you'll want streaming so the frontend can show progress as the agent works through steps — see /docs/streaming for how streamed tool-use events are structured. For teams running several agents in parallel, per-member API keys and usage breakdowns matter more than they seem to at first; that's one of the reasons SubToAPI issues separate sub_live_... keys per seat rather than one shared key — see /pricing for Solo, Team, and Scale plans, or start with a free trial at /signup.
Quick checklist before going live
- Step limit and token budget enforced
- Tool inputs validated before execution, not trusted blindly
- Every tool call and result logged
- Memory stored outside the model call, not assumed
- Clear fallback response when the agent hits its step limit
What's the difference between an autonomous agent and a simple tool-calling app?
A tool-calling app makes one model call, executes one tool, and returns an answer. An autonomous agent repeats that cycle multiple times without human input between steps, letting the model decide the next action based on previous tool results.
Do I need a framework like LangChain to build this?
No. The loop above is the entire mechanism most agent frameworks implement internally. A framework adds convenience (built-in retries, memory stores, tracing UIs) but isn't required to get a working agent running.
How do I stop an agent from running forever on a bad plan?
Enforce a hard step limit and a token budget in your loop, as shown above, and always have a defined fallback response for when either limit is hit instead of letting the loop exit silently.