← Blog

Claude API Autonomous Agent Tutorial: Step by Step

2026-10-08 · 5 min read · SubToAPI Team

What "autonomous agent" actually means with Claude

If you've searched for a Claude API autonomous agent tutorial, you're probably trying to build something that doesn't just answer one question and stop — it should plan, call tools, read the results, decide what to do next, and keep going until the task is done. That's the whole trick: an agent is a loop around Claude, not a single API call.

This tutorial walks through that loop end to end: defining tools, running the think-act-observe cycle, managing state across turns, and knowing when to stop. The code uses plain HTTP so it works whether you're calling Claude directly or through a proxy like SubToAPI that exposes the same Messages-style API.

The core building block: the agent loop

An autonomous agent built on the Claude API is fundamentally a while loop with four steps:

  1. Send the conversation history + available tools to Claude.
  2. If Claude's response contains a tool call, execute it in your code.
  3. Append the tool result to the conversation.
  4. Repeat until Claude returns a final text answer (no more tool calls) or you hit a safety limit.

Here's the minimal version in JavaScript:

async function runAgent(userTask, tools, executeTool) {
  const messages = [{ role: "user", content: userTask }];
  const MAX_TURNS = 10;

  for (let turn = 0; turn < MAX_TURNS; turn++) {
    const res = await fetch("https://api.subtoapi.app/v1/messages", {
      method: "POST",
      headers: {
        "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
        model: "claude-sonnet-4-5",
        max_tokens: 1024,
        messages,
        tools,
      }),
    });

    const data = await res.json();
    messages.push({ role: "assistant", content: data.content });

    const toolCalls = data.content.filter(b => b.type === "tool_use");
    if (toolCalls.length === 0) {
      return data.content.find(b => b.type === "text")?.text;
    }

    const toolResults = [];
    for (const call of toolCalls) {
      const result = await executeTool(call.name, call.input);
      toolResults.push({
        type: "tool_result",
        tool_use_id: call.id,
        content: JSON.stringify(result),
      });
    }
    messages.push({ role: "user", content: toolResults });
  }

  throw new Error("Agent did not finish within MAX_TURNS");
}

That's the entire skeleton. Everything else — better planning, memory, retries — is built on top of this loop. Full request/response shapes are in the docs and /docs/messages.

Defining tools the agent can actually use

Claude decides when to call a tool based on the tool's name, description, and input schema — so the quality of your schema directly determines how autonomous the agent can be. For a research agent, you might define:

{
  "name": "search_web",
  "description": "Search the web and return the top results with titles, URLs, and snippets.",
  "input_schema": {
    "type": "object",
    "properties": {
      "query": { "type": "string" }
    },
    "required": ["query"]
  }
}
{
  "name": "fetch_page",
  "description": "Download and extract the readable text content of a given URL.",
  "input_schema": {
    "type": "object",
    "properties": {
      "url": { "type": "string" }
    },
    "required": ["url"]
  }
}

Give Claude both tools and a task like "research the top 3 competitors to X and summarize their pricing," and it will chain search_web → fetch_page → search_web again on its own, deciding the order without you hardcoding it. See /docs/tools for the full schema reference and multi-tool examples.

Managing state across turns

Autonomous agents fail most often not because the model is wrong, but because the surrounding state is mismanaged. Three things matter:

Keep the full tool_use/tool_result pairing. Every tool_use block from the assistant must be matched with a tool_result block in the next user message with the same tool_use_id. If you drop or reorder these, the next call will error or hallucinate missing context.

Trim history before you hit context limits. Long-running agents accumulate tool outputs fast (page scrapes, search results). Summarize or truncate older tool results before appending new ones rather than letting the array grow unbounded.

Persist state outside the conversation for anything that matters. Don't rely on the model "remembering" a decision from ten turns ago — write key facts to a scratch object in your own code and re-inject them into the prompt if needed.

Stopping conditions: the part most tutorials skip

An agent without a stopping condition is a bug, not a feature. Build in at least three:

tools.push({
  name: "finish_task",
  description: "Call this when the task is fully complete, with a final summary.",
  input_schema: {
    type: "object",
    properties: { summary: { type: "string" } },
    required: ["summary"],
  },
});

Streaming for long-running agents

If your agent's tasks take more than a few seconds — common once it's chaining multiple tool calls — stream the response so your UI can show progress instead of a blank spinner. Claude's streaming format sends incremental content_block_delta events you can render as they arrive; see /docs/streaming for the event shapes and a working example.

Where SubToAPI fits

The agent loop itself is provider-agnostic — it's just HTTP calls and your own control flow. What changes between projects is usually the operational side: issuing separate API keys per agent or environment, tracking token usage per project, and rotating keys without redeploying code. SubToAPI wraps your existing Claude access in an HTTPS API with scoped sub_live_... keys, streaming, and usage metadata per key, which is useful once you have more than one autonomous agent running in production and need to know which one is costing what. Start with the quickstart, check pricing, or sign up to get a key.

questions

Do I need a special "agent" API or framework to build this? No. An autonomous agent is a loop you write around standard Claude API calls with tools — the Messages API is enough; frameworks just add abstractions on top of the same loop.

How do I prevent an agent from running forever or overspending? Combine a hard max-turn limit, an explicit completion tool the model calls when done, and a token or time budget you check on every iteration before continuing.

Can an agent call multiple tools in a single turn? Yes — Claude can return several tool_use blocks in one response; execute them (in parallel if independent) and return all corresponding tool_result blocks together in the next message.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →