← Blog

Claude API Agent Framework Integration Guide

2026-10-10 · 5 min read · SubToAPI Team

Integrating Claude into an agent framework means connecting Claude's messages and tool-use API to the loop that your framework uses to plan, call tools, and decide when to stop. Most frameworks — LangChain, LlamaIndex, CrewAI, AutoGen, or a hand-rolled agent loop — expect three things from a model provider: a chat completion endpoint, structured tool/function calling, and streaming output. Claude supports all three natively, so integration is mostly a matter of mapping the framework's model adapter to Claude's request and response shapes correctly.

This guide walks through the practical integration patterns: wiring Claude into popular frameworks, handling tool calls in an agent loop, managing streaming inside multi-step reasoning, and the authentication layer that ties it together. It also covers where a gateway like SubToAPI fits if you want one API key, usage metadata, and team access instead of juggling raw provider credentials across every agent you build.

Why agent frameworks need special handling for Claude

Agent frameworks treat the LLM as a decision-maker inside a loop: observe state, pick an action (often a tool call), execute it, feed the result back, repeat until done. Claude's tool use API is built for exactly this pattern — you send a list of tool definitions, Claude returns a tool_use content block with structured input, you run the tool, and you send a tool_result block back in the next message. The integration work is making sure your framework's adapter:

Get any of these wrong and you'll see agents that either loop forever, hallucinate tool outputs, or silently drop context mid-task. See /docs/tools for the exact request/response shape.

Wiring Claude into LangChain or LlamaIndex

Both frameworks expose a model abstraction layer where you plug in an API endpoint and key. The integration pattern is the same regardless of provider:

const agent = new AgentExecutor({
  llm: new ClaudeChatModel({
    apiKey: process.env.SUBTOAPI_KEY,
    baseURL: "https://api.subtoapi.app/v1/messages",
    model: "claude-sonnet-4-5",
  }),
  tools: [searchTool, calculatorTool, dbQueryTool],
  maxIterations: 8,
});

const result = await agent.invoke({ input: "Find last quarter's churn and summarize it." });

The key detail: the framework's "chat model" wrapper needs to speak Claude's message format (system, messages array with role: user/assistant, and content blocks), not OpenAI's. If you're using a framework's built-in Claude adapter, this is handled for you. If you're writing a custom adapter, model your request body after /docs/messages exactly — mismatched field names are the most common cause of silent agent failures.

Building a custom agent loop from scratch

If you're not using a heavy framework, a minimal agent loop looks like this:

let messages = [{ role: "user", content: task }];

while (true) {
  const response = await fetch("https://api.subtoapi.app/v1/messages", {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "claude-sonnet-4-5",
      max_tokens: 1024,
      messages,
      tools: toolDefinitions,
    }),
  }).then(r => r.json());

  messages.push({ role: "assistant", content: response.content });

  const toolCalls = response.content.filter(b => b.type === "tool_use");
  if (toolCalls.length === 0) break; // agent finished reasoning

  const toolResults = await Promise.all(
    toolCalls.map(async (call) => ({
      type: "tool_result",
      tool_use_id: call.id,
      content: await runTool(call.name, call.input),
    }))
  );

  messages.push({ role: "user", content: toolResults });
}

This pattern is framework-agnostic and works whether you're calling Claude directly or through a gateway. The loop terminates when Claude stops requesting tools — a cleaner signal than regex-matching output text, which is how a lot of early agent implementations worked.

Streaming inside an agent loop

Streaming and tool use aren't mutually exclusive, but most agent frameworks buffer the full response before parsing tool calls because partial JSON in a tool input isn't usable mid-stream. If your agent needs to show live "thinking" text to a user while still supporting tool calls, stream the text portions and buffer only the tool-use blocks until they close. /docs/streaming covers the event types (content_block_start, content_block_delta, content_block_stop) you'll need to distinguish text deltas from tool-input deltas.

Multi-agent setups and API key management

Frameworks like CrewAI or AutoGen spin up multiple agents that may each need their own model credentials, rate limits, and usage tracking — especially if one agent is a "manager" calling Claude more often than worker agents. Running every agent off a shared raw provider key makes it hard to see which agent is burning tokens or to rotate credentials when one agent's tool integration misbehaves.

This is where routing agent traffic through SubToAPI helps in practice: each agent (or each environment — dev, staging, prod) gets its own sub_live_... key, usage and latency show up per key in the dashboard, and you can revoke one agent's access without touching the others. Setup is the same /v1/messages endpoint your framework already expects, so no adapter rewrite is needed — see /docs/quickstart to generate a key and swap the base URL.

Common integration mistakes

Getting started

For a first integration, start with a single-tool agent loop before wiring in a full framework — it's much easier to debug a tool_use_id mismatch in 30 lines of code than inside a framework's internals. Once that loop works reliably, swapping in LangChain, LlamaIndex, or a custom orchestrator is mostly configuration. Check /pricing if you're evaluating per-seat costs for a team running several agents in parallel, and /signup to get a key and start testing.

questions

Does Claude's tool use work the same across all agent frameworks? The underlying API is identical, but each framework has its own adapter layer for serializing tools and parsing responses. Bugs are usually in the adapter, not the API — check field mapping against /docs/messages first.

Can I stream responses while an agent is making tool calls? Yes, but most implementations stream the text segments live and buffer tool-use JSON blocks until they're complete, since partial tool input isn't executable mid-stream.

Do I need separate API keys for each agent in a multi-agent system? It's not required, but it's recommended. Separate keys per agent make usage tracking, rate limiting, and revocation far easier when something misbehaves — SubToAPI supports issuing multiple keys under one team account.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →