Claude API Code Generation Tool: A Practical Example
What this article covers
If you're searching for a Claude API code generation tool example, you're probably trying to figure out how to wire Claude into a real product — not just paste a prompt into a chat window. This article walks through a working example: sending a code generation request to the Claude API, structuring the response, handling streaming output for long files, and using tool calls to let Claude request additional context (like reading a file schema) before it writes code.
The short answer: you send a prompt describing the code you want, Claude returns text (often in a fenced code block), and if you need Claude to interact with your environment — fetch a schema, run a linter, check a file — you define tools it can call mid-generation. Below is a concrete setup you can adapt.
Basic code generation request
The simplest pattern is a single request/response. You describe the task, specify the language and constraints, and parse the code block out of the response.
const res = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 2000,
messages: [{
role: "user",
content: "Write a Python function that validates an IBAN number. Include a docstring and type hints. Return only the code."
}]
})
});
const data = await res.json();
console.log(data.content[0].text);
For predictable extraction, instruct Claude explicitly: "respond with a single fenced code block, no explanation." This keeps downstream parsing simple — you just strip the triple backticks instead of handling arbitrary prose.
Using tool calls for context-aware generation
Pure text-in, text-out works for isolated snippets, but real code generation tools usually need context: an existing file, a database schema, a style guide. The reliable way to give Claude that context on demand is tool use, where you define a function Claude can call, and you execute it on your side.
const tools = [{
name: "get_file_contents",
description: "Fetch the current contents of a file in the project",
input_schema: {
type: "object",
properties: {
path: { type: "string", description: "Relative file path" }
},
required: ["path"]
}
}];
const response = await fetch("https://api.anthropic.com/v1/messages", {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": "2023-06-01"
},
body: JSON.stringify({
model: "claude-sonnet-4-5",
max_tokens: 2000,
tools,
messages: [{
role: "user",
content: "Add input validation to the existing createUser function in src/users.js. Read the file first."
}]
})
});
Claude will respond with a tool_use block asking for get_file_contents with path: "src/users.js". Your application reads the file, sends it back as a tool_result, and Claude produces the updated code with the validation added — grounded in the actual file rather than a guess. This is the pattern behind most "AI coding assistant" integrations you see in editors and CI bots.
Streaming for longer generations
Code generation for whole files or multi-file scaffolds can take several seconds. Streaming avoids a blank UI during that time and lets you render code incrementally in an editor pane.
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-5",
"max_tokens": 4000,
"stream": true,
"messages": [{"role": "user", "content": "Generate a complete Express.js CRUD API for a todo resource, with routes, controller, and model."}]
}'
You'll get a sequence of content_block_delta events with text fragments. Append them in order and you have the full file as it's generated — useful for showing live progress in a code-gen tool's UI rather than a frozen loading spinner.
Where SubToAPI fits in
If you're building a code generation feature on top of Claude but don't want to manage billing, per-user API keys, or usage tracking yourself, SubToAPI turns your existing Claude access into a standard HTTPS API with sub_live_... application keys. You get the same /v1/messages shape, streaming for incremental code output, and tool use for context-aware generation — plus per-key usage metadata so you can see which feature or customer is consuming tokens. Setup is in the quickstart, and plans start at €9/month on the pricing page.
Practical tips for code generation prompts
- Pin the language and version. "Python 3.11" or "TypeScript with strict mode" avoids ambiguous syntax choices.
- Ask for a single code block. It makes extraction deterministic and avoids mixing explanation with code in your output pipeline.
- Give Claude the error, not just the task, when fixing bugs. Paste the stack trace or failing test output directly into the prompt.
- Cap
max_tokenssensibly. Large files need more headroom; short functions don't, and a lower cap reduces runaway generations. - Validate before shipping. Run generated code through a linter or test suite automatically — treat Claude's output as a draft, not a merge-ready PR.
Questions
Can the Claude API generate entire files, not just snippets? Yes. Increase max_tokens to accommodate the expected output length and use streaming for anything beyond a few hundred lines so your UI doesn't appear frozen.
Does Claude need tool use to generate code, or is a plain prompt enough? A plain prompt is enough for isolated, self-contained code. Tool use becomes necessary when the generated code must reference real project state — existing files, schemas, or APIs — that Claude doesn't have in its context window.
What's the fastest way to add a hosted API layer on top of Claude for a code generation product? Use SubToAPI to get application API keys, streaming, and usage tracking without building that infrastructure yourself — see the docs for the full API reference.