Claude API Code Review Automation Tool: Build Guide
A Claude API code review automation tool is a script or CI job that sends a pull request's diff to Claude, gets back structured feedback on bugs, style, and security issues, and posts that feedback as comments on the PR. It replaces (or augments) a first-pass human reviewer by catching obvious problems — unhandled errors, SQL injection risks, missing null checks, inconsistent naming — before a human ever opens the diff.
This article shows exactly how to build one: the architecture, the prompt structure that actually produces useful reviews instead of generic praise, how to handle large diffs that exceed context limits, and how to wire it into GitHub Actions or GitLab CI. If you'd rather skip the API key management and billing setup, SubToAPI gives you a standard HTTPS endpoint for Claude that drops straight into any of the examples below.
How automated code review with Claude works
The core loop is simple:
- A CI pipeline triggers on pull request creation or update.
- The pipeline fetches the diff (not the full repo — diffs keep token usage manageable).
- The diff, plus relevant context (file paths, PR description, coding standards), is sent to Claude with a review-focused system prompt.
- Claude returns structured findings — ideally JSON with file, line, severity, and comment.
- A script posts those findings as inline PR comments via the GitHub/GitLab API.
The hard parts aren't the Claude call itself — it's steps 2, 4, and 5: getting a clean diff, forcing structured output, and mapping line numbers back to the actual file.
Why diffs, not full files
Sending whole files burns tokens and dilutes the model's attention. A unified diff with a few lines of surrounding context is usually enough for Claude to spot real issues, and it keeps each PR review well within a single request even for large pull requests. For genuinely massive diffs (500+ changed lines), chunk by file and review each file separately, then merge the results.
Prompt design that produces useful reviews
Generic prompts like "review this code" produce generic output ("looks good, consider adding comments"). A reviewer prompt needs three things: a narrow scope, a required output format, and explicit severity levels so the tool doesn't flood the PR with nitpicks.
You are a senior code reviewer. Review ONLY the diff below.
Flag issues in these categories: correctness bugs, security
vulnerabilities, resource leaks, and broken error handling.
Do NOT comment on formatting, naming style, or missing comments
unless they cause a real bug.
For each issue, output a JSON object with:
- file (string)
- line (number, from the diff's new line numbers)
- severity ("blocker" | "warning" | "suggestion")
- comment (string, one or two sentences)
If there are no issues, return an empty array.
Diff:
<<<DIFF_CONTENT>>>
This constraint-heavy prompt matters more than model choice. Without the "do not comment on formatting" line, Claude will cheerfully point out every missing semicolon, which trains the team to ignore the bot entirely.
Forcing structured output
Ask for JSON explicitly and validate the response before posting anything. Claude is generally reliable at JSON formatting when instructed clearly, but always wrap the parse in a try/catch and skip posting on malformed responses rather than crashing the pipeline. For stricter guarantees, use tool use so the review findings are returned as a structured tool call instead of free text.
Example: GitHub Actions workflow
name: claude-code-review
on:
pull_request:
types: [opened, synchronize]
jobs:
review:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Get diff
run: git diff origin/${{ github.base_ref }}...HEAD > diff.txt
- name: Run Claude review
run: node scripts/review.js
env:
SUBTOAPI_KEY: ${{ secrets.SUBTOAPI_KEY }}
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
And the review script, calling Claude through a standard API key:
import fs from "fs";
const diff = fs.readFileSync("diff.txt", "utf8");
const res = await fetch("https://api.subtoapi.app/v1/messages", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
"content-type": "application/json",
},
body: JSON.stringify({
model: "claude-opus-4",
max_tokens: 2000,
messages: [
{ role: "user", content: `${REVIEW_SYSTEM_PROMPT}\n\nDiff:\n${diff}` },
],
}),
});
const data = await res.json();
const findings = JSON.parse(data.content[0].text);
for (const finding of findings) {
// post each finding as a PR comment via the GitHub API
}
This is the full shape of the tool: fetch diff, call the model, parse structured findings, post comments. Everything else — rate limiting, retries, caching results across force-pushes — is refinement on top of this loop.
Handling scale and cost
Once this runs on every push to every open PR, volume adds up fast. A few practical controls:
- Only review the delta. On
synchronizeevents, diff against the previous commit, not the base branch, so you don't re-review unchanged code. - Cache by commit SHA. Skip the Claude call entirely if that commit was already reviewed.
- Set severity thresholds. Only post comments for
blockerandwarning; logsuggestion-level findings without cluttering the PR. - Track usage per repo or team. If multiple teams share the same integration, usage metadata lets you see which repos are driving cost, which matters once you're running this across dozens of repositories. SubToAPI's dashboard exposes this per API key, which is useful when a single key is shared across a CI org.
Build vs. buy
Building this tool yourself is a few hundred lines of code — the example above is close to the whole thing. The part that takes longer is operational: managing an API key securely in CI secrets, handling rate limits across concurrent PR reviews, and giving each repo or team its own key so you can see where usage and cost come from without reading raw provider logs.
That's the gap SubToAPI fills: it turns your Claude access into application API keys (sub_live_...) you can issue per repo or per team, with streaming, tool use, and usage metadata available through a single dashboard. If you're reviewing dozens of pull requests a day across several repositories, that per-key visibility is worth more than it sounds — it's the difference between "our Claude bill went up" and "repo X's review bot is driving 80% of the spend." Check /pricing for plan details or start with the free trial at /signup.
Getting started quickly
If you want to prototype the review loop today rather than build CI plumbing first, start with a plain script that reads a diff file and prints findings to the console. Once the prompt is producing reviews you'd actually trust, wire it into CI. The quickstart guide and messages API reference cover the request format if you're setting this up against SubToAPI's endpoint.
questions
Does Claude catch real bugs, or just style issues? With a scoped prompt that explicitly excludes style and formatting, Claude reliably catches logic errors, missing error handling, and common security patterns like injection risks. It's not a replacement for static analysis tools, but it complements them by reasoning about intent, not just syntax.
How do I avoid reviewing the same code repeatedly on every push? Diff against the last reviewed commit SHA instead of the PR's full history, and cache results per commit so force-pushes or rebases without code changes don't trigger a new API call.
Can this run in GitLab CI instead of GitHub Actions? Yes — the logic is identical. Swap the diff-fetching step for GitLab's merge request diff endpoint and post comments through GitLab's discussions API instead of GitHub's PR comments API.