Claude API Code Review Automation: A Practical Guide
Why automate code review with Claude
Automating code review with the Claude API means sending pull request diffs to Claude and getting back structured feedback — bugs, security issues, style violations, missing tests — before a human reviewer ever opens the PR. This doesn't replace human review, but it catches obvious issues fast, enforces consistency across a team, and shortens the feedback loop from hours to seconds.
The typical setup is a CI job or GitHub Action that triggers on pull_request, fetches the diff, sends it to Claude with a review-focused prompt, and posts the result as a PR comment. Below is a working implementation, plus the design decisions that actually matter: diff chunking, prompt structure, output format, and cost control.
The basic architecture
A code review automation pipeline has four parts:
- Trigger — a webhook or CI step fires on PR open/update.
- Diff extraction — pull the changed files and their diffs via the Git provider's API.
- Model call — send the diff (and relevant context) to Claude with review instructions.
- Output delivery — post the response as a PR comment or inline annotations.
The part developers usually get wrong is step 3 — sending too much context (whole repo), too little structure (free-text prose instead of actionable findings), or ignoring token limits on large diffs.
Writing a review prompt that works
Vague prompts like "review this code" produce vague output. A good review prompt constrains the model to specific categories and a specific output format:
You are a senior code reviewer. Review the following diff for:
1. Bugs or logic errors
2. Security vulnerabilities (injection, auth, secrets)
3. Missing error handling
4. Missing or inadequate tests
5. Style/convention violations
Only flag real issues — do not comment on subjective style preferences
unless they violate the project's existing conventions shown in the diff.
Respond in this exact JSON format:
{
"issues": [
{"file": "path", "line": 42, "severity": "high|medium|low", "comment": "..."}
],
"summary": "one paragraph overall assessment"
}
Forcing JSON output makes it trivial to post inline comments at specific lines rather than dumping a wall of text into the PR.
Handling large diffs
Claude's context window is large, but PRs with hundreds of changed files or generated code (lockfiles, migrations) will blow through useful context and cost. Practical mitigations:
- Exclude generated files — lockfiles,
dist/,vendor/, snapshot tests. Filter these before sending. - Chunk by file, not by byte count — send each file's diff as a separate request when the PR is large, and merge results afterward. This also makes retries cheaper when one file fails.
- Send surrounding context sparingly — a few lines before/after the changed hunk is usually enough; you rarely need the whole file.
async function reviewDiff(fileDiff, filename) {
const res = await fetch('https://api.subtoapi.app/v1/messages', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.SUBTOAPI_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'claude-opus-4',
max_tokens: 1024,
messages: [{
role: 'user',
content: `Review this diff from ${filename}:\n\n${fileDiff}`,
}],
}),
});
return res.json();
}
Running this per-file in parallel (with a concurrency cap) keeps review latency reasonable even on PRs with dozens of changed files.
Posting results back to the PR
Once you have structured JSON findings, map them to inline review comments using your Git provider's review API (GitHub's pulls/{pr}/reviews endpoint accepts a list of comments with path and line). Severity levels are useful here: you might auto-request-changes on "high" severity issues but just comment on "low" ones, keeping the bot from blocking merges over nitpicks.
A simple policy that works well in practice:
highseverity → post as a blocking review commentmedium→ post as a regular commentlow→ bundle into the summary only, don't clutter the diff
Streaming vs single-shot for review jobs
CI jobs don't benefit much from streaming since there's no human watching output live — a single-shot call with a reasonable max_tokens is simpler to parse and retry. Streaming is more useful if you're building an interactive review assistant inside an IDE or chat tool, where partial output improves perceived latency. If you're building that kind of interactive layer on top of Claude, see /docs/streaming for the relevant request format.
Where SubToAPI fits
If you're calling Claude from CI, cost visibility and access control matter more than they do for a one-off script. SubToAPI turns your existing Claude access into a standard HTTPS API with sub_live_... keys you can scope per project — so your "code review bot" key is separate from keys used elsewhere, and you can see exactly how many tokens the review pipeline consumes per run in the dashboard. Tool use is also supported if your review pipeline needs Claude to call out to a linter or test runner mid-review — see /docs/tools for details.
Setup is the same /v1/messages call shown above; swap your API key and you're running. Check /docs/quickstart for the full setup, or /pricing if you want to compare Solo, Team, and Scale seat pricing before rolling this out to a whole engineering org.
Rolling it out without annoying your team
A few practical tips from teams that have shipped this:
- Start in comment-only mode. Don't let the bot block merges until you trust its false-positive rate.
- Tune the prompt per repo, not globally. A backend service and a frontend app have different common failure modes.
- Log every review call. When the bot flags something wrong, you want the exact prompt/response to debug and improve the system prompt.
- Cap cost per PR. Set a hard limit on diff size sent to the model; huge auto-generated PRs (dependency bumps) should skip AI review entirely.
Questions
Does Claude API code review automation replace human reviewers? No. It's best used as a first pass that catches obvious bugs, security issues, and missing tests before a human reviewer spends time on the PR — it reduces review load, it doesn't eliminate it.
How do I avoid high token costs on large pull requests? Filter out generated/vendor files, chunk the diff by file instead of sending the whole PR in one call, and set a size cap that skips AI review for abnormally large auto-generated PRs like dependency bumps.
Can I get inline PR comments instead of a single text blob? Yes — have Claude return structured JSON with file, line, severity, and comment fields, then map that output to your Git provider's review API to post inline comments automatically.