← Blog

Claude API Code Review Assistant Setup Guide

2026-09-30 · 5 min read · SubToAPI Team

Setting up a Claude API code review assistant means wiring three things together correctly: a way to pull the diff you want reviewed, a prompt that turns Claude into a strict but useful reviewer, and an API layer that handles authentication, streaming, and rate limits without breaking your CI pipeline. Most teams get stuck not on the AI part but on the plumbing — getting a real API key, keeping it out of shared CI secrets chaos, and making sure a single noisy PR doesn't burn your whole monthly budget.

This guide walks through a complete setup: choosing where the assistant runs (local CLI, GitHub Action, or PR bot), structuring the diff so Claude reviews it accurately instead of hallucinating context, and the exact API calls needed to get consistent, actionable review comments back.

Decide Where the Assistant Runs

There are three common setups, and the right one depends on your workflow:

Start with the CI job. It's the easiest to test, doesn't require hosting anything, and gives every reviewer the same baseline feedback before a human even opens the PR.

Get an API Key

If you already pay for Claude through a Pro or Team seat but don't have direct API access, SubToAPI turns that subscription into a standard HTTPS API with an sub_live_... key, streaming support, and usage metadata — which matters here because you'll want to track exactly how many tokens each PR review consumes. Sign up, grab a key from the dashboard, and store it as SUBTOAPI_KEY in your CI secrets.

Quick sanity check before wiring it into CI:

curl https://api.subtoapi.app/v1/messages \
  -H "Authorization: Bearer $SUBTOAPI_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4",
    "max_tokens": 200,
    "messages": [{"role": "user", "content": "Reply with OK if this works."}]
  }'

Full request/response shape is in the messages docs and the quickstart if this is your first API call.

Prepare the Diff, Not the Whole Repo

The most common mistake in a code review assistant setup is dumping entire files into the prompt. Claude reviews diffs far more accurately when it sees the actual change plus a small amount of surrounding context, not the whole codebase. Pull the diff with git directly in CI:

git fetch origin main
git diff origin/main...HEAD > diff.patch

Keep the patch under a reasonable size (roughly 3,000–4,000 lines of diff per request). For larger PRs, split by file or by directory and send multiple requests rather than one giant prompt — this also keeps failures isolated, so one unparseable file doesn't kill the whole review.

Write a Review-Specific System Prompt

A generic "review this code" prompt produces vague output. Be explicit about scope, severity, and format:

{
  "model": "claude-sonnet-4",
  "max_tokens": 1500,
  "system": "You are a senior code reviewer. Review only the diff provided, not the whole file. Flag: bugs, security issues, unhandled errors, and clear style violations of the existing codebase. Ignore formatting Prettier/Black would fix. For each issue, output: file:line, severity (blocker/warning/nit), and a one-sentence fix. If nothing is wrong, say so explicitly — do not invent issues.",
  "messages": [
    {"role": "user", "content": "Review this diff:\n\n<diff contents>"}
  ]
}

That last instruction — "do not invent issues" — matters more than it looks. Without it, Claude will sometimes pad output with minor nitpicks to seem thorough, which trains developers to ignore the bot.

Stream Output for Faster CI Feedback

For local CLI use, streaming makes the review feel instant instead of a 20-second silent wait:

const res = await fetch("https://api.subtoapi.app/v1/messages", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.SUBTOAPI_KEY}`,
    "content-type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-sonnet-4",
    max_tokens: 1500,
    stream: true,
    system: reviewSystemPrompt,
    messages: [{ role: "user", content: diffText }],
  }),
});
// stream chunks to stdout as they arrive

Details on parsing the event stream are in the streaming docs. For a CI job posting a single PR comment, non-streaming is simpler — you only need the final text.

Give It Tools for Deeper Reviews

A review limited to the diff text alone can't check whether a function is actually used elsewhere or whether a test covers the change. If you want the assistant to look things up — search the repo, run the test suite, check for existing usages of a renamed function — define tools it can call and let Claude request them mid-conversation rather than trying to stuff the entire repo into context. This turns a surface-level diff review into something closer to what a senior engineer actually does before approving a PR. The tools documentation covers the request/response format for defining and handling tool calls.

Post the Review Back to the PR

Whatever CI system you use, the pattern is the same: run the API call, capture the text response, and post it as a PR comment via your git host's API (GitHub, GitLab, Bitbucket all support comment creation via REST). Keep the comment scoped to blockers and warnings at the top, nits collapsed or omitted, so developers don't tune out the bot after the third PR.

Control Cost Per PR

Every review call costs tokens proportional to diff size, and PR volume is unpredictable — a big refactor PR can be 10x the tokens of a typical fix. Set a max_tokens cap, chunk large diffs, and check the pricing page to pick a plan that matches your team's PR volume: Solo works for a single maintainer running local reviews, while Team and Scale add per-seat API keys so each contributor's usage is tracked separately in the same dashboard.

questions

Do I need a separate API key per developer? Not required, but recommended once you have more than a couple of contributors — separate keys make it easy to see whose PRs are driving token usage and to revoke access individually without rotating a shared secret.

Should the assistant block merges or just comment? Start with comment-only. Treat it as a second opinion, not a gate — false positives on a hard block will erode trust in the tool fast. Once you've tuned the prompt against a few weeks of real PRs, you can wire blocker-severity findings into a required check.

How do I keep it from reviewing generated or vendored files? Filter the diff before sending it — exclude paths like dist/, vendor/, *.lock, and generated protobuf/GraphQL files with a .gitattributes or a simple path filter in your CI script. Sending generated code wastes tokens and produces irrelevant comments.

Turn your Claude access into an HTTPS API

SubToAPI gives you application API keys, streaming, tool use and usage insights on top of your existing Claude access — set up in minutes.

Start free  Read the quickstart →