Специализированный AI-ревьюер против универсального агента: практический тест

Ask a general-purpose AI agent to review a 60-file change and you will usually get a confident summary. What you often don't get is a comment on file 47, a comment attached to the line that actually contains the bug, or a token bill proportional to the work. This article gives you a test you can run on your own code to find out whether a dedicated reviewer such as Open Code Review does better — and by how much — instead of trusting either vendor's demo.
The test plants known defects into a large diff, runs both reviewers under matched conditions, and scores six things: precision, recall, file coverage, anchor validity, wall-clock time and tokens. Everything is scripted, so a second person can repeat it and get comparable numbers.
The problem in one paragraph
Code review of a large change is a coverage task disguised as a reasoning task. A general agent has to decide which files to open, keep enough of each one in its context window, and then translate “this is wrong” into a pull request comment with a valid path, line and side. Three failure modes follow directly: files that are never opened, comments that land on the wrong line (or are rejected by the platform and silently become a general comment), and a high token count from re-reading the same files. A specialised reviewer claims to fix these by owning the iteration over files and the mapping to diff positions. The claim is testable.
The concrete case
Use one service you know well — the example below assumes a mid-sized Python web service — and build one deliberately large change on top of a fixed base commit:
- Noise: 40–60 files of realistic but harmless edits: renames, type hints, log message wording, an extracted helper. This is what makes the diff “large”.
- Seeded defects: 12–20 bugs you insert on purpose and record precisely (file, side, line range, category). This is your ground truth.
- Decoys: 3–5 changes that look suspicious but are correct (for example, an intentionally broad
exceptwith a comment explaining why). These measure false alarms.
Place defects deliberately to probe the known weaknesses:
| Placement | Why | Example defect |
|---|---|---|
| Last files in path order | Tests whether the reviewer reaches the end of the diff | zz_reports/export.py: off-by-one in pagination |
| Deep inside a 600-line file | Tests attention within a long file | Missing await on an async DB call |
| Deleted lines only (left side) | Tests anchoring to removed code | Removed permission check in a view |
| Two-file interaction | Tests cross-file reasoning | Function now returns cents, caller still assumes rupees/dollars |
| Test file | Tests whether tests are reviewed at all | Assertion changed to always pass |
| Config / non-code file | Tests coverage beyond source code | Timeout changed from seconds to milliseconds without unit change |
Step 0. Pin the environment
The comparison is only meaningful if the reviewers differ in tooling, not in model, temperature or input. Write the environment down before the first run.
# bench/ENV.md — fill in before the first run
base_sha: <git rev-parse base>
head_sha: <git rev-parse bench/head>
open_code_review: <version or commit you installed>
general_agent: <name + version, e.g. output of `claude --version`>
model (both): <same model id for both tools, if OCR lets you choose>
temperature: <value, or "tool default — not configurable">
machine / network: <runner type, region>
date: <YYYY-MM-DD>
If Open Code Review cannot use the same model as the agent, say so in the results. You are then comparing two systems, not two harnesses, and the conclusion must be narrower.
Step 1. Build the benchmark change and the ground truth
git switch -c bench/head base
# ... apply noise edits, seeded defects and decoys ...
git commit -am "bench: large change with seeded defects"
mkdir -p bench/runs
git diff --no-color --no-ext-diff -U3 base bench/head > bench/change.diff
git diff --stat base bench/head | tail -1
Record every seeded defect as one line of JSONL. Line numbers refer to the new file for RIGHT and the old file for LEFT, matching how GitHub addresses review comments.
{"id":"D01","path":"zz_reports/export.py","side":"RIGHT","start":88,"end":90,"category":"logic","severity":"high","summary":"page offset uses page*size+1"}
{"id":"D02","path":"app/orders/service.py","side":"RIGHT","start":412,"end":412,"category":"async","severity":"high","summary":"missing await on repo.save"}
{"id":"D03","path":"app/views/admin.py","side":"LEFT","start":57,"end":59,"category":"security","severity":"critical","summary":"permission check removed"}
{"id":"D04","path":"app/billing/money.py","side":"RIGHT","start":21,"end":24,"category":"cross-file","severity":"high","summary":"returns minor units; see D04b caller"}
Keep bench/truth.jsonl out of the reviewed branch. If a reviewer can read the answer key, the test is void. Also strip commit messages and code comments that hint at the bugs (“TODO: fix pagination”).
Step 2. Use a sandbox PR as the common output channel
Comparing free-text output from two tools is where most informal tests go wrong. The cleanest neutral format is the platform itself: let each run post review comments to its own pull request in a private sandbox repository, then pull the comments back through the API. This also tests anchoring honestly, because GitHub rejects review comments whose line is not part of the diff.
# one branch (and one PR) per tool per run, all pointing at the same commit
for run in ocr-r1 ocr-r2 ocr-r3 agent-r1 agent-r2 agent-r3; do
git branch "bench/$run" bench/head
git push sandbox "bench/$run"
gh pr create --repo OWNER/SANDBOX --base base --head "bench/$run" \
--title "bench $run" --body "benchmark run, do not merge"
done
Do not run this against a public or shared repository: posted comments are visible to everyone with access and may be indexed or emailed to watchers.
Step 3. Run Open Code Review
Install it following the project's own README and record the exact version. Because invocation details differ between releases, the script below wraps it in a variable rather than guessing flags.
# bench/run_ocr.sh — example wrapper; set OCR_CMD from the README of your version
set -euo pipefail
RUN="$1"; PR="$2"
start=$(date +%s)
$OCR_CMD "$PR" 2>&1 | tee "bench/runs/$RUN.log"
end=$(date +%s)
echo "{\"run\":\"$RUN\",\"wall_seconds\":$((end-start))}" > "bench/runs/$RUN.time.json"
Capture token usage from whatever the tool reports. If it reports nothing, read usage from your model provider's console for an isolated API key used only for this run, and note that the figure was taken from the provider side.
Step 4. Run the general-purpose agent under matched conditions
Give the agent the same input (the PR, or bench/change.diff plus read access to the checkout), the same model if possible, and a prompt that asks for the same output. Do not coach it with hints you would not give a human reviewer.
# bench/prompts/review.md
You are reviewing the changes between `base` and `bench/head` in this repository.
Review every changed file. For each real problem, output one JSON object per line:
{"path": "...", "line": N, "side": "RIGHT|LEFT", "severity": "low|medium|high|critical", "body": "..."}
`line` is the new-file line for RIGHT and the old-file line for LEFT, and must be a line
that appears in the diff. Output nothing except these JSON lines.
Example for Claude Code in headless mode, with read-only tools. Verify the flags with claude --help for your installed version:
start=$(date +%s)
claude -p "$(cat bench/prompts/review.md)" \
--output-format json \
--allowedTools "Read,Grep,Glob,Bash(git diff:*)" \
> bench/runs/agent-r1.raw.json
end=$(date +%s)
echo "{\"run\":\"agent-r1\",\"wall_seconds\":$((end-start))}" > bench/runs/agent-r1.time.json
The JSON envelope contains the agent's final text and a usage block; extract the comment lines and the token counts from it. Then post the comments to the agent's sandbox PR with the same API the specialised tool uses, so both are judged by the platform's anchoring rules:
gh api repos/OWNER/SANDBOX/pulls/$PR/reviews --method POST --input review.json
# review.json: {"event":"COMMENT","commit_id":"<head_sha>","comments":[{"path":..,"line":..,"side":..,"body":..}]}
If the API rejects the review because one comment has an invalid line, do not fix it by hand. Record the rejection, drop the invalid comment, post the rest, and count the dropped one as an anchor failure. A real integration would have faced the same error.
Step 5. Normalize all comments to one file per run
gh api repos/OWNER/SANDBOX/pulls/$PR/comments --paginate \
--jq '.[] | {path, line, side, body}' -c > bench/runs/$RUN.comments.jsonl
# general (non-line) comments: kept separately, they cannot be anchored
gh api repos/OWNER/SANDBOX/pulls/$PR/reviews --paginate \
--jq '.[] | select(.body != "") | {body}' -c > bench/runs/$RUN.general.jsonl
Findings that only appear in a general review body count toward recall only after manual adjudication, and never toward anchor validity.
Step 6. Score
The script parses the unified diff to learn which lines are commentable on each side, then matches comments to ground truth by path, side and line with a small tolerance.
#!/usr/bin/env python3
"""Score normalized review comments against seeded ground truth."""
import argparse
import collections
import json
import re
HUNK = re.compile(r"^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@")
def commentable_lines(diff_path):
lines = collections.defaultdict(lambda: {"LEFT": set(), "RIGHT": set()})
path, old, new, in_hunk = None, 0, 0, False
with open(diff_path, encoding="utf-8", errors="replace") as f:
for raw in f:
line = raw.rstrip("\n")
if line.startswith("diff --git "):
path, in_hunk = None, False
continue
if not in_hunk and line.startswith(("--- ", "+++ ")):
p = line[4:]
if p != "/dev/null":
path = p[2:] if p[:2] in ("a/", "b/") else p
continue
m = HUNK.match(line)
if m:
old, new, in_hunk = int(m[1]), int(m[2]), True
continue
if not in_hunk or path is None:
continue
if line.startswith("+"):
lines[path]["RIGHT"].add(new)
new += 1
elif line.startswith("-"):
lines[path]["LEFT"].add(old)
old += 1
elif line.startswith(" "):
lines[path]["LEFT"].add(old)
lines[path]["RIGHT"].add(new)
old += 1
new += 1
return lines
def load_jsonl(path):
with open(path, encoding="utf-8") as f:
return [json.loads(l) for l in f if l.strip()]
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--diff", required=True)
ap.add_argument("--truth", required=True)
ap.add_argument("--comments", required=True)
ap.add_argument("--adjudicated", help="JSONL of {comment_idx, verdict: valid|invalid}")
ap.add_argument("--tolerance", type=int, default=3)
a = ap.parse_args()
valid = commentable_lines(a.diff)
truth = load_jsonl(a.truth)
comments = load_jsonl(a.comments)
verdicts = {}
if a.adjudicated:
verdicts = {r["comment_idx"]: r["verdict"] for r in load_jsonl(a.adjudicated)}
found, useful, anchored, pending = set(), 0, 0, []
for i, c in enumerate(comments):
side, line = c.get("side") or "RIGHT", c.get("line")
if isinstance(line, int) and line in valid.get(c["path"], {}).get(side, set()):
anchored += 1
hit = next((d["id"] for d in truth
if d["path"] == c["path"] and d.get("side", "RIGHT") == side
and isinstance(line, int)
and d["start"] - a.tolerance <= line <= d["end"] + a.tolerance), None)
if hit and verdicts.get(i) != "invalid":
found.add(hit)
useful += 1
elif verdicts.get(i) == "valid":
useful += 1
elif i not in verdicts:
pending.append(i)
n = len(comments)
defect_files = {d["path"] for d in truth}
commented_files = {c["path"] for c in comments}
print(json.dumps({
"comments": n,
"precision": round(useful / n, 3) if n else None,
"recall": round(len(found) / len(truth), 3),
"anchor_validity": round(anchored / n, 3) if n else None,
"diff_files_commented": f"{len(commented_files & set(valid))}/{len(valid)}",
"defect_files_commented": f"{len(commented_files & defect_files)}/{len(defect_files)}",
"missed": sorted(d["id"] for d in truth if d["id"] not in found),
"pending_adjudication": pending,
}, ensure_ascii=False, indent=2))
if pending:
print("WARNING: precision is a lower bound until pending comments are adjudicated")
if __name__ == "__main__":
main()
python3 bench/score.py --diff bench/change.diff --truth bench/truth.jsonl \
--comments bench/runs/ocr-r1.comments.jsonl --adjudicated bench/runs/ocr-r1.adj.jsonl
What each number means:
- Precision — share of comments that are useful: either they hit a seeded defect or a human judged them a real issue. A comment on a decoy counts as a false positive unless the adjudicator disagrees with your decoy.
- Recall — share of seeded defects found by at least one line comment.
- Anchor validity — share of comments sitting on a line that exists in the diff on the stated side. Comments the API rejected in Step 4 must be added back to the denominator by hand.
- File coverage — two ratios: changed files that received any comment, and defect-bearing files that received any comment. The second is the one that exposes skipped files.
- Time and tokens — from the
.time.jsonfiles and the usage figures captured per run.
Step 7. Adjudicate blind
Location matching is a filter, not a judgement: a comment three lines from D02 about a typo will match D02. Have a second person read every matched and pending comment without knowing which tool wrote it and record a verdict:
{"comment_idx": 0, "verdict": "valid"}
{"comment_idx": 4, "verdict": "invalid", "note": "near D02 but about naming, not the missing await"}
Shuffle comments from all runs into one sheet with tool names removed, then split verdicts back per run. This takes longer than the automated part and is the step that makes the numbers trustworthy.
Step 8. Repeat and report variance
Both tools are non-deterministic. Run each at least three times on the same commit and report the median plus min–max. A difference smaller than the spread between runs of the same tool is not a difference.
Results template
Fill this in from your own runs. The empty cells are intentional: we are not publishing numbers we did not measure on your code.
| Metric (median, min–max over runs) | Open Code Review | General agent |
|---|---|---|
| Precision | — | — |
| Recall (seeded defects) | — | — |
| Defect files commented | — | — |
| Anchor validity (incl. API rejections) | — | — |
| Decoys flagged | — | — |
| Wall-clock time, s | — | — |
| Input / output tokens | — | — |
| Model used | — | — |
Report per-category recall as well (security, async, cross-file, left-side). An overall recall can hide the fact that one tool never comments on deleted lines.
How to verify your own setup before trusting it
- Score the answer key against itself. Convert
truth.jsonlinto a comments file (path, side, line = start) and run the scorer. You should see recall 1.0 and anchor validity 1.0. Anything less means a ground-truth line is outside the diff — fix the key, not the reviewers. - Score an empty run. An empty comments file should give recall 0 and no crash.
- Post one hand-written comment on a known-bad line (outside any hunk) to the sandbox PR and confirm the API rejects it. This proves your anchor test reflects the platform's real behaviour.
- Diff the inputs. Confirm every run's PR shows the same head SHA:
gh pr view $PR --json headRefOid.
Failure cases you will likely hit
- Agent output is not valid JSONL. Do not repair it with a second model. Count the run as producing zero line comments and note the format failure; that is a genuine operational result.
- Lines are off by the hunk header. A common agent mistake is counting lines from the top of the hunk, or giving the old-file number with
side: RIGHT. The scorer will show low anchor validity with comments clustered a few lines away from defects — look at the raw lines before concluding the tool “missed” bugs. - Rename-only files. A file renamed without edits has no hunk; any comment on it is unanchorable. Keep seeded defects out of pure renames.
- Context truncation. If the agent's log shows it never read a file, that is a coverage failure, not a reasoning failure. Log which files each run opened if your agent exposes tool calls.
- Rate limits mid-run. Retries inflate wall-clock time and sometimes tokens. Mark affected runs and rerun rather than averaging them in.
- Seed leakage. Variable names like
buggy_offsetor a branch namedadd-bugsleak the answer. Review your own seeds like an attacker would. - Untrusted content in the diff. A reviewer that reads the repository also reads any instructions hidden in it. If you add a prompt injection probe as one of the seeds, score it separately and keep the agent's tools read-only.
Limitations
- Seeded bugs are easier and more local than many real ones. High recall here is necessary, not sufficient.
- One repository and one language say little about others. Repeat on at least two codebases before generalising.
- If the two tools use different models, the test cannot separate model quality from harness quality.
- The prompt for the general agent is a variable. A better prompt may close part of the gap; publish the prompt you used so others can challenge it.
- Precision depends on the adjudicator. Two adjudicators with a disagreement log are better than one.
- Token counts from different sources (tool log vs. provider console) are not always counted the same way. State the source for each figure.
Checklist
- Base and head SHAs, tool versions and model IDs recorded in
bench/ENV.md. - Ground truth outside the reviewed branch; self-score gives 1.0 / 1.0.
- Private sandbox repository; one PR per run on the same head SHA.
- Same input, same model where possible, read-only tools for the agent.
- API rejections recorded as anchor failures, not fixed by hand.
- Blind adjudication of every matched and pending comment.
- At least three runs per tool; median and range reported.
- Prompts, scorer and truth file published alongside the numbers.
When the table is filled, the decision is usually simple: if the specialised reviewer wins on defect-file coverage and anchor validity with similar recall, it is doing what it promises; if the general agent matches it once the prompt asks for per-file iteration, the extra tool may not be worth maintaining. Either way, you will know from your own code rather than from a demo.
More practical material is in the guides; definitions of the terms used here are collected in the glossary.