Специализированный AI-ревьюер против универсального агента: практический тест

Специализированный AI-ревьюер против универсального агента: практический тест
Temporary fallback cover; replace in editorial pass.

Intermediate · 40 min read · Updated 7 October 2026

Ask a general-purpose AI agent to review a 60-file change and you will usually get a confident summary. What you often don't get is a comment on file 47, a comment attached to the line that actually contains the bug, or a token bill proportional to the work. This article gives you a test you can run on your own code to find out whether a dedicated reviewer such as Open Code Review does better — and by how much — instead of trusting either vendor's demo.

The test plants known defects into a large diff, runs both reviewers under matched conditions, and scores six things: precision, recall, file coverage, anchor validity, wall-clock time and tokens. Everything is scripted, so a second person can repeat it and get comparable numbers.

The problem in one paragraph

Code review of a large change is a coverage task disguised as a reasoning task. A general agent has to decide which files to open, keep enough of each one in its context window, and then translate “this is wrong” into a pull request comment with a valid path, line and side. Three failure modes follow directly: files that are never opened, comments that land on the wrong line (or are rejected by the platform and silently become a general comment), and a high token count from re-reading the same files. A specialised reviewer claims to fix these by owning the iteration over files and the mapping to diff positions. The claim is testable.

The concrete case

Use one service you know well — the example below assumes a mid-sized Python web service — and build one deliberately large change on top of a fixed base commit:

Place defects deliberately to probe the known weaknesses:

PlacementWhyExample defect
Last files in path orderTests whether the reviewer reaches the end of the diffzz_reports/export.py: off-by-one in pagination
Deep inside a 600-line fileTests attention within a long fileMissing await on an async DB call
Deleted lines only (left side)Tests anchoring to removed codeRemoved permission check in a view
Two-file interactionTests cross-file reasoningFunction now returns cents, caller still assumes rupees/dollars
Test fileTests whether tests are reviewed at allAssertion changed to always pass
Config / non-code fileTests coverage beyond source codeTimeout changed from seconds to milliseconds without unit change

Step 0. Pin the environment

The comparison is only meaningful if the reviewers differ in tooling, not in model, temperature or input. Write the environment down before the first run.

# bench/ENV.md — fill in before the first run
base_sha:            <git rev-parse base>
head_sha:            <git rev-parse bench/head>
open_code_review:    <version or commit you installed>
general_agent:       <name + version, e.g. output of `claude --version`>
model (both):        <same model id for both tools, if OCR lets you choose>
temperature:         <value, or "tool default — not configurable">
machine / network:   <runner type, region>
date:                <YYYY-MM-DD>

If Open Code Review cannot use the same model as the agent, say so in the results. You are then comparing two systems, not two harnesses, and the conclusion must be narrower.

Step 1. Build the benchmark change and the ground truth

git switch -c bench/head base
# ... apply noise edits, seeded defects and decoys ...
git commit -am "bench: large change with seeded defects"

mkdir -p bench/runs
git diff --no-color --no-ext-diff -U3 base bench/head > bench/change.diff
git diff --stat base bench/head | tail -1

Record every seeded defect as one line of JSONL. Line numbers refer to the new file for RIGHT and the old file for LEFT, matching how GitHub addresses review comments.

{"id":"D01","path":"zz_reports/export.py","side":"RIGHT","start":88,"end":90,"category":"logic","severity":"high","summary":"page offset uses page*size+1"}
{"id":"D02","path":"app/orders/service.py","side":"RIGHT","start":412,"end":412,"category":"async","severity":"high","summary":"missing await on repo.save"}
{"id":"D03","path":"app/views/admin.py","side":"LEFT","start":57,"end":59,"category":"security","severity":"critical","summary":"permission check removed"}
{"id":"D04","path":"app/billing/money.py","side":"RIGHT","start":21,"end":24,"category":"cross-file","severity":"high","summary":"returns minor units; see D04b caller"}

Keep bench/truth.jsonl out of the reviewed branch. If a reviewer can read the answer key, the test is void. Also strip commit messages and code comments that hint at the bugs (“TODO: fix pagination”).

Step 2. Use a sandbox PR as the common output channel

Comparing free-text output from two tools is where most informal tests go wrong. The cleanest neutral format is the platform itself: let each run post review comments to its own pull request in a private sandbox repository, then pull the comments back through the API. This also tests anchoring honestly, because GitHub rejects review comments whose line is not part of the diff.

# one branch (and one PR) per tool per run, all pointing at the same commit
for run in ocr-r1 ocr-r2 ocr-r3 agent-r1 agent-r2 agent-r3; do
  git branch "bench/$run" bench/head
  git push sandbox "bench/$run"
  gh pr create --repo OWNER/SANDBOX --base base --head "bench/$run" \
     --title "bench $run" --body "benchmark run, do not merge"
done

Do not run this against a public or shared repository: posted comments are visible to everyone with access and may be indexed or emailed to watchers.

Step 3. Run Open Code Review

Install it following the project's own README and record the exact version. Because invocation details differ between releases, the script below wraps it in a variable rather than guessing flags.

# bench/run_ocr.sh — example wrapper; set OCR_CMD from the README of your version
set -euo pipefail
RUN="$1"; PR="$2"
start=$(date +%s)
$OCR_CMD "$PR" 2>&1 | tee "bench/runs/$RUN.log"
end=$(date +%s)
echo "{\"run\":\"$RUN\",\"wall_seconds\":$((end-start))}" > "bench/runs/$RUN.time.json"

Capture token usage from whatever the tool reports. If it reports nothing, read usage from your model provider's console for an isolated API key used only for this run, and note that the figure was taken from the provider side.

Step 4. Run the general-purpose agent under matched conditions

Give the agent the same input (the PR, or bench/change.diff plus read access to the checkout), the same model if possible, and a prompt that asks for the same output. Do not coach it with hints you would not give a human reviewer.

# bench/prompts/review.md
You are reviewing the changes between `base` and `bench/head` in this repository.
Review every changed file. For each real problem, output one JSON object per line:
{"path": "...", "line": N, "side": "RIGHT|LEFT", "severity": "low|medium|high|critical", "body": "..."}
`line` is the new-file line for RIGHT and the old-file line for LEFT, and must be a line
that appears in the diff. Output nothing except these JSON lines.

Example for Claude Code in headless mode, with read-only tools. Verify the flags with claude --help for your installed version:

start=$(date +%s)
claude -p "$(cat bench/prompts/review.md)" \
  --output-format json \
  --allowedTools "Read,Grep,Glob,Bash(git diff:*)" \
  > bench/runs/agent-r1.raw.json
end=$(date +%s)
echo "{\"run\":\"agent-r1\",\"wall_seconds\":$((end-start))}" > bench/runs/agent-r1.time.json

The JSON envelope contains the agent's final text and a usage block; extract the comment lines and the token counts from it. Then post the comments to the agent's sandbox PR with the same API the specialised tool uses, so both are judged by the platform's anchoring rules:

gh api repos/OWNER/SANDBOX/pulls/$PR/reviews --method POST --input review.json
# review.json: {"event":"COMMENT","commit_id":"<head_sha>","comments":[{"path":..,"line":..,"side":..,"body":..}]}

If the API rejects the review because one comment has an invalid line, do not fix it by hand. Record the rejection, drop the invalid comment, post the rest, and count the dropped one as an anchor failure. A real integration would have faced the same error.

Step 5. Normalize all comments to one file per run

gh api repos/OWNER/SANDBOX/pulls/$PR/comments --paginate \
  --jq '.[] | {path, line, side, body}' -c > bench/runs/$RUN.comments.jsonl

# general (non-line) comments: kept separately, they cannot be anchored
gh api repos/OWNER/SANDBOX/pulls/$PR/reviews --paginate \
  --jq '.[] | select(.body != "") | {body}' -c > bench/runs/$RUN.general.jsonl

Findings that only appear in a general review body count toward recall only after manual adjudication, and never toward anchor validity.

Step 6. Score

The script parses the unified diff to learn which lines are commentable on each side, then matches comments to ground truth by path, side and line with a small tolerance.

#!/usr/bin/env python3
"""Score normalized review comments against seeded ground truth."""
import argparse
import collections
import json
import re

HUNK = re.compile(r"^@@ -(\d+)(?:,\d+)? \+(\d+)(?:,\d+)? @@")


def commentable_lines(diff_path):
    lines = collections.defaultdict(lambda: {"LEFT": set(), "RIGHT": set()})
    path, old, new, in_hunk = None, 0, 0, False
    with open(diff_path, encoding="utf-8", errors="replace") as f:
        for raw in f:
            line = raw.rstrip("\n")
            if line.startswith("diff --git "):
                path, in_hunk = None, False
                continue
            if not in_hunk and line.startswith(("--- ", "+++ ")):
                p = line[4:]
                if p != "/dev/null":
                    path = p[2:] if p[:2] in ("a/", "b/") else p
                continue
            m = HUNK.match(line)
            if m:
                old, new, in_hunk = int(m[1]), int(m[2]), True
                continue
            if not in_hunk or path is None:
                continue
            if line.startswith("+"):
                lines[path]["RIGHT"].add(new)
                new += 1
            elif line.startswith("-"):
                lines[path]["LEFT"].add(old)
                old += 1
            elif line.startswith(" "):
                lines[path]["LEFT"].add(old)
                lines[path]["RIGHT"].add(new)
                old += 1
                new += 1
    return lines


def load_jsonl(path):
    with open(path, encoding="utf-8") as f:
        return [json.loads(l) for l in f if l.strip()]


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--diff", required=True)
    ap.add_argument("--truth", required=True)
    ap.add_argument("--comments", required=True)
    ap.add_argument("--adjudicated", help="JSONL of {comment_idx, verdict: valid|invalid}")
    ap.add_argument("--tolerance", type=int, default=3)
    a = ap.parse_args()

    valid = commentable_lines(a.diff)
    truth = load_jsonl(a.truth)
    comments = load_jsonl(a.comments)
    verdicts = {}
    if a.adjudicated:
        verdicts = {r["comment_idx"]: r["verdict"] for r in load_jsonl(a.adjudicated)}

    found, useful, anchored, pending = set(), 0, 0, []
    for i, c in enumerate(comments):
        side, line = c.get("side") or "RIGHT", c.get("line")
        if isinstance(line, int) and line in valid.get(c["path"], {}).get(side, set()):
            anchored += 1
        hit = next((d["id"] for d in truth
                    if d["path"] == c["path"] and d.get("side", "RIGHT") == side
                    and isinstance(line, int)
                    and d["start"] - a.tolerance <= line <= d["end"] + a.tolerance), None)
        if hit and verdicts.get(i) != "invalid":
            found.add(hit)
            useful += 1
        elif verdicts.get(i) == "valid":
            useful += 1
        elif i not in verdicts:
            pending.append(i)

    n = len(comments)
    defect_files = {d["path"] for d in truth}
    commented_files = {c["path"] for c in comments}
    print(json.dumps({
        "comments": n,
        "precision": round(useful / n, 3) if n else None,
        "recall": round(len(found) / len(truth), 3),
        "anchor_validity": round(anchored / n, 3) if n else None,
        "diff_files_commented": f"{len(commented_files & set(valid))}/{len(valid)}",
        "defect_files_commented": f"{len(commented_files & defect_files)}/{len(defect_files)}",
        "missed": sorted(d["id"] for d in truth if d["id"] not in found),
        "pending_adjudication": pending,
    }, ensure_ascii=False, indent=2))
    if pending:
        print("WARNING: precision is a lower bound until pending comments are adjudicated")


if __name__ == "__main__":
    main()
python3 bench/score.py --diff bench/change.diff --truth bench/truth.jsonl \
  --comments bench/runs/ocr-r1.comments.jsonl --adjudicated bench/runs/ocr-r1.adj.jsonl

What each number means:

Step 7. Adjudicate blind

Location matching is a filter, not a judgement: a comment three lines from D02 about a typo will match D02. Have a second person read every matched and pending comment without knowing which tool wrote it and record a verdict:

{"comment_idx": 0, "verdict": "valid"}
{"comment_idx": 4, "verdict": "invalid", "note": "near D02 but about naming, not the missing await"}

Shuffle comments from all runs into one sheet with tool names removed, then split verdicts back per run. This takes longer than the automated part and is the step that makes the numbers trustworthy.

Step 8. Repeat and report variance

Both tools are non-deterministic. Run each at least three times on the same commit and report the median plus min–max. A difference smaller than the spread between runs of the same tool is not a difference.

Results template

Fill this in from your own runs. The empty cells are intentional: we are not publishing numbers we did not measure on your code.

Metric (median, min–max over runs)Open Code ReviewGeneral agent
Precision——
Recall (seeded defects)——
Defect files commented——
Anchor validity (incl. API rejections)——
Decoys flagged——
Wall-clock time, s——
Input / output tokens——
Model used——

Report per-category recall as well (security, async, cross-file, left-side). An overall recall can hide the fact that one tool never comments on deleted lines.

How to verify your own setup before trusting it

  1. Score the answer key against itself. Convert truth.jsonl into a comments file (path, side, line = start) and run the scorer. You should see recall 1.0 and anchor validity 1.0. Anything less means a ground-truth line is outside the diff — fix the key, not the reviewers.
  2. Score an empty run. An empty comments file should give recall 0 and no crash.
  3. Post one hand-written comment on a known-bad line (outside any hunk) to the sandbox PR and confirm the API rejects it. This proves your anchor test reflects the platform's real behaviour.
  4. Diff the inputs. Confirm every run's PR shows the same head SHA: gh pr view $PR --json headRefOid.

Failure cases you will likely hit

Limitations

Checklist

When the table is filled, the decision is usually simple: if the specialised reviewer wins on defect-file coverage and anchor validity with similar recall, it is doing what it promises; if the general agent matches it once the prompt asks for per-file iteration, the extra tool may not be worth maintaining. Either way, you will know from your own code rather than from a demo.

More practical material is in the guides; definitions of the terms used here are collected in the glossary.

We publish what works for us—and implement the same solutions for your business. We design AI automation, Telegram bots, chats, and AI agents for real-world processes. Discuss your project →