OpenCodeReview: Why Alibaba Ships Deterministic Code Review, Not Another Agent — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

OpenCodeReview: Why Alibaba Ships Deterministic Code Review, Not Another Agent

Alibaba open-sourced the internal code-review tool that ran across its monorepo for two years. It beats Claude Code on F1 at one-fourteenth the tokens by replacing prompt-driven reasoning with hard engineering constraints on file selection, bundling, rule matching, and comment positioning.

The interesting thing about Alibaba’s open-code-review isn’t that it does code review. Every agent does code review now. The interesting thing is what its authors concluded after two years of running AI code review against their own internal monorepo: the prompt is not the bottleneck. The pipeline is.

The repo, tagged v1.13.1 on September 16 (the day I’m writing this) and Apache-2.0, is the open-source release of AoneCodeReview — the assistant that “over the past two years… has served tens of thousands of developers and identified millions of code defects” inside Alibaba Group. They shipped it as a CLI called ocr, with plugins for Claude Code, Codex, Cursor, and OpenCode, and a benchmark called AACR-Bench (also Apache-2.0, also on Hugging Face at Alibaba-Aone/aacr-bench) that lets you actually measure how well any of these tools do. The headline number on the README is “~1/9 of the tokens” compared to general-purpose agents, with significantly higher precision and F1. On their own leaderboard the best ocr-driven row uses ~1/14.7 the tokens of the equivalent Claude Code row on the same Opus 4.6 model. The trade-off is explicit and accepted: lower recall, on purpose.

That’s a posture worth reading carefully.

The bet, in one sentence

Code review at scale is not a reasoning problem; it is a context-shaping problem. A general-purpose agent asked to “review this diff” has to re-derive from first principles, on every invocation, which files matter, which lines matter, what the rules of the language are, and where to place the comment on the diff. The model is asked to do all of that and understand the change and judge its quality. OpenCodeReview’s bet is that you can hand-engineer the first three questions out of the prompt and let the agent focus on the last one.

The README is unusually direct about this. From the English section:

The root cause: a purely language-driven architecture lacks hard constraints on the review process.

The fix is structured as four deterministic pieces, each one a thing an agent normally gets wrong on big diffs:

  1. Precise file selection — engineering code, not the model, decides which files in a changeset actually need review and which to filter.
  2. Smart file bundling — related files (think message_en.properties and message_zh.properties) get packed into a single review unit. Each bundle runs as a sub-agent with isolated context, which makes the run stable on very large changesets and naturally parallel.
  3. Fine-grained rule matching — a template-engine pass attaches the right review rules to each file’s characteristics before the model sees the file. The model’s attention is narrow on purpose.
  4. External positioning and reflection modules — separate components that take the model’s output and re-anchor it to the correct file/line, then run a reflection pass on the content.

The agent layer is kept for the parts that genuinely require dynamic decision-making: scenario-tuned prompts and a scenario-tuned toolset distilled from large-scale call-trace analysis (call frequency distributions, per-tool repetition rates, the cost of adding a new tool to the overall call chain). The README specifically calls out that the toolset was pruned from a generic agent toolkit, not augmented.

What the AACR-Bench numbers actually show

The public leaderboard ranks OCR and Claude Code across the same set of models. I’m reading the raw imgs/benchmark-en.png directly because the GitHub README links to the image rather than the table, and the image is more current than the prose.

The dataset is concrete: 2,145 review comments extracted from 200 real pull requests across 50 active open-source projects and 10 mainstream programming languages (C++, Rust, Go, Java, C#, TypeScript, Python, JavaScript, Ruby, PHP). Of those 2,145 comments, 1,505 are expert-verified correct and 640 are deliberately-included incorrect — the latter is what makes this a “can you tell a bad review from a good one?” benchmark rather than a “how many comments can you generate?” benchmark. The 80+ senior engineers who annotated the data each had two or more years of experience, and the pipeline was three rounds of cross-validation. The numbers in the chart are precision, recall, line-precision, F1, average wall-clock time, and average tokens.

The interesting rows:

RankSourceModelF1Precision (matches/generated)Recall (matches/1505)Avg TimeAvg Token
1OCRClaude-4.6-Opus25.10%33.90% (301/889)20.00% (301/1505)1m23s385K
2OCRQwen3.8-Max23.00%33.90% (262/774)17.40% (262/1505)5m14s334K
3OCRGLM-5.221.30%32.30% (239/741)15.90% (239/1505)7m58s682K
9Claude CodeClaude-4.8-Opus14.13%15.93% (191/1200)12.70% (191/1505)5m38s2,062K
12Claude CodeClaude-4.6-Opus11.57%7.23% (435/5980)28.90% (435/1505)13m6s5,664K
14CodexGPT-5.58.36%27.82% (74/266)4.92% (74/1505)2m58s525K

The Opus-4.6 head-to-head is the cleanest comparison: identical model, same dataset. OCR: 25.10% F1, 1m23s, 385K tokens. Claude Code: 11.57% F1, 13m6s, 5,664K tokens. That’s a 2.17× lift on F1, 9.4× faster wall-clock, and 14.7× fewer tokens — for a tool whose internal claim is “~1/9”. The “~1/9” is conservative; in this specific row it’s actually ~1/15.

Notice what OCR doesn’t optimize for: recall. The Opus-4.6 OCR row catches 301/1505 = 20.00%. Claude Code catches 435/1505 = 28.90%. Claude Code is finding more valid comments — but it’s also generating 5,980 of them (vs OCR’s 889), so only 7.23% of its output is correct. The noise ratio is 92.77% for Claude Code on that row. OCR’s noise ratio is 66.10%. OCR is what you get when you decide that a code review tool that requires a human to triage a 5,980-comment PR is not a code review tool. The whole pipeline is downstream of that decision.

The Qwen3.7-Max row at rank 10 — Claude Code + Qwen3.7-Max — is also worth staring at: F1 12.17%, recall 23.37% (highest recall on the entire board), 5,153K tokens of output. A model asked to “be thorough” can be thorough, and the cost is on the page.

How the deterministic pieces actually work

I dug into the source tree to see what “deterministic engineering” means in code, not in README prose. The Go layout is:

internal/
  agent/      config/     delegate/   diff/        gitcmd/
  llm/        llmloop/    mcp/        model/       pathutil/
  release/    scan/       session/    stdout/      suggestdiff/
  telemetry/  tool/       viewer/

A few of the packages give the bet away by their names. internal/diff is the git-diff-to-review-unit pipeline — it owns file selection, filtering, and bundling. internal/agent is the agent loop, but importantly it consumes the structured output of diff/ rather than raw diffs. internal/suggestdiff is the comment-positioning module — the README’s “external positioning and reflection modules” are real, separately compiled Go packages, not prompt fragments. internal/llmloop is the layer that calls the model and handles tool use; it’s comparatively thin because most of the prompt-shaping work happens upstream.

The CLI surface is deliberately small. The full install-and-review flow is six commands:

npm install -g @alibaba-group/open-code-review
ocr config provider     # interactive, picks OpenAI-compatible / Anthropic / Bedrock / Azure / Gemini
ocr config model
ocr review                                  # workspace mode: staged + unstaged + untracked
ocr review --from main --to feature-branch  # branch range, merge-base mode
ocr review --commit abc123                  # single commit
ocr scan                                    # full-file audit (no diff)

There are two more modes that aren’t in the quickstart but matter: ocr review --resume <session-id> resumes an interrupted review with the same context (the session viewer at the official docs replays sessions in a browser), and ocr delegate preview / ocr delegate rule <files…> are the entry points for delegation mode, which I’ll come back to.

The output format is structured JSON when you ask for it (ocr review --format json --output result.json), which is what makes the tool composable with AI host agents — the host can ingest the result without parsing free-form review prose.

Delegate mode — the bet extends to the agent

The most interesting thing on the roadmap is Delegate Mode, which is also shipping now (it’s in the v1.13 release), not just planned. The pitch: in delegation mode, ocr itself doesn’t call any LLM. It does the deterministic pieces — file selection, bundling, rule matching, exclude resolution, background-context injection, diff collection — and hands the assembled review task to the host agent (Claude Code, Codex, Cursor) as a structured prompt. The host uses its own agent loop and its own subscription budget.

This is a sharp move. The current OCR tool has to handle a separate billing relationship with Anthropic or whoever — you configure a provider, you set an API key, you pay per token. Delegate mode means OCR is useful to people who already pay Claude Code $20/month and don’t want a second API key. The pipeline that handles context shaping runs locally and free; the model that does the actual reasoning runs on a subscription the user already has.

The two extremes this opens up: a small team that doesn’t want a code-review LLM bill at all can run delegate mode and get most of the precision/F1 win (the deterministic pipeline still applies) on top of their existing Claude Code subscription. A large team that wants OCR’s full pipeline and its own model routing can run the normal OCR-managed mode and pick the best model on the AACR-Bench leaderboard.

The roadmap also flags two more modes worth noting. Ultra Mode is opt-in for security-sensitive changesets: higher recall at the cost of more tokens and time. Domain-Specific Long-Term Memory (planned for H1 2027) is the long-game answer to one of the genuine hard problems in code review — “this codebase already decided X about Y, the agent keeps flagging it.” Without memory, every review re-derives the project’s conventions from scratch.

What the benchmark is honest about

Reading the AACR-Bench README and metrics carefully, a few things stand out as worth taking seriously rather than as marketing:

The noise-injection design. 640 of the 2,145 comments in the test set are known-incorrect. A benchmark that only contains correct comments rewards models that confidently emit many comments; a benchmark that includes deliberate noise rewards models that know when to be quiet. OCR’s row-by-row pattern (high precision, lower recall) is exactly what you’d expect a system that’s been optimized against a noise-aware benchmark to do. Claude Code’s row-by-row pattern (low precision, higher recall) is what you get from a system that hasn’t.

The line-precision metric is separate from the content metric. positive_line_match_rate measures whether the comment is anchored to the right lines; positive_match_rate measures whether the content is right. The two can diverge sharply: a model can produce a correct observation attached to the wrong line, which in a real PR review is worse than no comment. The “external positioning and reflection modules” are the OCR-specific answer to this — the README explicitly says these modules “systematically improve both the location accuracy and content accuracy of AI feedback.”

The evaluation harness is published as code, not a screenshot. The evaluation/ directory under aacr-bench ships a pipeline.py + a judge.py + a per-reviewer adapter (reviewers/ocr.py, reviewers/claude.py, reviewers/codex.py), all JSONL-based with a documented schema. Anyone can rerun any row. The data is on Hugging Face under Apache-2.0. This is closer to a real benchmark than most “benchmarks” in the LLM space.

The data is multilingual on purpose. Ten languages, including system-level (C++, Rust, Go), enterprise (Java, C#, TypeScript), and scripting (Python, JavaScript, Ruby, PHP). Most code-review benchmarks are heavily Python-weighted; AACR-Bench’s C++ and Rust coverage is what lets the leaderboard separate code-review skill from “good at Python” skill.

There is an ASSURANCE_CASE.md. This is unusual. It’s a 200+ line document that lays out the threat model (five trust boundaries: Git→CLI, CLI→LLM, CLI→output, Browser→Viewer, Network), the threat summary table, and the mitigation table. For an OSS CLI that reads your diffs and sends them to an external API, this is the document you actually want before you run it on a private repo. The mitigation list — strict format validation on parsed diffs, HTTPS-only API calls, file-write scope constrained to the working directory, host-header allowlist on the local viewer to block DNS rebinding — is the same set of constraints you’d want to see in any production code-review tool, and the fact that they shipped as a markdown file rather than a blog post matters.

What it doesn’t fix

Three honest limits that the README and benchmark together make clear.

Recall is the deliberate sacrifice. Across the OCR leaderboard rows, the highest recall is 20.00% (Opus 4.6). That means 1,204 of the 1,505 expert-verified defects in the test set are not caught by OCR on its best model. For high-stakes reviews (security-sensitive changesets, kernel code, anything where a missed defect is expensive), the Ultra Mode roadmap item is the response — but at the time of writing it’s planned, not shipped. Until Ultra Mode lands, teams that need high recall should pair OCR with a broader scan or a human reviewer on the categories OCR is known to miss.

The benchmark is one slice of “code review.” AACR-Bench measures whether a tool can produce a correct, well-anchored review comment on a real PR diff. It does not measure whether a tool can keep up with a 50-file refactor PR, whether its comments remain useful across a long-running branch, whether it understands the intent of a change versus just its diff, or whether it can integrate with a team’s existing comment-resolution workflow (the Session Viewer helps here, but it’s a viewer, not a workflow tool). The OpenSSF Best Practices Gold badge the project carries is a sign of hygiene, not a guarantee of fit.

OCR’s competitive advantage is the pipeline, not the model. Re-rank the leaderboard by token efficiency and OCR wins across all the models it appears on. Re-rank by raw F1, and a Claude Code row on a stronger model could in principle beat OCR’s row 1 — though no such row currently exists on the published leaderboard. The bet is that the pipeline will keep OCR ahead of any single-model configuration. If you took the OCR pipeline and ran it against a hypothetical frontier model that materially beats Opus 4.6, you’d expect OCR’s row 1 to move up. If you took Claude Code’s agent loop and ran it against the same frontier model, you’d expect Claude Code to move up too — but slower, because the agent loop has more tokens to spend per review. The pipeline is the multiplier, and OCR is betting on the multiplier rather than on the model.

Why this matters for the rest of the agent ecosystem

Most of the agent-harness discussion in 2026 has been about better prompts, longer context, stronger tool use, or stronger models. OpenCodeReview is a counterpoint: a substantial fraction of the agent-quality problem in code review is upstream of the model. File selection, bundling, rule matching, and comment positioning can be done by deterministic code that doesn’t drift, doesn’t burn tokens, and doesn’t need a GPU. The agent layer is then free to focus on the parts that genuinely require language understanding — interpreting the change, judging its quality, deciding what to say.

This generalizes. Any domain where the agent’s failure mode is “the prompt didn’t include the right scaffolding” is a candidate for the OCR pattern: pull the scaffolding out of the prompt and into code. The list of candidate domains is long — code review, security audit, PR triage, test generation, log analysis, on-call summarization. Most of them currently run as “ask the model to do the whole thing.” Most of them have the same shape of failure: too many false positives, too many tokens, comments landing on the wrong lines.

OpenCodeReview is a clean, measurable instance of that pattern working. Two years of internal validation, a published benchmark with concrete numbers, an Apache-2.0 license, an MCP server, plugins for the four major coding agents, and a roadmap that takes the bet seriously enough to plan memory and ultra-recall modes for 2027. The “1/9 the tokens” claim is conservative; the real number is closer to 1/15 on the head-to-head model. The precision/recall trade-off is explicit and accepted. The threat model is documented.

Whether you should adopt it depends on whether your team’s pain with code review is “we generate too much noise to triage” (OCR is a great fit), “we miss too many real defects” (wait for Ultra Mode), or “we want one tool that does both” (no such tool exists yet — not even, frankly, the agents this one is being measured against).

I keep coming back to the line in the README: “the root cause: a purely language-driven architecture lacks hard constraints on the review process.” That’s a sentence that could have been written about any of a dozen agent tasks. Alibaba wrote it about code review, built a tool around the principle, and shipped a benchmark that lets you check whether they were right. So far, on the public leaderboard, they were.

Where to dig further

  • alibaba/open-code-review — the CLI, plugins for Claude Code / Codex / Cursor / OpenCode, MCP server, the ASSURANCE_CASE.md, and the ROADMAP.md that explains what’s next.
  • alibaba/aacr-bench — the benchmark repo, including the evaluation/ pipeline that reproduces any leaderboard row and the JSONL schema for the dataset.
  • Alibaba-Aone/aacr-bench on Hugging Face — the 2,145-comment test set under Apache-2.0, with the 80+-engineer annotation methodology.
  • arXiv:2601.19494 — AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context, the underlying paper by Lei Zhang, Yongda Yu, Minghui Yu, et al.
  • open-codereview.ai/docs — official docs for CLI Reference, Review Rules, MCP server, CI/CD integration, and the Session Viewer.
  • npm: @alibaba-group/open-code-review — the install target if you want to run it locally against your own repo.

One last thing worth noting: the project’s CLAUDE.md, AGENTS.md, and .claude-plugin/ directory are all present in the repo root. The README explicitly says OCR ships a Claude Code plugin with review slash commands, a Codex plugin with callable review skills, and a portable agent skill that works on any skill-compatible host. The point isn’t just that the tool exists — it’s that the tool is designed to be the code-review layer of whatever agent stack you already run. Whether that layer is worth its slot is now a measurable question, and the benchmark to answer it is on Hugging Face.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.