Nobody’s agreed on how to score an AI code reviewer. Not even close.
Qodo says it straight out: “There is no standard benchmark for AI code review yet.” What exists instead is a scatter of open-source and vendor-run tests. Each one measures something different. Each one scores it other.
And, as Qodo points out about the public tests from 2025 and 2026, they keep pulling from the same overlapping repositories.
If you’re a small team weighing whether a review bot earns its seat cost, that’s the whole situation in one line: the yardsticks contradict each other. So when a vendor says “we caught X% of issues,” that number is empty until you know which yardstick produced it.
Here’s what each benchmark genuinely measures, where the problems hide.
And the way I’d test review tooling on my own builds without getting sold anything.
Nobody agrees on the yardstick
I run AI automation across my own projects, and code review is the bit I trust least.
A generation benchmark proves a model can write code. That’s it. It won’t tell you whether the bot reading your pull request catches the bug that ships at 2 a.m. on a Friday.
Reviewing is a separate skill. Pattern-matching across a diff. Deciding if a “finding” is real or noise. Knowing when saying nothing is the correct answer.
That last one matters more than vendors let on. A bot that flags everything is worse than no bot at all.
Your human reviewers start skimming, and skimming is how stuff slips through.
Every benchmark below is an attempt to quantify this. Each one makes its own tradeoffs doing it.
What each test measures
Martian’s open-source benchmark gives you the whole thing: pull requests, golden comments, LLM judge prompts, evaluation pipeline. All of it out in the open. It’s built on 50 PRs from five major open-source projects, carrying 173 human-verified golden comments labeled by severity — Low, Medium, High, Critical. And tagged by category: bug, security, concurrency, data, API, performance, test gap, documentation defect, style, and speculative. The tools it’s put through the wringer run from Augment, Claude Code, CodeRabbit, and Cursor Bugbot through Devin, Gemini, GitHub Copilot, GitLab Duo, and Sourcery. Smartest move in its design: a continuously updated online benchmark that samples fresh real-world PRs off GitHub, ones where review bots already left comments. Built specifically to dodge training-data leakage.
Qodo’s Code Review Benchmark 1.0 inverts the approach. Starts with 100 real merged PRs, modifies them with 580 human-validated injected issues spread across eight repositories and seven languages. Scores product-level precision, recall, and F1 for bugs plus repository-specific standards. Qodo names its own two weak spots in the glossary: it’s vendor-run, and injected issues might not perfectly match the defects developers actually write. That second one’s serious. Coming back to it.
SWE-PRBench works with 350 merged PRs that have substantive human review comments as ground truth, run under three frozen context configurations. Tests base-model issue detection, hallucination behavior, and how models handle separate context formats.
CodeReviewBench is the smallest of the bunch: 30 merged PRs from five open-source projects, yielding 95 confirmed bugs, versioned outputs, everything scored under a single review runner. Tracks precision, recall, F1, and cost per PR. Qodo’s notes list its limits. Small sample, one run per model, and it compares models inside one runner rather than complete products.
GitHub’s ReviewBench holds 219 public pull requests across 19 languages, aligned to GitHub-wide distributions while keeping substantive review cases intact.
And then there’s c-CRAB, an arXiv-published dataset for evaluating agents on code-review tasks, if the academic angle’s your entry point.
Planted bugs aren’t real bugs
Here’s where the tests split, hard: ground truth quality beats sample size, every time.
Qodo plants 580 issues into merged PRs. Gets scale from it.
Gets severity control from it.
But a planted bug is its own creature. Think cooking-show prep versus a working kitchen. The show vegetable is cut clean, uniform, camera-ready. The working cook’s board is a mess of odd shapes and peels. Planted issues come out clean, self-contained, easy to spot. Real defects tangle up with context, style, intent.
A bot can nail planted issues and still miss the subtle thing a senior human flags in one line, in passing.
Martian, SWE-PRBench, and CodeReviewBench all use human review comments as ground truth.
Messier to score, sure. Human comments don’t line up tidy the way planted bugs do. But it measures what you actually care about. Did the bot catch what a real reviewer caught?
So when a vendor throws a benchmark number at you, check the ground truth method first. Injected issues or real reviewer comments. Decide from that. A 90% catch rate on planted bugs and a 40% catch rate on organic defects can both be true simultaneously. Both numbers can sit in the same sales deck.
Vendors grade their own homework
Now the overlap.
Qodo sells review tooling and runs its own benchmark, and to its credit says “vendor-run” out loud. Martian’s benchmark evaluates Qodo’s product. GitHub runs ReviewBench while selling Copilot. And Copilot shows up in Martian’s evaluated list too. No neutral referee anywhere in this picture.
Not an accusation of bad faith. Open-sourcing the datasets, judge prompts, and pipeline. Which Martian does — is genuinely good practice, and it means you can rerun scores yourself. Still, the benchmark worth trusting is the one you can audit. Not the one with the flashiest leaderboard.
Vendor won’t say which benchmark their numbers came from? That number is marketing. Not measurement.
Run one benchmark yourself
You don’t have budget for five benchmarks. Fine. Here’s the stripped-down version.
Ground truth method comes first.
Human-verified comments beat injected issues for predicting real-world performance — that’s my call. And I’d put a tooling decision on it.
Then pick one open benchmark and run your candidates yourself. Martian’s is fully open, pipeline included.
Score your shortlist against those same 50 PRs and 173 golden comments instead of swallowing self-reported figures.
Weight severity, not raw counts. A tool that catches 90% of style nits while missing Critical security findings is a liability wearing a nice dashboard. The severity labels exist precisely so you can slice scores that way.
Measure noise yourself.
Flip the bot on for two weeks of PRs, then count how many flagged comments your human reviewers waved off. That’s your personal false-positive rate, and it’s worth more than any public score you’ll ever read.
And push on leakage.
A static dataset goes stale and seeps into training data. A continuously updated online benchmark is the design that stays honest over time.
Stop trusting numbers you didn’t verify
AI code review benchmarking is young, fragmented, and part self-graded.
That isn’t a reason to skip review bots.
It’s a reason to grade them yourself with the open material that already exists.
Pull the Martian benchmark’s datasets and pipeline. Check whether any vendor claim traces back to injected issues or real reviewer comments. Measure your own false-positive rate in production. If you want the complete picture, read Qodo’s benchmark glossary for its honest limitation list, GitHub’s ReviewBench announcement for the distribution argument. And the c-CRAB paper for the academic framing.
The tools will improve. The benchmarks will keep fighting it out.
Your job stays the same either way: quit accepting numbers you never verified, ’cause nobody’s running that check for you.
