ReviewBench
We show this table for reference; we do not rank on it.
Measures how reliably code-review agents identify useful issues in real pull requests while avoiding false positives.
Evaluation results published by GitHub.
Grounded F1 on ReviewBench — 2026-10-07 capture
We mirror the published grounded f1 view for ReviewBench. Copilot Code Review · Balanced leads the public snapshot at 40.1%, followed by Devin AI (37.0%) and Qodo (35.1%). We do not use these results to rank models overall.
Copilot Code Review · Balanced
Copilot Code Review
Balanced
Devin AI
Devin AI
Qodo
Qodo
28 reviewer configurationsAgenticCurrentDisplay onlyUpdated 2026-10-07 capture
| Source rank | Reviewer / configuration | Grounded F1 | Grounded precision | Grounded recall | Augmented F1 | Augmented precision | Augmented recall | Completed PRs / repeats |
|---|---|---|---|---|---|---|---|---|
| 1 | Copilot Code Review · Balancedeffort=mediumRun date: · official-2026-09 | 40.1% | 87.8% ± 0.5 | 26.0% ± 0.6 | 49.7% | 87.4% ± 2.0 | 34.7% ± 0.2 | 219 / 2193 repeats |
| 2 | Devin AIRun date: · official-2026-09 | 37.0% | 84.0% ± 0.7 | 23.8% ± 0.1 | 47.3% | 80.4% ± 0.2 | 33.5% ± 0.3 | 219 / 2193 repeats |
| 3 | QodoRun date: · official-2026-09 | 35.1% | 85.3% ± 0.8 | 22.1% ± 0.9 | 44.5% | 86.8% ± 1.1 | 29.9% ± 1.1 | 218-219 / 2193 repeats |
| 4 | Codex · GPT-5.6 Sol · Maxmodel=gpt-5.6-sol · reasoning=maxRun date: · official-2026-07 | 32.2% | 86.8% ± 0.3 | 19.7% ± 0.9 | 43.1% | 89.1% ± 1.3 | 28.4% ± 1.3 | 215 / 2193 repeats |
| 5 | Codex · GPT-5.6 Sol · Ultramodel=gpt-5.6-sol · reasoning=ultraRun date: · official-2026-09 | 29.6% | 85.9% ± 1.6 | 17.9% ± 0.4 | 38.4% | 87.3% ± 1.2 | 24.6% ± 0.1 | 219 / 2193 repeats |
| 6 | Codex · GPT-5.6 Sol · xHighmodel=gpt-5.6-sol · reasoning=xhighRun date: · official-2026-07 | 29.5% | 88.8% ± 0.9 | 17.7% ± 0.5 | 40.4% | 90.5% ± 0.8 | 26.0% ± 0.5 | 215 / 2193 repeats |
| 7 | Copilot Code Review · Liteeffort=lowRun date: · official-2026-07 | 28.6% | 83.8% ± 0.9 | 17.2% ± 0.4 | 37.2% | 83.7% ± 0.3 | 23.9% ± 0.9 | 216-218 / 2193 repeats |
| 8 | Codex · GPT-5.6 Luna · Maxmodel=gpt-5.6-luna · reasoning=maxRun date: · official-2026-07 | 27.4% | 87.7% ± 1.5 | 16.2% ± 0.4 | 36.0% | 88.5% ± 1.3 | 22.6% ± 0.2 | 215 / 2193 repeats |
| 9 | CubicRun date: · official-2026-07 | 27.3% | 85.5% | 16.3% | 37.4% | 85.1% | 24.0% | 219 / 2191 repeat |
| 10 | GreptileRun date: · official-2026-09 | 27.2% | 86.1% | 16.2% | 32.8% | 85.6% | 20.3% | 218 / 2191 repeat |
| 11 | Codex · GPT-5.6 Sol · Highmodel=gpt-5.6-sol · reasoning=highRun date: · official-2026-07 | 26.8% | 89.0% ± 2.3 | 15.8% ± 0.0 | 36.5% | 90.8% ± 0.7 | 22.8% ± 1.2 | 215 / 2193 repeats |
| 12 | Codex · GPT-5.6 Terra · Ultramodel=gpt-5.6-terra · reasoning=ultraRun date: · official-2026-07 | 25.1% | 90.0% ± 1.9 | 14.6% ± 0.8 | 33.8% | 90.1% ± 1.4 | 20.8% ± 0.9 | 215 / 2193 repeats |
| 13 | Codex · GPT-5.6 Terra · Maxmodel=gpt-5.6-terra · reasoning=maxRun date: · official-2026-07 | 25.0% | 90.2% ± 1.3 | 14.5% ± 0.2 | 34.1% | 90.0% ± 1.3 | 21.1% ± 0.3 | 215 / 2193 repeats |
| 14 | Codex · GPT-5.6 Sol · Mediummodel=gpt-5.6-sol · reasoning=mediumRun date: · official-2026-07 | 23.7% | 88.5% ± 2.0 | 13.7% ± 0.4 | 30.1% | 90.5% ± 1.9 | 18.0% ± 0.4 | 215 / 2193 repeats |
| 15 | Codex · GPT-5.4 · xHighmodel=gpt-5.4 · reasoning=xhighRun date: · official-2026-07 | 21.8% | 88.7% ± 1.5 | 12.4% ± 0.2 | 30.9% | 86.5% ± 0.6 | 18.8% ± 0.1 | 219 / 2193 repeats |
| 16 | Codex · GPT-5.4 · Highmodel=gpt-5.4 · reasoning=highRun date: · official-2026-07 | 21.0% | 88.0% ± 0.4 | 11.9% ± 0.4 | 28.3% | 86.7% ± 0.9 | 16.9% ± 0.2 | 219 / 2193 repeats |
| 17 | Codex · GPT-5.6 Terra · xHighmodel=gpt-5.6-terra · reasoning=xhighRun date: · official-2026-07 | 19.8% | 90.1% ± 1.7 | 11.1% ± 0.7 | 25.6% | 89.3% ± 1.3 | 14.9% ± 0.9 | 215 / 2193 repeats |
| 18 | Codex · GPT-5.5 · xHighmodel=gpt-5.5 · reasoning=xhighRun date: · official-2026-07 | 18.8% | 88.6% ± 1.1 | 10.5% ± 0.8 | 24.8% | 87.1% ± 1.2 | 14.4% ± 0.2 | 219 / 2193 repeats |
| 19 | Codex · GPT-5.6 Sol · Lowmodel=gpt-5.6-sol · reasoning=lowRun date: · official-2026-07 | 17.6% | 90.2% ± 0.9 | 9.7% ± 0.4 | 21.1% | 91.4% ± 1.1 | 11.9% ± 0.3 | 215 / 2193 repeats |
| 20 | CursorRun date: · official-2026-09 | 17.3% | 87.7% ± 0.8 | 9.6% ± 0.3 | 21.7% | 89.3% ± 0.8 | 12.3% ± 0.5 | 219 / 2193 repeats |
| 21 | Codex · GPT-5.4 · Mediummodel=gpt-5.4 · reasoning=mediumRun date: · official-2026-07 | 17.1% | 88.0% ± 0.8 | 9.5% ± 0.5 | 23.5% | 85.1% ± 1.4 | 13.6% ± 0.3 | 219 / 2193 repeats |
| 22 | Codex · GPT-5.6 Terra · Highmodel=gpt-5.6-terra · reasoning=highRun date: · official-2026-07 | 16.6% | 88.0% ± 1.5 | 9.2% ± 0.3 | 20.3% | 85.4% ± 1.3 | 11.5% ± 0.3 | 215 / 2193 repeats |
| 23 | Codex · GPT-5.5 · Highmodel=gpt-5.5 · reasoning=highRun date: · official-2026-07 | 16.1% | 85.7% ± 0.3 | 8.9% ± 0.1 | 20.3% | 82.4% ± 1.5 | 11.6% ± 0.2 | 219 / 2193 repeats |
| 24 | Codex · GPT-5.4 · Lowmodel=gpt-5.4 · reasoning=lowRun date: · official-2026-07 | 14.0% | 88.7% ± 0.5 | 7.6% ± 0.3 | 19.3% | 82.4% ± 0.5 | 10.9% ± 0.9 | 219 / 2193 repeats |
| 25 | Codex · GPT-5.5 · Mediummodel=gpt-5.5 · reasoning=mediumRun date: · official-2026-07 | 14.0% | 83.7% ± 1.4 | 7.6% ± 0.3 | 16.8% | 77.8% ± 1.7 | 9.4% ± 0.3 | 219 / 2193 repeats |
| 26 | Codex · GPT-5.6 Terra · Mediummodel=gpt-5.6-terra · reasoning=mediumRun date: · official-2026-07 | 12.1% | 87.6% ± 0.7 | 6.5% ± 0.4 | 14.7% | 83.4% ± 0.2 | 8.0% ± 0.8 | 215 / 2193 repeats |
| 27 | Codex · GPT-5.5 · Lowmodel=gpt-5.5 · reasoning=lowRun date: · official-2026-07 | 10.8% | 83.8% ± 1.2 | 5.7% ± 0.5 | 13.0% | 75.2% ± 3.2 | 7.1% ± 0.6 | 219 / 2193 repeats |
| 28 | Codex · GPT-5.6 Terra · Lowmodel=gpt-5.6-terra · reasoning=lowRun date: · official-2026-07 | 10.6% | 88.5% ± 3.0 | 5.7% ± 0.5 | 13.0% | 83.7% ± 0.3 | 7.0% ± 0.5 | 215 / 2193 repeats |
Source ranks follow the default grounded F1 view. F1 uses the website formula on published mean precision and recall. ± shows reported standard deviation across repeats, not a confidence interval. Completed PR counts can differ between runs. Augmented recall uses a different denominator for each reviewer and is a per-system diagnostic.
How to read these review results
ReviewBench tests complete code-review agents on 219 pull requests. Grounded precision measures valid findings among those matched to the reference set; grounded recall measures how many known valid issues the reviewer finds. We retain every published configuration, source rank, run date, profile label, repeat count, and completion count.
The default source order uses grounded F1, which balances precision and recall. Augmented metrics also credit valid findings outside the reference set. Augmented recall changes its denominator for each reviewer, so it is a per-system diagnostic rather than a direct comparison between agents.
GitHub publishes this benchmark and evaluates its own Copilot Code Review product alongside other reviewers. The judge uses Claude Sonnet 5, and the reference findings can be incomplete. These results describe the listed agent configurations on this task set; they stay outside overall and category scores.
Snapshot
The published ReviewBench snapshot places Copilot Code Review · Balanced first at 40.1%. The third row is 5.0 points behind. The broader top-10 range is 12.8 points, so the table still separates the published systems.
28 reviewer configurations are shown for ReviewBench. The benchmark falls in the Agentic category. ReviewBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ReviewBench
Year
2026
Tasks
219 pull requests across 187 repositories and 19 languages
Format
Grounded precision, recall, and F1; augmented metrics
Each row is a complete reviewer configuration, including its disclosed reasoning settings. The source ranks by grounded F1. We retain source profiles, run dates, repeats, and completed PR counts. GitHub develops the benchmark and Copilot Code Review; Claude Sonnet 5 judges findings. These agent results are display only and do not change overall or category scores.
Freshness and provenance
Version
ReviewBench research preview; source profiles retained per configuration
Refresh cadence
Published leaderboard snapshots
Staleness state
Current
Question availability
Public upstream corpus; only aggregate results mirrored
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ReviewBench measure?
ReviewBench measures how reliably complete code-review agents find useful issues in 219 real pull requests. Grounded metrics compare findings with a reference set; augmented metrics also credit valid new findings. Each result belongs to the listed reviewer configuration, including its model and reasoning settings when disclosed.
How should I compare ReviewBench results?
Start with grounded precision and recall, then check the exact configuration, run date, repeats, and completed pull requests. Grounded F1 balances precision and recall in the default source order. Augmented recall uses a different denominator for each reviewer, so it is a diagnostic rather than a direct comparison.
Does ReviewBench change model rankings?
These results stay outside overall and category scores. ReviewBench evaluates complete review agents, and their setup affects the result. We preserve every published configuration instead of converting one into a base-model score. GitHub publishes the benchmark and its own Copilot results, which is useful context when reading the table.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.