# ReviewBench

> Measures how reliably code-review agents identify useful issues in real pull requests while avoiding false positives.

Canonical page: https://benchlm.ai/benchmarks/reviewbench

- Category: [Agentic](/agentic)
- Last updated: 2026-10-07 capture

Evaluation results published by [GitHub](https://review-bench.ai/). undefined

ReviewBench tests complete code-review agents on 219 pull requests. Grounded precision measures valid findings among those matched to the reference set; grounded recall measures how many known valid issues the reviewer finds. We retain every published configuration, source rank, run date, profile label, repeat count, and completion count.

The default source order uses grounded F1, which balances precision and recall. Augmented metrics also credit valid findings outside the reference set. Augmented recall changes its denominator for each reviewer, so it is a per-system diagnostic rather than a direct comparison between agents.

GitHub publishes this benchmark and evaluates its own Copilot Code Review product alongside other reviewers. The judge uses Claude Sonnet 5, and the reference findings can be incomplete. These results describe the listed agent configurations on this task set; they stay outside overall and category scores.

- [Official ReviewBench leaderboard](https://review-bench.ai/)
- [Published aggregate results](https://review-bench.ai/api/leaderboard)
- [Scoring and limitations](https://review-bench.ai/methodology)
- [Benchmark repository](https://github.com/review-bench/ReviewBench)

## About ReviewBench

- Year: 2026
- Tasks: 219 pull requests across 187 repositories and 19 languages
- Format: Grounded precision, recall, and F1; augmented metrics
- Paper: [ReviewBench: An open benchmark for AI code review](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/)

Each row is a complete reviewer configuration, including its disclosed reasoning settings. The source ranks by grounded F1. We retain source profiles, run dates, repeats, and completed PR counts. GitHub develops the benchmark and Copilot Code Review; Claude Sonnet 5 judges findings. These agent results are display only and do not change overall or category scores.

ReviewBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (28 reviewer configurations)

| Source rank | Reviewer / configuration | Grounded F1 | Grounded precision | Grounded recall | Augmented F1 | Augmented precision | Augmented recall | Run date | Source profile | Completed PRs per repeat |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | [Copilot Code Review · Balanced](https://review-bench.ai/); effort=medium | 40.1% | 87.8% ± 0.5 | 26.0% ± 0.6 | 49.7% | 87.4% ± 2.0 | 34.7% ± 0.2 | 2026-10-01 | official-2026-09 | 219/219, 219/219, 219/219 |
| 2 | [Devin AI](https://review-bench.ai/); Not disclosed | 37.0% | 84.0% ± 0.7 | 23.8% ± 0.1 | 47.3% | 80.4% ± 0.2 | 33.5% ± 0.3 | 2026-09-28 | official-2026-09 | 219/219, 219/219, 219/219 |
| 3 | [Qodo](https://review-bench.ai/); Not disclosed | 35.1% | 85.3% ± 0.8 | 22.1% ± 0.9 | 44.5% | 86.8% ± 1.1 | 29.9% ± 1.1 | 2026-09-28 | official-2026-09 | 219/219, 218/219, 219/219 |
| 4 | [Codex · GPT-5.6 Sol · Max](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=max | 32.2% | 86.8% ± 0.3 | 19.7% ± 0.9 | 43.1% | 89.1% ± 1.3 | 28.4% ± 1.3 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 5 | [Codex · GPT-5.6 Sol · Ultra](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=ultra | 29.6% | 85.9% ± 1.6 | 17.9% ± 0.4 | 38.4% | 87.3% ± 1.2 | 24.6% ± 0.1 | 2026-10-05 | official-2026-09 | 219/219, 219/219, 219/219 |
| 6 | [Codex · GPT-5.6 Sol · xHigh](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=xhigh | 29.5% | 88.8% ± 0.9 | 17.7% ± 0.5 | 40.4% | 90.5% ± 0.8 | 26.0% ± 0.5 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 7 | [Copilot Code Review · Lite](https://review-bench.ai/); effort=low | 28.6% | 83.8% ± 0.9 | 17.2% ± 0.4 | 37.2% | 83.7% ± 0.3 | 23.9% ± 0.9 | 2026-07-20 | official-2026-07 | 218/219, 216/219, 217/219 |
| 8 | [Codex · GPT-5.6 Luna · Max](https://review-bench.ai/); model=gpt-5.6-luna · reasoning=max | 27.4% | 87.7% ± 1.5 | 16.2% ± 0.4 | 36.0% | 88.5% ± 1.3 | 22.6% ± 0.2 | 2026-07-10 | official-2026-07 | 215/219, 215/219, 215/219 |
| 9 | [Cubic](https://review-bench.ai/); Not disclosed | 27.3% | 85.5% | 16.3% | 37.4% | 85.1% | 24.0% | 2026-06-29 | official-2026-07 | 219/219 |
| 10 | [Greptile](https://review-bench.ai/); Not disclosed | 27.2% | 86.1% | 16.2% | 32.8% | 85.6% | 20.3% | 2026-06-16 | official-2026-09 | 218/219 |
| 11 | [Codex · GPT-5.6 Sol · High](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=high | 26.8% | 89.0% ± 2.3 | 15.8% ± 0.0 | 36.5% | 90.8% ± 0.7 | 22.8% ± 1.2 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 12 | [Codex · GPT-5.6 Terra · Ultra](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=ultra | 25.1% | 90.0% ± 1.9 | 14.6% ± 0.8 | 33.8% | 90.1% ± 1.4 | 20.8% ± 0.9 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 13 | [Codex · GPT-5.6 Terra · Max](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=max | 25.0% | 90.2% ± 1.3 | 14.5% ± 0.2 | 34.1% | 90.0% ± 1.3 | 21.1% ± 0.3 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 14 | [Codex · GPT-5.6 Sol · Medium](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=medium | 23.7% | 88.5% ± 2.0 | 13.7% ± 0.4 | 30.1% | 90.5% ± 1.9 | 18.0% ± 0.4 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 15 | [Codex · GPT-5.4 · xHigh](https://review-bench.ai/); model=gpt-5.4 · reasoning=xhigh | 21.8% | 88.7% ± 1.5 | 12.4% ± 0.2 | 30.9% | 86.5% ± 0.6 | 18.8% ± 0.1 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 16 | [Codex · GPT-5.4 · High](https://review-bench.ai/); model=gpt-5.4 · reasoning=high | 21.0% | 88.0% ± 0.4 | 11.9% ± 0.4 | 28.3% | 86.7% ± 0.9 | 16.9% ± 0.2 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 17 | [Codex · GPT-5.6 Terra · xHigh](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=xhigh | 19.8% | 90.1% ± 1.7 | 11.1% ± 0.7 | 25.6% | 89.3% ± 1.3 | 14.9% ± 0.9 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 18 | [Codex · GPT-5.5 · xHigh](https://review-bench.ai/); model=gpt-5.5 · reasoning=xhigh | 18.8% | 88.6% ± 1.1 | 10.5% ± 0.8 | 24.8% | 87.1% ± 1.2 | 14.4% ± 0.2 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 19 | [Codex · GPT-5.6 Sol · Low](https://review-bench.ai/); model=gpt-5.6-sol · reasoning=low | 17.6% | 90.2% ± 0.9 | 9.7% ± 0.4 | 21.1% | 91.4% ± 1.1 | 11.9% ± 0.3 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 20 | [Cursor](https://review-bench.ai/); Not disclosed | 17.3% | 87.7% ± 0.8 | 9.6% ± 0.3 | 21.7% | 89.3% ± 0.8 | 12.3% ± 0.5 | 2026-09-27 | official-2026-09 | 219/219, 219/219, 219/219 |
| 21 | [Codex · GPT-5.4 · Medium](https://review-bench.ai/); model=gpt-5.4 · reasoning=medium | 17.1% | 88.0% ± 0.8 | 9.5% ± 0.5 | 23.5% | 85.1% ± 1.4 | 13.6% ± 0.3 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 22 | [Codex · GPT-5.6 Terra · High](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=high | 16.6% | 88.0% ± 1.5 | 9.2% ± 0.3 | 20.3% | 85.4% ± 1.3 | 11.5% ± 0.3 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 23 | [Codex · GPT-5.5 · High](https://review-bench.ai/); model=gpt-5.5 · reasoning=high | 16.1% | 85.7% ± 0.3 | 8.9% ± 0.1 | 20.3% | 82.4% ± 1.5 | 11.6% ± 0.2 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 24 | [Codex · GPT-5.4 · Low](https://review-bench.ai/); model=gpt-5.4 · reasoning=low | 14.0% | 88.7% ± 0.5 | 7.6% ± 0.3 | 19.3% | 82.4% ± 0.5 | 10.9% ± 0.9 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 25 | [Codex · GPT-5.5 · Medium](https://review-bench.ai/); model=gpt-5.5 · reasoning=medium | 14.0% | 83.7% ± 1.4 | 7.6% ± 0.3 | 16.8% | 77.8% ± 1.7 | 9.4% ± 0.3 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 26 | [Codex · GPT-5.6 Terra · Medium](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=medium | 12.1% | 87.6% ± 0.7 | 6.5% ± 0.4 | 14.7% | 83.4% ± 0.2 | 8.0% ± 0.8 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |
| 27 | [Codex · GPT-5.5 · Low](https://review-bench.ai/); model=gpt-5.5 · reasoning=low | 10.8% | 83.8% ± 1.2 | 5.7% ± 0.5 | 13.0% | 75.2% ± 3.2 | 7.1% ± 0.6 | 2026-06-29 | official-2026-07 | 219/219, 219/219, 219/219 |
| 28 | [Codex · GPT-5.6 Terra · Low](https://review-bench.ai/); model=gpt-5.6-terra · reasoning=low | 10.6% | 88.5% ± 3.0 | 5.7% ± 0.5 | 13.0% | 83.7% ± 0.3 | 7.0% ± 0.5 | 2026-07-19 | official-2026-07 | 215/219, 215/219, 215/219 |

Source ranks follow the default grounded F1 view. F1 uses the website formula on published mean precision and recall. ± shows standard deviation across repeats, not a confidence interval. Augmented recall uses a reviewer-specific denominator. These agent results stay outside overall and category scores.

## FAQ

### What does ReviewBench measure?

ReviewBench measures how reliably complete code-review agents find useful issues in 219 real pull requests. Grounded metrics compare findings with a reference set; augmented metrics also credit valid new findings. Each result belongs to the listed reviewer configuration, including its model and reasoning settings when disclosed.

### How should I compare ReviewBench results?

Start with grounded precision and recall, then check the exact configuration, run date, repeats, and completed pull requests. Grounded F1 balances precision and recall in the default source order. Augmented recall uses a different denominator for each reviewer, so it is a diagnostic rather than a direct comparison.

### Does ReviewBench change model rankings?

These results stay outside overall and category scores. ReviewBench evaluates complete review agents, and their setup affects the result. We preserve every published configuration instead of converting one into a base-model score. GitHub publishes the benchmark and its own Copilot results, which is useful context when reading the table.
