Skip to main content
BenchLM
Data

ReviewBench

We show this table for reference; we do not rank on it.

Data verified 40 confirmed releases in the last 30 daysFollow model changes

Measures how reliably code-review agents identify useful issues in real pull requests while avoiding false positives.

Evaluation results published by GitHub.

Grounded F1 on ReviewBench — 2026-10-07 capture

We mirror the published grounded f1 view for ReviewBench. Copilot Code Review · Balanced leads the public snapshot at 40.1%, followed by Devin AI (37.0%) and Qodo (35.1%). We do not use these results to rank models overall.

28 reviewer configurationsAgenticCurrentDisplay onlyUpdated 2026-10-07 capture

Review agents (28 configurations)
Source rankReviewer / configurationGrounded F1Grounded precisionGrounded recallAugmented F1Augmented precisionAugmented recallCompleted PRs / repeats
1Copilot Code Review · Balancedeffort=mediumRun date: · official-2026-0940.1%87.8% ± 0.526.0% ± 0.649.7%87.4% ± 2.034.7% ± 0.2219 / 2193 repeats
2Devin AIRun date: · official-2026-0937.0%84.0% ± 0.723.8% ± 0.147.3%80.4% ± 0.233.5% ± 0.3219 / 2193 repeats
3QodoRun date: · official-2026-0935.1%85.3% ± 0.822.1% ± 0.944.5%86.8% ± 1.129.9% ± 1.1218-219 / 2193 repeats
4Codex · GPT-5.6 Sol · Maxmodel=gpt-5.6-sol · reasoning=maxRun date: · official-2026-0732.2%86.8% ± 0.319.7% ± 0.943.1%89.1% ± 1.328.4% ± 1.3215 / 2193 repeats
5Codex · GPT-5.6 Sol · Ultramodel=gpt-5.6-sol · reasoning=ultraRun date: · official-2026-0929.6%85.9% ± 1.617.9% ± 0.438.4%87.3% ± 1.224.6% ± 0.1219 / 2193 repeats
6Codex · GPT-5.6 Sol · xHighmodel=gpt-5.6-sol · reasoning=xhighRun date: · official-2026-0729.5%88.8% ± 0.917.7% ± 0.540.4%90.5% ± 0.826.0% ± 0.5215 / 2193 repeats
7Copilot Code Review · Liteeffort=lowRun date: · official-2026-0728.6%83.8% ± 0.917.2% ± 0.437.2%83.7% ± 0.323.9% ± 0.9216-218 / 2193 repeats
8Codex · GPT-5.6 Luna · Maxmodel=gpt-5.6-luna · reasoning=maxRun date: · official-2026-0727.4%87.7% ± 1.516.2% ± 0.436.0%88.5% ± 1.322.6% ± 0.2215 / 2193 repeats
9CubicRun date: · official-2026-0727.3%85.5%16.3%37.4%85.1%24.0%219 / 2191 repeat
10GreptileRun date: · official-2026-0927.2%86.1%16.2%32.8%85.6%20.3%218 / 2191 repeat
11Codex · GPT-5.6 Sol · Highmodel=gpt-5.6-sol · reasoning=highRun date: · official-2026-0726.8%89.0% ± 2.315.8% ± 0.036.5%90.8% ± 0.722.8% ± 1.2215 / 2193 repeats
12Codex · GPT-5.6 Terra · Ultramodel=gpt-5.6-terra · reasoning=ultraRun date: · official-2026-0725.1%90.0% ± 1.914.6% ± 0.833.8%90.1% ± 1.420.8% ± 0.9215 / 2193 repeats
13Codex · GPT-5.6 Terra · Maxmodel=gpt-5.6-terra · reasoning=maxRun date: · official-2026-0725.0%90.2% ± 1.314.5% ± 0.234.1%90.0% ± 1.321.1% ± 0.3215 / 2193 repeats
14Codex · GPT-5.6 Sol · Mediummodel=gpt-5.6-sol · reasoning=mediumRun date: · official-2026-0723.7%88.5% ± 2.013.7% ± 0.430.1%90.5% ± 1.918.0% ± 0.4215 / 2193 repeats
15Codex · GPT-5.4 · xHighmodel=gpt-5.4 · reasoning=xhighRun date: · official-2026-0721.8%88.7% ± 1.512.4% ± 0.230.9%86.5% ± 0.618.8% ± 0.1219 / 2193 repeats
16Codex · GPT-5.4 · Highmodel=gpt-5.4 · reasoning=highRun date: · official-2026-0721.0%88.0% ± 0.411.9% ± 0.428.3%86.7% ± 0.916.9% ± 0.2219 / 2193 repeats
17Codex · GPT-5.6 Terra · xHighmodel=gpt-5.6-terra · reasoning=xhighRun date: · official-2026-0719.8%90.1% ± 1.711.1% ± 0.725.6%89.3% ± 1.314.9% ± 0.9215 / 2193 repeats
18Codex · GPT-5.5 · xHighmodel=gpt-5.5 · reasoning=xhighRun date: · official-2026-0718.8%88.6% ± 1.110.5% ± 0.824.8%87.1% ± 1.214.4% ± 0.2219 / 2193 repeats
19Codex · GPT-5.6 Sol · Lowmodel=gpt-5.6-sol · reasoning=lowRun date: · official-2026-0717.6%90.2% ± 0.99.7% ± 0.421.1%91.4% ± 1.111.9% ± 0.3215 / 2193 repeats
20CursorRun date: · official-2026-0917.3%87.7% ± 0.89.6% ± 0.321.7%89.3% ± 0.812.3% ± 0.5219 / 2193 repeats
21Codex · GPT-5.4 · Mediummodel=gpt-5.4 · reasoning=mediumRun date: · official-2026-0717.1%88.0% ± 0.89.5% ± 0.523.5%85.1% ± 1.413.6% ± 0.3219 / 2193 repeats
22Codex · GPT-5.6 Terra · Highmodel=gpt-5.6-terra · reasoning=highRun date: · official-2026-0716.6%88.0% ± 1.59.2% ± 0.320.3%85.4% ± 1.311.5% ± 0.3215 / 2193 repeats
23Codex · GPT-5.5 · Highmodel=gpt-5.5 · reasoning=highRun date: · official-2026-0716.1%85.7% ± 0.38.9% ± 0.120.3%82.4% ± 1.511.6% ± 0.2219 / 2193 repeats
24Codex · GPT-5.4 · Lowmodel=gpt-5.4 · reasoning=lowRun date: · official-2026-0714.0%88.7% ± 0.57.6% ± 0.319.3%82.4% ± 0.510.9% ± 0.9219 / 2193 repeats
25Codex · GPT-5.5 · Mediummodel=gpt-5.5 · reasoning=mediumRun date: · official-2026-0714.0%83.7% ± 1.47.6% ± 0.316.8%77.8% ± 1.79.4% ± 0.3219 / 2193 repeats
26Codex · GPT-5.6 Terra · Mediummodel=gpt-5.6-terra · reasoning=mediumRun date: · official-2026-0712.1%87.6% ± 0.76.5% ± 0.414.7%83.4% ± 0.28.0% ± 0.8215 / 2193 repeats
27Codex · GPT-5.5 · Lowmodel=gpt-5.5 · reasoning=lowRun date: · official-2026-0710.8%83.8% ± 1.25.7% ± 0.513.0%75.2% ± 3.27.1% ± 0.6219 / 2193 repeats
28Codex · GPT-5.6 Terra · Lowmodel=gpt-5.6-terra · reasoning=lowRun date: · official-2026-0710.6%88.5% ± 3.05.7% ± 0.513.0%83.7% ± 0.37.0% ± 0.5215 / 2193 repeats

Source ranks follow the default grounded F1 view. F1 uses the website formula on published mean precision and recall. ± shows reported standard deviation across repeats, not a confidence interval. Completed PR counts can differ between runs. Augmented recall uses a different denominator for each reviewer and is a per-system diagnostic.

How to read these review results

ReviewBench tests complete code-review agents on 219 pull requests. Grounded precision measures valid findings among those matched to the reference set; grounded recall measures how many known valid issues the reviewer finds. We retain every published configuration, source rank, run date, profile label, repeat count, and completion count.

The default source order uses grounded F1, which balances precision and recall. Augmented metrics also credit valid findings outside the reference set. Augmented recall changes its denominator for each reviewer, so it is a per-system diagnostic rather than a direct comparison between agents.

GitHub publishes this benchmark and evaluates its own Copilot Code Review product alongside other reviewers. The judge uses Claude Sonnet 5, and the reference findings can be incomplete. These results describe the listed agent configurations on this task set; they stay outside overall and category scores.

Snapshot

28 configurations219 pull requestsGrounded and augmented metricsDisplay only

The published ReviewBench snapshot places Copilot Code Review · Balanced first at 40.1%. The third row is 5.0 points behind. The broader top-10 range is 12.8 points, so the table still separates the published systems.

28 reviewer configurations are shown for ReviewBench. The benchmark falls in the Agentic category. ReviewBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ReviewBench

Year

2026

Tasks

219 pull requests across 187 repositories and 19 languages

Format

Grounded precision, recall, and F1; augmented metrics

Each row is a complete reviewer configuration, including its disclosed reasoning settings. The source ranks by grounded F1. We retain source profiles, run dates, repeats, and completed PR counts. GitHub develops the benchmark and Copilot Code Review; Claude Sonnet 5 judges findings. These agent results are display only and do not change overall or category scores.

Freshness and provenance

Version

ReviewBench research preview; source profiles retained per configuration

Refresh cadence

Published leaderboard snapshots

Staleness state

Current

Question availability

Public upstream corpus; only aggregate results mirrored

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ReviewBench measure?

ReviewBench measures how reliably complete code-review agents find useful issues in 219 real pull requests. Grounded metrics compare findings with a reference set; augmented metrics also credit valid new findings. Each result belongs to the listed reviewer configuration, including its model and reasoning settings when disclosed.

How should I compare ReviewBench results?

Start with grounded precision and recall, then check the exact configuration, run date, repeats, and completed pull requests. Grounded F1 balances precision and recall in the default source order. Augmented recall uses a different denominator for each reviewer, so it is a diagnostic rather than a direct comparison.

Does ReviewBench change model rankings?

These results stay outside overall and category scores. ReviewBench evaluates complete review agents, and their setup affects the result. We preserve every published configuration instead of converting one into a base-model score. GitHub publishes the benchmark and its own Copilot results, which is useful context when reading the table.

Last updated: 2026-10-07 capture · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.