Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

ReactBench v1 (ReactBench)

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

How we show ReactBench

We mirror the official ReactBench table from Million: 54 effort variants across 39 React tasks. The source score is a weighted rubric aggregate, and solutions that miss a blocking criterion receive zero.

ReactBench tests production React work that ordinary behavior checks can miss, including performance, accessibility, and code quality. The rows also include an agent setup and reasoning-effort choice, so we preserve the published variants and costs without adding the scores to weighted coding rankings.

Snapshot

54 model variants20 base models39 tasksCost per rollout preservedDisplay only

ReactBench score (Pass@1) on ReactBench — September 15, 2026 snapshot

We mirror the published reactbench score (pass@1) view for ReactBench. GPT-5.6 Sol leads the public snapshot at 46.7%, followed by Fable 5.1 · Max (45.6%) and Fable 5.1 · XHigh (44.1%). We do not use these results to rank models overall.

54 modelsCodingCurrentDisplay onlyUpdated September 15, 2026 snapshot

ReactBench score (Pass@1) table (54 models)

Score
1
GPT-5.6 SolOpenAI · Closed
46.7%
2
45.6%
3
44.1%
4
Claude Fable 5Anthropic · Closed
43.6%
5
Claude Fable 5Anthropic · Closed
42.1%
6
Claude Opus 5Anthropic · Closed
42.1%
7
Claude Opus 5Anthropic · Closed
42.1%
8
GPT-5.6 SolOpenAI · Closed
42.1%
9
40.5%
10
GPT-5.6 TerraOpenAI · Closed
40.5%
11
Claude Fable 5Anthropic · Closed
39.0%
12
GPT-5.6 SolOpenAI · Closed
37.4%
13
GPT-5.6 TerraOpenAI · Closed
37.4%
14
GPT-5.6 TerraOpenAI · Closed
36.9%
15
Claude Opus 5Anthropic · Closed
36.4%
16
Claude Opus 5Anthropic · Closed
34.9%
17
GPT-5.6 LunaOpenAI · Closed
33.8%
18
32.3%
19
Grok 4.6xAI · Closed
32.3%
20
GPT-5.6 SolOpenAI · Closed
32.3%
21
GPT-5.6 SolOpenAI · Closed
32.3%
22
Claude Fable 5Anthropic · Closed
31.8%
23
GLM-5.3-FlashZ.AI · Open weight
31.3%
24
Claude Opus 4.8Anthropic · Closed
30.8%
25
Claude Opus 4.8Anthropic · Closed
30.8%
26
Grok 4.6xAI · Closed
30.8%
27
Gemini 3.8 FlashGoogle · Closed
30.8%
28
GPT-5.6 LunaOpenAI · Closed
30.8%
29
Kimi K3Moonshot AI · Closed
30.8%
30
GPT-5.6 LunaOpenAI · Closed
30.3%
31
Claude Sonnet 5Anthropic · Closed
29.7%
32
DeepSeek V4 Pro 0813DeepSeek · Closed
29.7%
33
GPT-5.6 TerraOpenAI · Closed
29.7%
34
Claude Sonnet 5Anthropic · Closed
29.2%
35
28.7%
36
Claude Sonnet 5Anthropic · Closed
28.2%
37
GPT-5.6 TerraOpenAI · Closed
28.2%
38
GLM-5.3Z.AI · Open weight
27.7%
39
27.2%
40
Claude Opus 4.8Anthropic · Closed
26.7%
41
Claude Opus 5Anthropic · Closed
26.2%
42
Claude Opus 4.8Anthropic · Closed
25.6%
43
Grok 4.6xAI · Closed
25.6%
44
Grok 4.6xAI · Closed
25.6%
45
GLM-5.2Z.AI · Open weight
24.1%
46
Claude Sonnet 5Anthropic · Closed
22.6%
47
Claude Opus 4.8Anthropic · Closed
21.0%
48
GPT-5.6 LunaOpenAI · Closed
19.5%
49
Muse Spark 1.2Meta · Closed
19.5%
50
DeepSeek V4 Flash 0731DeepSeek · Closed
18.5%
51
Claude Sonnet 5Anthropic · Closed
16.9%
52
Composer 2.5Cursor · Closed
13.3%
53
GPT-5.6 LunaOpenAI · Closed
8.7%
54
Inkling-SmallThinking Machines Lab · Open weight
6.7%

The published ReactBench snapshot places GPT-5.6 Sol first at 46.7%. The third row is 2.6 points behind. The broader top-10 range is 6.2 points, so many of the published results sit in a relatively narrow band.

54 models have been evaluated on ReactBench. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. ReactBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ReactBench

Year

2026

Tasks

51 production React tasks

Format

Pass@1 weighted rubric score

Difficulty

Production frontend engineering

ReactBench uses 51 tasks drawn from real React repositories. Its weighted rubric score assigns zero to solutions that miss a blocking criterion. We mirror each published reasoning-effort variant and its average rollout cost, but keep the results out of weighted coding scores because the agent setup and effort level are part of the row.

BenchLM freshness & provenance

Version

ReactBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does ReactBench measure?

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

Which model leads the published ReactBench snapshot?

GPT-5.6 Sol currently leads the published ReactBench snapshot with 46.7% reactbench score (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on ReactBench?

The September 15, 2026 snapshot snapshot contains 54 AI models.

Last updated: September 15, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.