Skip to main content
BenchLM

BridgeBench

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

BridgeBench rates 25 models across nine software-building axes, including reasoning, front end, back end, security, speed, and cost. Its overall rating is the equal-weight mean of those axes.

Overall editorial rating on BridgeBench — September 23, 2026

We mirror the published overall editorial rating view for BridgeBench. Claude Opus 5.5 leads the public snapshot at 760, followed by GPT-6 Astra (734) and Claude Fable 5.1 (711). We do not use these results to rank models overall.

25 modelsCodingCurrentDisplay onlyUpdated September 23, 2026

Overall editorial rating table (25 models)

Score
1
Claude Opus 5.5Anthropic · Closed
760
2
GPT-6 AstraOpenAI · Closed
734
3
Claude Fable 5.1Anthropic · Closed
711
4
GPT-6 SolOpenAI · Closed
643
5
Claude Fable 5Anthropic · Closed
628
6
MiMo-V2.6-ProXiaomi · Open weight
617
7
Muse Spark 1.3Meta · Closed
606
8
GPT-5.6 SolOpenAI · Closed
600
9
Gemini 3.8 FlashGoogle · Closed
584
10
Grok 4.7xAI · Closed
583
11
Step 5 PreviewStepFun · Closed
572
12
DeepSeek V4.1 FlashDeepSeek · Open weight
564
13
Grok 4.6xAI · Closed
536
14
Qwen3.8 MaxAlibaba · Open weight
532
15
Qwen3.8-Flash-NextAlibaba · Open weight
522
16
GPT-6 LunaOpenAI · Closed
513
17
Claude Opus 5Anthropic · Closed
512
18
GPT-5.6 TerraOpenAI · Closed
511
19
Kimi K3Moonshot AI · Closed
501
20
Claude Sonnet 5Anthropic · Closed
501
21
GPT-5.6 LunaOpenAI · Closed
476
22
468
23
461
24
460
25
450

How to read this leaderboard

Compare the source's software-building assessments as one editorial view. The overall mean includes speed and cost alongside capability axes, so a higher rating is not a pure capability claim.

Operator receipt: 25 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5.5 at 760.

Honest limit: BridgeBench does not publish its proprietary prompts, task sets, or detailed rubrics. Confidence varies by model and axis, and the source says small rating gaps do not prove a meaningful difference on every task. These ratings do not enter BenchLM overall or category rankings.

How we show BridgeBench

We captured 25 overall ratings from the public BridgeBench leaderboard on September 23, 2026. BridgeBench averages nine axes, including reasoning, front end, back end, security, speed, and cost, with equal weight.

These are BridgeBench editorial assessments expressed as Elo-style points. They are not head-to-head match records or pass rates from a shared test suite. The source keeps its prompts, task sets, and detailed rubrics private.

We show this table for reference only. Its ratings do not enter our overall or category rankings.

Snapshot

25 model ratings9 equally weighted axesEditorial ratingsDisplay only

The published BridgeBench snapshot places Claude Opus 5.5 first at 760. The third row is 49 score units behind. The broader top-10 range is 177 score units, so the table still separates the published systems.

25 models have been evaluated on BridgeBench. The benchmark falls in the Coding category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. BridgeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About BridgeBench

Year

2026

Tasks

No public task count

Format

Editorial Elo-style rating

Difficulty

Software-building assessments

The published ratings combine editorial judgment with available performance evidence. BridgeBench converts comparative assessments to Elo-style points; they are not head-to-head match results or pass rates from one shared test suite. We mirror the 25 published overall ratings as display-only context.

Freshness and provenance

Version

September 2026 editorial ratings

Refresh cadence

Manual snapshot

Staleness state

Current

Question availability

Prompts, task sets, and detailed rubrics private

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does BridgeBench measure?

BridgeBench rates 25 models across nine software-building axes, including reasoning, front end, back end, security, speed, and cost. Its overall rating is the equal-weight mean of those axes.

Which model leads the published BridgeBench snapshot?

Claude Opus 5.5 currently leads the published BridgeBench snapshot with 760 overall editorial rating. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on BridgeBench?

The September 23, 2026 snapshot contains 25 AI models.

Last updated: September 23, 2026 · mirrored from published editorial ratings

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.