# BridgeBench

> BridgeBench rates 25 models across nine software-building axes, including reasoning, front end, back end, security, speed, and cost. Its overall rating is the equal-weight mean of those axes.

Canonical page: https://benchlm.ai/benchmarks/bridgebench

- Category: [Coding](/coding)
- Last updated: September 23, 2026

## About BridgeBench

- Year: 2026
- Tasks: No public task count
- Format: Editorial Elo-style rating
- Difficulty: Software-building assessments

The published ratings combine editorial judgment with available performance evidence. BridgeBench converts comparative assessments to Elo-style points; they are not head-to-head match results or pass rates from one shared test suite. We mirror the 25 published overall ratings as display-only context.

BridgeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 760 |
| 2 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 734 |
| 3 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 711 |
| 4 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 643 |
| 5 | [Claude Fable 5](/models/claude-fable) | Anthropic | 628 |
| 6 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | Xiaomi | 617 |
| 7 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 606 |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 600 |
| 9 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 584 |
| 10 | [Grok 4.7](/models/grok-4-7) | xAI | 583 |
| 11 | [Step 5 Preview](/models/step-5-preview) | StepFun | 572 |
| 12 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 564 |
| 13 | [Grok 4.6](/models/grok-4-6) | xAI | 536 |
| 14 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 532 |
| 15 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 522 |
| 16 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | 513 |
| 17 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 512 |
| 18 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 511 |
| 19 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 501 |
| 20 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 501 |
| 21 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 476 |
| 22 | [Qwen3.8 2.4T A95B](https://www.bridgebench.ai/leaderboard) | Alibaba | 468 |
| 23 | [GLM 5.3](https://www.bridgebench.ai/leaderboard) | Z.ai | 461 |
| 24 | [DeepSeek V4 Pro](https://www.bridgebench.ai/leaderboard) | DeepSeek | 460 |
| 25 | [GLM 5.3 Flash](https://www.bridgebench.ai/leaderboard) | Z.ai | 450 |

## FAQ

### What does BridgeBench measure?

BridgeBench rates 25 models across nine software-building axes, including reasoning, front end, back end, security, speed, and cost. Its overall rating is the equal-weight mean of those axes.

### Which model leads the published BridgeBench snapshot?

Claude Opus 5.5 currently leads the published BridgeBench snapshot with a score of 760.

### How many models are evaluated on BridgeBench?

The September 23, 2026 contains 25 AI models.

### Does BridgeBench affect BenchLM's overall score?

Not directly. BridgeBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
