LiveBench
A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.
Data verified 37 confirmed releases in the last 30 daysSee provider release alertsHow we show LiveBench
We mirror the official LiveBench 2026-06-25 release: 57 model variants across 23 objective tasks in 7 categories. The visible score is the mean of the seven category averages, matching the source leaderboard.
LiveBench refreshes its questions to limit contamination, but each row still carries a specific model version and reasoning-effort setting. We keep the mirror display only and preserve the category and cost fields in the snapshot instead of collapsing them into weighted model scores.
Snapshot
LiveBench overall score on LiveBench — June 25, 2026
We mirror the published livebench overall score view for LiveBench. Claude Fable 5.1 leads the public snapshot at 83.4%, followed by Claude Fable 5 (83.0%) and GPT-6 Astra (82.2%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Fable 5
Anthropic
GPT-6 Astra
OpenAI
57 modelsExternal benchmark mirrorsRefreshingDisplay onlyUpdated June 25, 2026
LiveBench overall score table (57 models)
ScoreThe published LiveBench snapshot places Claude Fable 5.1 first at 83.4%. The third row is 1.3 points behind. The broader top-10 range is 4.2 points, so many of the published results sit in a relatively narrow band.
57 models have been evaluated on LiveBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. LiveBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About LiveBench
Year
2024
Tasks
23 objective tasks across 7 categories
Format
Mean of category averages
Difficulty
Broad frontier-model evaluation
LiveBench rotates questions to limit contamination and scores answers against objective ground truth. We mirror the current release-specific table, including category scores and cost fields. The overall score is the mean of category averages. Model versions and reasoning-effort settings remain separate, and the mirror does not feed weighted rankings.
BenchLM freshness & provenance
Version
LiveBench 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does LiveBench measure?
A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.
Which model leads the published LiveBench snapshot?
Claude Fable 5.1 currently leads the published LiveBench snapshot with 83.4% livebench overall score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on LiveBench?
The June 25, 2026 snapshot contains 57 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.