Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

LiveBench

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

Data verified 37 confirmed releases in the last 30 daysSee provider release alerts

How we show LiveBench

We mirror the official LiveBench 2026-06-25 release: 57 model variants across 23 objective tasks in 7 categories. The visible score is the mean of the seven category averages, matching the source leaderboard.

LiveBench refreshes its questions to limit contamination, but each row still carries a specific model version and reasoning-effort setting. We keep the mirror display only and preserve the category and cost fields in the snapshot instead of collapsing them into weighted model scores.

Snapshot

57 model variants23 tasks7 categoriesRelease 2026-06-25Display only

LiveBench overall score on LiveBench — June 25, 2026

We mirror the published livebench overall score view for LiveBench. Claude Fable 5.1 leads the public snapshot at 83.4%, followed by Claude Fable 5 (83.0%) and GPT-6 Astra (82.2%). We do not use these results to rank models overall.

57 modelsExternal benchmark mirrorsRefreshingDisplay onlyUpdated June 25, 2026

LiveBench overall score table (57 models)

Score
1
Claude Fable 5.1Anthropic · Closed
83.4%
2
Claude Fable 5Anthropic · Closed
83.0%
3
GPT-6 AstraOpenAI · Closed
82.2%
4
Muse Spark 1.3Meta · Closed
81.6%
6
GPT-5.6 SolOpenAI · Closed
81.1%
7
GPT-5.5OpenAI · Closed
80.2%
8
Claude Opus 5Anthropic · Closed
80.1%
9
79.5%
10
Kimi K3Moonshot AI · Closed
79.2%
11
Gemini 3.7 FlashGoogle · Closed
78.8%
12
Qwen3.8 MaxAlibaba · Open weight
78.5%
13
Grok 4.6xAI · Closed
78.0%
14
GPT-5.4OpenAI · Closed
78.0%
15
Muse Spark 1.2Meta · Closed
78.0%
16
GPT-5.6 TerraOpenAI · Closed
77.9%
17
DeepSeek V4 Pro 0813DeepSeek · Closed
77.4%
18
77.4%
19
Gemini 3.1 ProGoogle · Closed
77.0%
20
Smaug MiniUnknown
76.9%
22
Claude Opus 4.7Anthropic · Closed
76.5%
23
Claude Opus 4.8Anthropic · Closed
76.2%
24
Qwen3.8-Flash-NextAlibaba · Open weight
76.2%
25
GLM-5.3Z.AI · Open weight
76.1%
26
Claude Sonnet 5Anthropic · Closed
76.0%
27
Gemini 3.8 FlashGoogle · Closed
75.8%
28
Grok 4.5xAI · Closed
75.8%
29
Muse Spark 1.1Meta · Closed
75.3%
30
Qwen3.8-27BAlibaba · Open weight
75.3%
31
Gemini 3.5 FlashGoogle · Closed
74.6%
32
GPT-5.2OpenAI · Closed
74.6%
33
Claude Opus 4.6Anthropic · Closed
74.5%
34
DeepSeek V4 Flash 0731DeepSeek · Closed
74.2%
35
GPT-5.2-CodexOpenAI · Closed
74.0%
36
Gemini 3.6 FlashGoogle · Closed
73.6%
37
GPT-5.6 LunaOpenAI · Closed
73.6%
38
GLM-5.2Z.AI · Open weight
73.2%
39
Qwen3.7 MaxAlibaba · Closed
73.1%
40
Claude Sonnet 4.6Anthropic · Closed
73.0%
41
Claude Opus 4.5Anthropic · Closed
72.6%
42
InklingThinking Machines Lab · Open weight
71.9%
43
GLM-5.3-FlashZ.AI · Open weight
71.6%
44
DeepSeek V4 Pro 0813DeepSeek · Closed
71.6%
45
Kimi K2.6Moonshot AI · Open weight
70.5%
46
GPT-5.4 nanoOpenAI · Closed
69.6%
48
Qwen3.6 PlusAlibaba · Closed
68.9%
49
Kimi K2.7 CodeMoonshot AI · Open weight
68.4%
50
Grok Build 0.1xAI · Closed
67.8%
51
Nemotron 3 UltraNVIDIA · Open weight
67.4%
52
MiniMax M3MiniMax · Open weight
67.3%
53
GPT-5.4 miniOpenAI · Closed
66.4%
54
DeepSeek V4 Flash 0731DeepSeek · Closed
65.5%
55
Qwen3.6-27BAlibaba · Open weight
64.0%
56
Gemini 3.5 Flash-LiteGoogle · Closed
63.9%
57
Grok 4.3xAI · Closed
62.2%

The published LiveBench snapshot places Claude Fable 5.1 first at 83.4%. The third row is 1.3 points behind. The broader top-10 range is 4.2 points, so many of the published results sit in a relatively narrow band.

57 models have been evaluated on LiveBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. LiveBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LiveBench

Year

2024

Tasks

23 objective tasks across 7 categories

Format

Mean of category averages

Difficulty

Broad frontier-model evaluation

LiveBench rotates questions to limit contamination and scores answers against objective ground truth. We mirror the current release-specific table, including category scores and cost fields. The overall score is the mean of category averages. Model versions and reasoning-effort settings remain separate, and the mirror does not feed weighted rankings.

BenchLM freshness & provenance

Version

LiveBench 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LiveBench measure?

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

Which model leads the published LiveBench snapshot?

Claude Fable 5.1 currently leads the published LiveBench snapshot with 83.4% livebench overall score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LiveBench?

The June 25, 2026 snapshot contains 57 AI models.

Last updated: June 25, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.