# LiveBench

> A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

Canonical page: https://benchlm.ai/benchmarks/livebench

- Category: [external](/external)
- Last updated: June 25, 2026

## About LiveBench

- Year: 2024
- Tasks: 23 objective tasks across 7 categories
- Format: Mean of category averages
- Difficulty: Broad frontier-model evaluation
- Paper: [LiveBench: A Challenging, Contamination-Free LLM Benchmark](https://livebench.ai/)

LiveBench rotates questions to limit contamination and scores answers against objective ground truth. We mirror the current release-specific table, including category scores and cost fields. The overall score is the mean of category averages. Model versions and reasoning-effort settings remain separate, and the mirror does not feed weighted rankings.

LiveBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (57 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 83.4% |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 83.0% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 82.2% |
| 4 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 81.6% |
| 5 | [Deepseek V4.1 Flash Max (max effort)](https://livebench.ai/) | DeepSeek | 81.1% |
| 6 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 81.1% |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 80.2% |
| 8 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 80.1% |
| 9 | [Smaug Agentic](https://livebench.ai/) | Unknown | 79.5% |
| 10 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 79.2% |
| 11 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 78.8% |
| 12 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 78.5% |
| 13 | [Grok 4.6](/models/grok-4-6) | xAI | 78.0% |
| 14 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 78.0% |
| 15 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 78.0% |
| 16 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 77.9% |
| 17 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 77.4% |
| 18 | [Smaug Flash](https://livebench.ai/) | Unknown | 77.4% |
| 19 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 77.0% |
| 20 | [Smaug Mini](https://livebench.ai/) | Unknown | 76.9% |
| 21 | [Deepseek V4 Flash Vision Exp](https://livebench.ai/) | DeepSeek | 76.8% |
| 22 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 76.5% |
| 23 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 76.2% |
| 24 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 76.2% |
| 25 | [GLM-5.3](/models/glm-5-3) | Z.AI | 76.1% |
| 26 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 76.0% |
| 27 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 75.8% |
| 28 | [Grok 4.5](/models/grok-4-5) | xAI | 75.8% |
| 29 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 75.3% |
| 30 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 75.3% |
| 31 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 74.6% |
| 32 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 74.6% |
| 33 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 74.5% |
| 34 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 74.2% |
| 35 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 74.0% |
| 36 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 73.6% |
| 37 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 73.6% |
| 38 | [GLM-5.2](/models/glm-5-2) | Z.AI | 73.2% |
| 39 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 73.1% |
| 40 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 73.0% |
| 41 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 72.6% |
| 42 | [Inkling](/models/inkling) | Thinking Machines Lab | 71.9% |
| 43 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 71.6% |
| 44 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 71.6% |
| 45 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 70.5% |
| 46 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 69.6% |
| 47 | [Ox Alpha Max (max effort)](https://livebench.ai/) | Unknown | 69.2% |
| 48 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 68.9% |
| 49 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | Moonshot AI | 68.4% |
| 50 | [Grok Build 0.1](/models/grok-build-0-1) | xAI | 67.8% |
| 51 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 67.4% |
| 52 | [MiniMax M3](/models/minimax-m3) | MiniMax | 67.3% |
| 53 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 66.4% |
| 54 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 65.5% |
| 55 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 64.0% |
| 56 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 63.9% |
| 57 | [Grok 4.3](/models/grok-4-3) | xAI | 62.2% |

## FAQ

### What does LiveBench measure?

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

### Which model leads the published LiveBench snapshot?

Claude Fable 5.1 currently leads the published LiveBench snapshot with a score of 83.4%.

### How many models are evaluated on LiveBench?

The June 25, 2026 contains 57 AI models.

### Does LiveBench affect BenchLM's overall score?

Not directly. LiveBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
