LLM Leaderboard History: Arena Rankings Since 2023
claude-opus-4-6-thinking leads the 2026-07 Arena snapshot at 1512 rating points. Across 39 month-end snapshots since May 2023, the leader changed 18 times among 7 providers.
This is a history of measured user preference, not a universal capability curve. Ratings move with the model pool, prompt mix, voters, and Arena's methodology. Use the timeline to inspect leadership changes; use capability benchmarks for claims about coding, reasoning, or task accuracy.
The Numbers
Key stats from 39 monthly Arena preference snapshots.
+437
Leader Rating Change
1075 to 1512
18
Crown Changes
39 months tracked
5mo
Longest Reign
gemini-2.5-pro
37 pts
Current Open Gap
glm-5.1 trails the leader
11.5/mo
Leader Trend
Descriptive, not capability growth
7
Providers at #1
openai leads
Elo Rating Progression: May 2023 to Today
More
Month-end Arena preference leaders. Proprietary models use a solid line; openly licensed models use a dashed line. The series is not adjusted for changes in the model pool or voting population.
Elo Milestones
When the AI frontier crossed each Elo threshold for the first time.
gpt-4-0314
2023-12 · openai
gpt-4-0314
2024-01 · openai
gpt-4-1106-preview
2024-02 · openai
gpt-4-1106-preview
2024-02 · openai
chatgpt-4o-latest
2024-08 · openai
o1-2024-12-17
2024-12 · openai
chatgpt-4o-latest-20250326
2025-03 · openai
gemini-2.5-pro
2025-06 · google
claude-opus-4-6
2026-02 · anthropic
Key Breakthroughs
Slide through 25 significant moments in the AI race. Each dot marks a breakthrough.
UW enters the race
guanaco-33b is the first UW model to take #1
The Open-License Gap
More
Rating-point difference between the leading proprietary and openly licensed models. The latest gap is 37 points. The closest post-2023 snapshot was 0 points in 2025-01.
Crown Change Timeline
18 changesEvery time the #1 model changed hands since May 2023.
2023-06
vicuna-13b→guanaco-33b(UW)
2023-07
guanaco-33b→vicuna-33b(LMSYS)
2023-10
vicuna-33b→wizardlm-70b(microsoft)
2023-12
wizardlm-70b→gpt-4-0314(openai)
2024-02
gpt-4-0314→gpt-4-1106-preview(openai)
2024-03
gpt-4-1106-preview→claude-3-opus-20240229(anthropic)
2024-04
claude-3-opus-20240229→gpt-4-turbo-2024-04-09(openai)
2024-05
gpt-4-turbo-2024-04-09→gpt-4o-2024-05-13(openai)
Provider Dominance
Months at #1 on the Arena leaderboard since May 2023. Who has dominated the AI race?
Category Breakdown: Who Wins at What?
More
Arena tracks Elo by category. Here's who leads in coding, math, creative writing, and more — with Elo gain over time.
English
+290 Elo#1: claude-opus-4-6-thinking (anthropic) · 1525
Hard Prompts
+260 Elo#1: claude-fable-5 (anthropic) · 1534
Coding
+243 Elo#1: claude-fable-5 (anthropic) · 1553
Chinese
+235 Elo#1: claude-opus-4-6 (anthropic) · 1565
Multi-Turn
+217 Elo#1: gpt-5.2-high (openai) · 1525
Math
+200 Elo#1: claude-fable-5 (anthropic) · 1543
Creative Writing
+188 Elo#1: claude-fable-5 (anthropic) · 1508
Instruction Following
+159 Elo#1: claude-fable-5 (anthropic) · 1513
Korean
+135 Elo#1: gpt-5.5 (openai) · 1507
Japanese
+128 Elo#1: claude-fable-5 (anthropic) · 1531
Vision Arena: Multimodal Model Rankings
More
How multimodal (vision) models have improved over time. Currently led by claude-fable-5 at 1334 Elo.
Openly Licensed Leaders Over Time
Every model that led rows Arena classified with a non-proprietary license.
vicuna-13b
LMSYS · 2023-05
guanaco-33b
UW · 2023-06
vicuna-33b
LMSYS · 2023-07
wizardlm-70b
microsoft · 2023-10
mixtral-8x7b-instruct-v0.1
mistral · 2023-12
qwen1.5-72b-chat
alibaba · 2024-02
llama-3-70b-instruct
meta · 2024-04
gemma-2-27b-it
google · 2024-06
llama-3.1-405b-instruct
meta · 2024-07
llama-3.1-405b-instruct-bf16
meta · 2024-09
llama-3.1-nemotron-70b-instruct
nvidia · 2024-10
athene-v2-chat
NexusFlow · 2024-11
deepseek-v3
deepseek · 2024-12
deepseek-r1
deepseek · 2025-01
deepseek-v3-0324
deepseek · 2025-03
deepseek-r1-0528
deepseek · 2025-06
glm-4.5
zai · 2025-08
qwen3-vl-235b-a22b-instruct
alibaba · 2025-09
glm-4.6
zai · 2025-10
kimi-k2.5-thinking
moonshot · 2026-01
qwen3.5-397b-a17b
alibaba · 2026-02
glm-5.1
zai · 2026-04
Monthly Rankings
Full top-20 Arena rankings for any month. Scroll through 39 snapshots from 2023-05 to 2026-07.
| # | Model | Org | License | Elo | Votes |
|---|---|---|---|---|---|
| 1 | claude-opus-4-6-thinking | anthropic | Proprietary | 1512 | 62,772 |
| 2 | claude-opus-4-6 | anthropic | Proprietary | 1507 | 66,604 |
| 3 | claude-fable-5 | anthropic | Proprietary | 1504 | 14,575 |
| 4 | claude-opus-4-7-thinking | anthropic | Proprietary | 1499 | 50,483 |
| 5 | gpt-5.4 | openai | Proprietary | 1499 | 61,783 |
| 6 | gpt-5.4-mini-high | openai | Proprietary | 1496 | 57,600 |
| 7 | gpt-5.4-high | openai | Proprietary | 1494 | 58,808 |
| 8 | gpt-5.2-high | openai | Proprietary | 1493 | 47,421 |
| 9 | claude-opus-4-7 | anthropic | Proprietary | 1492 | 51,584 |
| 10 | muse-spark-1.1 | meta | Proprietary | 1491 | 7,777 |
| 11 | gemini-3.6-flash | Proprietary | 1491 | 4,747 | |
| 12 | gpt-5.5 | openai | Proprietary | 1490 | 47,031 |
| 13 | gemini-3.5-flash-high | Proprietary | 1490 | 10,092 | |
| 14 | gemini-3-pro | Proprietary | 1490 | 40,384 | |
| 15 | gemini-3.1-pro-preview | Proprietary | 1489 | 84,142 | |
| 16 | gpt-5.5-high | openai | Proprietary | 1488 | 45,597 |
| 17 | gpt-5.2 | openai | Proprietary | 1486 | 76,972 |
| 18 | gpt-5.6-sol-xhigh | openai | Proprietary | 1486 | 6,117 |
| 19 | qwen3.7-max-preview | alibaba | Proprietary | 1485 | 3,714 |
| 20 | gpt-5.1 | openai | Proprietary | 1485 | 42,489 |
Frequently Asked Questions
What is the Arena Elo leaderboard?
Arena is a crowdsourced preference test where users vote on anonymous side-by-side model responses. Its ratings summarize relative preference within the models, prompts, voters, and ranking method present in each published snapshot. They are useful evidence, but they do not measure every capability or convert directly into task success rates.
How much have AI models improved since 2023?
The published monthly leader moved from 1075 (vicuna-13b) in May 2023 to 1512 (claude-opus-4-6-thinking) in 2026-07, a difference of 437 rating points across 39 monthly snapshots. This is a historical preference trend, not a controlled estimate of capability growth.
Which company has dominated the AI leaderboard?
openai has held the #1 position for 15 out of 39 months (38%), followed by google (8 months) and anthropic (7 months). The crown has changed hands 18 times since May 2023.
Are openly licensed LLMs catching up to proprietary models?
The highest-rated openly licensed model is glm-5.1 at 1475. It trails claude-opus-4-6-thinking by 37 rating points in the 2026-07 snapshot. The closest post-2023 gap was 0 points in 2025-01.
How does Arena Elo differ from benchmark scores?
Arena ratings summarize blind user preferences across the prompts and models in its pool. Benchmarks instead test defined tasks under a specified scoring protocol. Neither is universally stronger: preference ratings are broad but population-dependent, while benchmarks are easier to interpret by capability but can saturate, leak, or reward a narrow setup.
Which AI model is best for coding?
According to Arena coding Elo, claude-fable-5 (anthropic) currently leads with an Elo of 1553. The coding category has seen 15 crown changes over 23 months.
What models are best for math and reasoning?
The current Arena math leader is claude-fable-5 (anthropic) at 1543 rating points. The category leader changed 15 times across 19 monthly snapshots.
Where does this data come from?
All ratings come from the Arena Leaderboard Dataset on Hugging Face, maintained by Arena Intelligence. We process the text, text_style_control, vision, and webdev subsets, keep the latest published board in each calendar month, deduplicate repeated display-model rows, and recompute the visible ordinal order from published ratings.
Data attribution: All Elo ratings on this page come from the Arena Leaderboard Dataset by Arena Intelligence (lmarena-ai), available on HuggingFace. Data covers text, text_style_control, vision, webdev arena subsets.
We process and visualize this data to provide historical comparisons. We do not generate the underlying Elo ratings. For BenchLM's own benchmark-based rankings, see the main leaderboard or the AI Race timeline.