Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

LLM Leaderboard History: Arena Rankings Since 2023

claude-opus-4-6-thinking leads the 2026-07 Arena snapshot at 1512 rating points. Across 39 month-end snapshots since May 2023, the leader changed 18 times among 7 providers.

This is a history of measured user preference, not a universal capability curve. Ratings move with the model pool, prompt mix, voters, and Arena's methodology. Use the timeline to inspect leadership changes; use capability benchmarks for claims about coding, reasoning, or task accuracy.

Data: Arena Leaderboard Dataset |Updated: 2026-07-27|4 arena subsets, 14 categories

The Numbers

Key stats from 39 monthly Arena preference snapshots.

+437

Leader Rating Change

1075 to 1512

18

Crown Changes

39 months tracked

5mo

Longest Reign

gemini-2.5-pro

37 pts

Current Open Gap

glm-5.1 trails the leader

11.5/mo

Leader Trend

Descriptive, not capability growth

7

Providers at #1

openai leads

Elo Rating Progression: May 2023 to Today

More

Month-end Arena preference leaders. Proprietary models use a solid line; openly licensed models use a dashed line. The series is not adjusted for changes in the model pool or voting population.

Proprietary #1 Openly licensed #1

Elo Milestones

When the AI frontier crossed each Elo threshold for the first time.

1100

gpt-4-0314

2023-12 · openai

1150

gpt-4-0314

2024-01 · openai

1200

gpt-4-1106-preview

2024-02 · openai

1250

gpt-4-1106-preview

2024-02 · openai

1300

chatgpt-4o-latest

2024-08 · openai

1350

o1-2024-12-17

2024-12 · openai

1400

chatgpt-4o-latest-20250326

2025-03 · openai

1450

gemini-2.5-pro

2025-06 · google

1500

claude-opus-4-6

2026-02 · anthropic

Key Breakthroughs

Slide through 25 significant moments in the AI race. Each dot marks a breakthrough.

2023-06New Challenger

UW enters the race

guanaco-33b is the first UW model to take #1

2023202420252026
1 / 25

The Open-License Gap

More

Rating-point difference between the leading proprietary and openly licensed models. The latest gap is 37 points. The closest post-2023 snapshot was 0 points in 2025-01.

Crown Change Timeline

18 changes

Every time the #1 model changed hands since May 2023.

2023-06

vicuna-13bguanaco-33b(UW)

2023-07

guanaco-33bvicuna-33b(LMSYS)

2023-10

vicuna-33bwizardlm-70b(microsoft)

2023-12

wizardlm-70bgpt-4-0314(openai)

2024-02

gpt-4-0314gpt-4-1106-preview(openai)

2024-03

gpt-4-1106-previewclaude-3-opus-20240229(anthropic)

2024-04

claude-3-opus-20240229gpt-4-turbo-2024-04-09(openai)

2024-05

gpt-4-turbo-2024-04-09gpt-4o-2024-05-13(openai)

Provider Dominance

Months at #1 on the Arena leaderboard since May 2023. Who has dominated the AI race?

openai15 months (38%)
google8 months (21%)
anthropic7 months (18%)
LMSYS4 months (10%)
microsoft2 months (5%)
deepseek2 months (5%)
UW1 month (3%)

Category Breakdown: Who Wins at What?

More

Arena tracks Elo by category. Here's who leads in coding, math, creative writing, and more — with Elo gain over time.

English

+290 Elo

#1: claude-opus-4-6-thinking (anthropic) · 1525

29 months tracked11 crown changes

Hard Prompts

+260 Elo

#1: claude-fable-5 (anthropic) · 1534

24 months tracked12 crown changes

Coding

+243 Elo

#1: claude-fable-5 (anthropic) · 1553

23 months tracked15 crown changes

Chinese

+235 Elo

#1: claude-opus-4-6 (anthropic) · 1565

29 months tracked18 crown changes

Multi-Turn

+217 Elo

#1: gpt-5.2-high (openai) · 1525

26 months tracked12 crown changes

Math

+200 Elo

#1: claude-fable-5 (anthropic) · 1543

19 months tracked15 crown changes

Creative Writing

+188 Elo

#1: claude-fable-5 (anthropic) · 1508

22 months tracked9 crown changes

Instruction Following

+159 Elo

#1: claude-fable-5 (anthropic) · 1513

19 months tracked10 crown changes

Korean

+135 Elo

#1: gpt-5.5 (openai) · 1507

26 months tracked12 crown changes

Japanese

+128 Elo

#1: claude-fable-5 (anthropic) · 1531

26 months tracked11 crown changes

Vision Arena: Multimodal Model Rankings

More

How multimodal (vision) models have improved over time. Currently led by claude-fable-5 at 1334 Elo.

Openly Licensed Leaders Over Time

Every model that led rows Arena classified with a non-proprietary license.

vicuna-13b

LMSYS · 2023-05

1075

guanaco-33b

UW · 2023-06

1065

vicuna-33b

LMSYS · 2023-07

1096

wizardlm-70b

microsoft · 2023-10

1099

mixtral-8x7b-instruct-v0.1

mistral · 2023-12

1059

qwen1.5-72b-chat

alibaba · 2024-02

1147

llama-3-70b-instruct

meta · 2024-04

1208

gemma-2-27b-it

google · 2024-06

1214

llama-3.1-405b-instruct

meta · 2024-07

1262

llama-3.1-405b-instruct-bf16

meta · 2024-09

1267

llama-3.1-nemotron-70b-instruct

nvidia · 2024-10

1271

athene-v2-chat

NexusFlow · 2024-11

1274

deepseek-v3

deepseek · 2024-12

1315

deepseek-r1

deepseek · 2025-01

1358

deepseek-v3-0324

deepseek · 2025-03

1370

deepseek-r1-0528

deepseek · 2025-06

1424

glm-4.5

zai · 2025-08

1429

qwen3-vl-235b-a22b-instruct

alibaba · 2025-09

1428

glm-4.6

zai · 2025-10

1439

kimi-k2.5-thinking

moonshot · 2026-01

1446

qwen3.5-397b-a17b

alibaba · 2026-02

1452

glm-5.1

zai · 2026-04

1468

Monthly Rankings

Full top-20 Arena rankings for any month. Scroll through 39 snapshots from 2023-05 to 2026-07.

#ModelOrgLicenseEloVotes
1claude-opus-4-6-thinkinganthropicProprietary151262,772
2claude-opus-4-6anthropicProprietary150766,604
3claude-fable-5anthropicProprietary150414,575
4claude-opus-4-7-thinkinganthropicProprietary149950,483
5gpt-5.4openaiProprietary149961,783
6gpt-5.4-mini-highopenaiProprietary149657,600
7gpt-5.4-highopenaiProprietary149458,808
8gpt-5.2-highopenaiProprietary149347,421
9claude-opus-4-7anthropicProprietary149251,584
10muse-spark-1.1metaProprietary14917,777
11gemini-3.6-flashgoogleProprietary14914,747
12gpt-5.5openaiProprietary149047,031
13gemini-3.5-flash-highgoogleProprietary149010,092
14gemini-3-progoogleProprietary149040,384
15gemini-3.1-pro-previewgoogleProprietary148984,142
16gpt-5.5-highopenaiProprietary148845,597
17gpt-5.2openaiProprietary148676,972
18gpt-5.6-sol-xhighopenaiProprietary14866,117
19qwen3.7-max-previewalibabaProprietary14853,714
20gpt-5.1openaiProprietary148542,489

Share These Insights

Found this useful? Share it with your team or cite it in your research.

Frequently Asked Questions

What is the Arena Elo leaderboard?

Arena is a crowdsourced preference test where users vote on anonymous side-by-side model responses. Its ratings summarize relative preference within the models, prompts, voters, and ranking method present in each published snapshot. They are useful evidence, but they do not measure every capability or convert directly into task success rates.

How much have AI models improved since 2023?

The published monthly leader moved from 1075 (vicuna-13b) in May 2023 to 1512 (claude-opus-4-6-thinking) in 2026-07, a difference of 437 rating points across 39 monthly snapshots. This is a historical preference trend, not a controlled estimate of capability growth.

Which company has dominated the AI leaderboard?

openai has held the #1 position for 15 out of 39 months (38%), followed by google (8 months) and anthropic (7 months). The crown has changed hands 18 times since May 2023.

Are openly licensed LLMs catching up to proprietary models?

The highest-rated openly licensed model is glm-5.1 at 1475. It trails claude-opus-4-6-thinking by 37 rating points in the 2026-07 snapshot. The closest post-2023 gap was 0 points in 2025-01.

How does Arena Elo differ from benchmark scores?

Arena ratings summarize blind user preferences across the prompts and models in its pool. Benchmarks instead test defined tasks under a specified scoring protocol. Neither is universally stronger: preference ratings are broad but population-dependent, while benchmarks are easier to interpret by capability but can saturate, leak, or reward a narrow setup.

Which AI model is best for coding?

According to Arena coding Elo, claude-fable-5 (anthropic) currently leads with an Elo of 1553. The coding category has seen 15 crown changes over 23 months.

What models are best for math and reasoning?

The current Arena math leader is claude-fable-5 (anthropic) at 1543 rating points. The category leader changed 15 times across 19 monthly snapshots.

Where does this data come from?

All ratings come from the Arena Leaderboard Dataset on Hugging Face, maintained by Arena Intelligence. We process the text, text_style_control, vision, and webdev subsets, keep the latest published board in each calendar month, deduplicate repeated display-model rows, and recompute the visible ordinal order from published ratings.

Data attribution: All Elo ratings on this page come from the Arena Leaderboard Dataset by Arena Intelligence (lmarena-ai), available on HuggingFace. Data covers text, text_style_control, vision, webdev arena subsets.

We process and visualize this data to provide historical comparisons. We do not generate the underlying Elo ratings. For BenchLM's own benchmark-based rankings, see the main leaderboard or the AI Race timeline.