Skip to main content
BenchLM

LLM Leaderboard History: Arena Rankings Since 2023

claude-opus-5-max leads the 2026-09 Arena snapshot at 1505 rating points. Across 41 month-end snapshots since May 2023, the leader changed 19 times among 7 providers.

Arena data fetched

This is a history of measured user preference, not a universal capability curve. Ratings move with the model pool, prompt mix, voters, and Arena's methodology. Use the timeline to inspect leadership changes; use capability benchmarks for claims about coding, reasoning, or task accuracy.

Data: Arena Leaderboard Dataset 4 arena subsets, 14 categories

The Numbers

Key stats from 41 monthly Arena preference snapshots.

+430

Leader Rating Change

1075 to 1505

19

Crown Changes

41 months tracked

5mo

Longest Reign

gemini-2.5-pro

29 pts

Current Open Gap

kimi-k3-max trails the leader

10.8/mo

Leader Trend

Descriptive, not capability growth

7

Providers at #1

openai leads

Elo rating progression: May 2023 to today

More

Month-end Arena preference leaders. Proprietary models use a solid line; openly licensed models use a dashed line. The series is not adjusted for changes in the model pool or voting population.

Proprietary #1 Openly licensed #1

Elo Milestones

When the AI frontier crossed each Elo threshold for the first time.

1100

gpt-4-0314

2023-12 · openai

1150

gpt-4-0314

2024-01 · openai

1200

gpt-4-1106-preview

2024-02 · openai

1250

gpt-4-1106-preview

2024-02 · openai

1300

chatgpt-4o-latest

2024-08 · openai

1350

o1-2024-12-17

2024-12 · openai

1400

chatgpt-4o-latest-20250326

2025-03 · openai

1450

gemini-2.5-pro

2025-06 · google

1500

claude-opus-4-6

2026-02 · anthropic

Key Breakthroughs

Slide through 25 significant moments in the AI race. Each dot marks a breakthrough.

2023-06New Challenger

UW enters the race

guanaco-33b is the first UW model to take #1

2023202420252026
1 / 25

The Open-License Gap

More

Rating-point difference between the leading proprietary and openly licensed models. The latest gap is 29 points. The closest post-2023 snapshot was 0 points in 2025-01.

Crown change timeline

19 changes

Every time the #1 model changed hands since May 2023.

2023-06

vicuna-13b→guanaco-33b(UW)

2023-07

guanaco-33b→vicuna-33b(LMSYS)

2023-10

vicuna-33b→wizardlm-70b(microsoft)

2023-12

wizardlm-70b→gpt-4-0314(openai)

2024-02

gpt-4-0314→gpt-4-1106-preview(openai)

2024-03

gpt-4-1106-preview→claude-3-opus-20240229(anthropic)

2024-04

claude-3-opus-20240229→gpt-4-turbo-2024-04-09(openai)

2024-05

gpt-4-turbo-2024-04-09→gpt-4o-2024-05-13(openai)

Provider Dominance

Months at #1 on the Arena leaderboard since May 2023. Who has dominated the AI race?

openai15 months (37%)
anthropic9 months (22%)
google8 months (20%)
LMSYS4 months (10%)
microsoft2 months (5%)
deepseek2 months (5%)
UW1 month (2%)

Category breakdown: who wins at what?

More

Arena tracks Elo by category. Here's who leads in coding, math, creative writing, and more — with Elo gain over time.

English

+276 Elo

#1: claude-opus-4-6-high (anthropic) · 1512

31 months tracked12 crown changes

Chinese

+264 Elo

#1: claude-opus-5-max (anthropic) · 1594

31 months tracked19 crown changes

Hard Prompts

+260 Elo

#1: claude-fable-5 (anthropic) · 1533

26 months tracked14 crown changes

Coding

+241 Elo

#1: claude-opus-4-7-high (anthropic) · 1552

25 months tracked17 crown changes

Multi-Turn

+204 Elo

#1: claude-opus-4-6-high (anthropic) · 1512

28 months tracked13 crown changes

Math

+187 Elo

#1: claude-fable-5 (anthropic) · 1530

21 months tracked17 crown changes

Creative Writing

+185 Elo

#1: claude-fable-5 (anthropic) · 1505

24 months tracked9 crown changes

Instruction Following

+160 Elo

#1: claude-opus-4-6-high (anthropic) · 1514

21 months tracked11 crown changes

Korean

+153 Elo

#1: claude-opus-5-max (anthropic) · 1525

28 months tracked13 crown changes

Japanese

+111 Elo

#1: claude-opus-5-max (anthropic) · 1514

28 months tracked13 crown changes

Vision Arena: multimodal model rankings

More

How multimodal (vision) models have improved over time. Currently led by claude-fable-5 at 1330 Elo.

Openly licensed leaders over time

Every model that led rows Arena classified with a non-proprietary license.

vicuna-13b

LMSYS · 2023-05

1075

guanaco-33b

UW · 2023-06

1065

vicuna-33b

LMSYS · 2023-07

1096

wizardlm-70b

microsoft · 2023-10

1099

mixtral-8x7b-instruct-v0.1

mistral · 2023-12

1059

qwen1.5-72b-chat

alibaba · 2024-02

1147

llama-3-70b-instruct

meta · 2024-04

1208

gemma-2-27b-it

google · 2024-06

1214

llama-3.1-405b-instruct

meta · 2024-07

1262

llama-3.1-405b-instruct-bf16

meta · 2024-09

1267

llama-3.1-nemotron-70b-instruct

nvidia · 2024-10

1271

athene-v2-chat

NexusFlow · 2024-11

1274

deepseek-v3

deepseek · 2024-12

1315

deepseek-r1

deepseek · 2025-01

1358

deepseek-v3-0324

deepseek · 2025-03

1370

deepseek-r1-0528

deepseek · 2025-06

1424

glm-4.5

zai · 2025-08

1429

qwen3-vl-235b-a22b-instruct

alibaba · 2025-09

1428

glm-4.6

zai · 2025-10

1439

kimi-k2.5-thinking

moonshot · 2026-01

1446

qwen3.5-397b-a17b

alibaba · 2026-02

1452

glm-5.1

zai · 2026-04

1468

kimi-k3-max

moonshot · 2026-07

1473

Monthly Rankings

Full top-20 Arena rankings for any month. Scroll through 41 snapshots from 2023-05 to 2026-09.

#ModelOrgLicenseEloVotes
1claude-opus-5-maxanthropicProprietary150516,839
2claude-opus-5-highanthropicProprietary150434,617
3claude-opus-4-6-highanthropicProprietary150372,110
4claude-opus-4-6anthropicProprietary149876,032
5claude-fable-5anthropicProprietary149426,977
6gemini-3.7-flash-highgoogleProprietary14915,685
7claude-opus-4-7-highanthropicProprietary149060,123
8muse-spark-1.2 (xHigh)metaProprietary14893,244
9gemini-3.5-flash-highgoogleProprietary148433,874
10claude-opus-4-7anthropicProprietary148361,256
11gemini-3.1-pro-previewgoogleProprietary1480102,763
12qwen3.8-maxalibabaProprietary148013,136
13gemini-3-progoogleProprietary148040,675
14muse-spark-1.1metaProprietary147923,768
15gemini-3.6-flash-highgoogleProprietary147622,009
16kimi-k3-maxmoonshotOpen147617,895
17gemini-3.5-flash-mediumgoogleProprietary147632,274
18glm-5.3-maxzaiOpen14757,401
19qwen3.7-max-previewalibabaProprietary14743,710
20muse-sparkmetaProprietary147413,571

Share these insights

Found this useful? Share it with your team or cite it in your research.

Questions

What is the Arena Elo leaderboard?

Arena is a crowdsourced preference test where users vote on anonymous side-by-side model responses. Its ratings summarize relative preference within the models, prompts, voters, and ranking method present in each published snapshot. They are useful evidence, but they do not measure every capability or convert directly into task success rates.

How much have AI models improved since 2023?

The published monthly leader moved from 1075 (vicuna-13b) in May 2023 to 1505 (claude-opus-5-max) in 2026-09, a difference of 430 rating points across 41 monthly snapshots. This is a historical preference trend, not a controlled estimate of capability growth.

Which company has dominated the AI leaderboard?

openai has held the #1 position for 15 out of 41 months (37%), followed by anthropic (9 months) and google (8 months). The crown has changed hands 19 times since May 2023.

Are openly licensed LLMs catching up to proprietary models?

The highest-rated openly licensed model is kimi-k3-max at 1476. It trails claude-opus-5-max by 29 rating points in the 2026-09 snapshot. The closest post-2023 gap was 0 points in 2025-01.

How does Arena Elo differ from benchmark scores?

Arena ratings summarize blind user preferences across the prompts and models in its pool. Benchmarks instead test defined tasks under a specified scoring protocol. Neither is universally stronger: preference ratings are broad but population-dependent, while benchmarks are easier to interpret by capability but can saturate, leak, or reward a narrow setup.

Which AI model is best for coding?

According to Arena coding Elo, claude-opus-4-7-high (anthropic) currently leads with an Elo of 1552. The coding category has seen 17 crown changes over 25 months.

What models are best for math and reasoning?

The current Arena math leader is claude-fable-5 (anthropic) at 1530 rating points. The category leader changed 17 times across 21 monthly snapshots.

Where does this data come from?

All ratings come from the Arena Leaderboard Dataset on Hugging Face, maintained by Arena Intelligence. We process the text, text_style_control, vision, and webdev subsets, keep the latest published board in each calendar month, deduplicate repeated display-model rows, and recompute the visible ordinal order from published ratings.

Data attribution: All Elo ratings on this page come from the Arena Leaderboard Dataset by Arena Intelligence (lmarena-ai), available on HuggingFace. Data covers text, text_style_control, vision, webdev arena subsets.

We process and visualize this data to provide historical comparisons. We do not generate the underlying Elo ratings. For BenchLM's own benchmark-based rankings, see the main leaderboard or the AI Race timeline.