Watch the AI race unfold
Scrub the timeline to travel through 18 months of model releases. See who held the crown, who climbed the ranks, and how the benchmark landscape evolved.
Monthly Race Snapshot
Export or embed the live race view for the currently selected month.
More
Grok 4.6
xAI
Releases this month
19 modelsProvider Race
Cumulative avg. top-3 score through Aug 2026Benchmark Health
More
How fresh are the benchmarks we use to score models? Green means the benchmark is actively separating models. Red means scores are bunching up or the benchmark is outdated.
Agentic
82 benchmarksCoding
64 benchmarksReasoning
28 benchmarksMultimodal
65 benchmarksKnowledge
46 benchmarksMultilingual
13 benchmarksInstruction Following
4 benchmarksMath
32 benchmarksWant deeper historical analysis?
Explore 21 months of Arena Elo ratings — crown changes, provider dominance, open-source gap tracking, and more.
LLM Leaderboard History