WildBench
We show this table for reference; we do not rank on it.
An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.
About WildBench
Year
2024
Tasks
1,024 real-world tasks
Format
Real-world task evaluation
Difficulty
Diverse real-world scenarios
WildBench bridges the gap between static benchmarks and human preference evaluations. Tasks are derived from real ChatGPT conversations, making it more representative of actual user needs than synthetic benchmarks.
Freshness and provenance
Version
WildBench 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does WildBench measure?
An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.
Which model scores highest on WildBench?
No models have been evaluated on WildBench yet.
How many models are evaluated on WildBench?
0 AI models have been evaluated on WildBench on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.