# WildBench

> An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.

Canonical page: https://benchlm.ai/benchmarks/wildbench

- Category: [Reasoning](/reasoning)
- Last updated: September 27, 2026

## About WildBench

- Year: 2024
- Tasks: 1,024 real-world tasks
- Format: Real-world task evaluation
- Difficulty: Diverse real-world scenarios
- Paper: [WildBench: Benchmarking Language Models with Challenging Tasks from Real Users in the Wild](https://arxiv.org/abs/2406.04770)

WildBench bridges the gap between static benchmarks and human preference evaluations. Tasks are derived from real ChatGPT conversations, making it more representative of actual user needs than synthetic benchmarks.

WildBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does WildBench measure?

An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.

### Which model scores highest on WildBench?

No models have been evaluated on WildBench yet.

### How many models are evaluated on WildBench?

0 AI models have been evaluated on WildBench on BenchLM.

### Does WildBench affect BenchLM's overall score?

Not directly. WildBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
