# PaperBench

> A research-reproduction benchmark that asks agents to recreate the contributions of AI papers from the paper alone.

Canonical page: https://benchlm.ai/benchmarks/paperbench

- Category: [Coding](/coding)
- Last updated: September 22, 2026

## About PaperBench

- Year: 2026
- Tasks: AI research-paper reproduction
- Format: Long-horizon agent evaluation
- Difficulty: Frontier autonomous research and engineering
- Paper: [Qwen3.8-Max: A New Bar for Coding and Cowork](https://qwen.ai/blog?id=qwen3.8)

Qwen evaluates Qwen3.8-Max in PaperBench's BasicAgent setting under Code-Dev mode, using Claude Opus 4.6 as judge and averaging three runs of up to 12 hours. We keep this provider-run score display-only because the agent setup and judge are part of the result.

PaperBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 93.0% |

## FAQ

### What does PaperBench measure?

A research-reproduction benchmark that asks agents to recreate the contributions of AI papers from the paper alone.

### Which model scores highest on PaperBench?

Qwen3.8 Max by Alibaba currently leads with a score of 93.0% on PaperBench.

### How many models are evaluated on PaperBench?

1 AI models have been evaluated on PaperBench on BenchLM.

### Does PaperBench affect BenchLM's overall score?

Not directly. PaperBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
