# LongBench v2

> A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

Canonical page: https://benchlm.ai/benchmarks/longbench-v2

- Category: [Reasoning](/reasoning)
- Last updated: September 10, 2026

## About LongBench v2

- Year: 2025
- Tasks: Long-context tasks
- Format: Extended-context retrieval and reasoning
- Difficulty: Hard long-context
- Paper: [LongBench v2](https://arxiv.org/abs/2412.15204)

LongBench v2 is useful because context-window size alone is not a capability. It measures whether a model can retain, retrieve, and reason over long inputs effectively.

LongBench v2 is currently weighted in BenchLM's scoring formula. The Reasoning category carries 17% of the overall score, and LongBench v2 contributes 25% of that category score.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 66.3% |
| 2 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 64.4% |
| 3 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 63.2% |
| 4 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 62% |
| 5 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 61.9% |
| 6 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 61% |
| 7 | [GLM-5](/models/glm-5) | Z.AI | 60.8% |
| 8 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 60.6% |
| 9 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 60.2% |
| 10 | [Agents-A1](/models/agents-a1) | InternScience | 60.2% |
| 11 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 59% |
| 12 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 52.1% |
| 13 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 43.7% |
| 14 | [LLaDA2.2-mini](/models/llada2-2-mini) | InclusionAI | 35.0% |

## FAQ

### What does LongBench v2 measure?

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

### Which model scores highest on LongBench v2?

Qwen3.8 Max by Alibaba currently leads with a score of 66.3% on LongBench v2.

### How many models are evaluated on LongBench v2?

14 AI models have been evaluated on LongBench v2 on BenchLM.

### Does LongBench v2 affect BenchLM's overall score?

Yes. LongBench v2 is a weighted benchmark inside the Reasoning category, which carries 17% of BenchLM's overall score. LongBench v2 itself contributes 25% of that category score.

## Compare Top Models on LongBench v2

- [Qwen3.8 Max vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-qwen3-8-max)
- [Claude Opus 4.5 vs Qwen3.5 397B](/compare/claude-opus-4-5-vs-qwen3-5-397b)
- [Qwen3.5 397B vs Qwen3.6 Plus](/compare/qwen3-5-397b-vs-qwen3-6-plus)
- [Qwen3.6 Plus vs Nemotron 3 Ultra](/compare/nemotron-3-ultra-vs-qwen3-6-plus)
