Skip to main content
BenchLM
Data

LongBench v2

Data verified 38 confirmed releases in the last 30 daysFollow model changes

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

Top models on LongBench v2 — October 7, 2026

As of October 7, 2026, Qwen3.8 Max leads the LongBench v2 leaderboard with 66.3% , followed by Beam (65.5%) and Claude Opus 4.5 (64.4%).

17 modelsReasoning25% of category scoreCurrentUpdated October 7, 2026

LongBench v2 leaderboard
RankModel / configurationScoreParameters (B)Open / closed
1Qwen3.8 MaxAlibaba
66.3%
Not reportedOpen
2BeamReflection AI
65.5%
Not reportedPending
3Claude Opus 4.5Anthropic
64.4%
Not reportedClosed
4Qwen3.5 397BAlibaba
63.2%
Not reportedOpen
5Qwen3.6 PlusAlibaba
62%
Not reportedClosed
6Nemotron 3 UltraNVIDIA
61.9%
Not reportedOpen
7Kimi K2.5Moonshot AI
61%
Not reportedOpen
8GLM-5Z.AI
60.8%
Not reportedOpen
9Qwen3.5-27BAlibaba
60.6%
Not reportedOpen
10Qwen3.5-122B-A10BAlibaba
60.2%
Not reportedOpen
11Agents-A1InternScience
60.2%
Not reportedOpen
12Qwen3.5-35B-A3BAlibaba
59%
Not reportedOpen
13Agents-A1-4BInternScience
52.1%
Not reportedOpen
14DeepSeek V4 Pro BaseDeepSeek
51.5%
Not reportedOpen
15DeepSeek V4 Flash BaseDeepSeek
44.7%
Not reportedOpen
16MiniCPM5-2BOpenBMB
43.7%
Not reportedOpen
17LLaDA2.2-miniInclusionAI
35.0%
Not reportedOpen

According to BenchLM.ai, Qwen3.8 Max leads the LongBench v2 benchmark with a score of 66.3%, followed by Beam (65.5%) and Claude Opus 4.5 (64.4%). The top models are clustered within 1.9 points, suggesting this benchmark is nearing saturation for frontier models.

17 models have been evaluated on LongBench v2. The benchmark falls in the Reasoning category. The Reasoning leaderboard ranks models by a weighted category score, and LongBench v2 contributes 25% of it. BenchAlign v5.8 also uses it in the overall ranking.

About LongBench v2

Year

2025

Tasks

Long-context tasks

Format

Extended-context retrieval and reasoning

Difficulty

Hard long-context

LongBench v2 is useful because context-window size alone is not a capability. It measures whether a model can retain, retrieve, and reason over long inputs effectively.

Freshness and provenance

Version

LongBench v2 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does LongBench v2 measure?

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

Which model scores highest on LongBench v2?

Qwen3.8 Max by Alibaba currently leads with a score of 66.3% on LongBench v2.

How many models are evaluated on LongBench v2?

17 AI models have published results on LongBench v2 in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version LongBench v2 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.