Benchmark profile
LVBench
A long-video understanding benchmark for retrieving and reasoning over information distributed across extended video inputs.
Data verified 21 confirmed releases in the last 30 daysSee provider release alertsBenchmark score on LVBench — August 3, 2026
BenchLM mirrors the published score view for LVBench. Qwen3.8 Max leads the public snapshot at 81.8%. BenchLM does not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout LVBench
Year
2026
Tasks
Long-form video question answering
Format
Long-video understanding score
Difficulty
Extended temporal reasoning
We store Qwen's standard LVBench result without the separate memory system. The provider also reports an LVBench-with-memory row, which remains distinct and is not collapsed into this value.
BenchLM freshness & provenance
Version
LVBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does LVBench measure?
A long-video understanding benchmark for retrieving and reasoning over information distributed across extended video inputs.
Which model scores highest on LVBench?
Qwen3.8 Max by Alibaba currently leads with a score of 81.8% on LVBench.
How many models are evaluated on LVBench?
1 AI models have been evaluated on LVBench on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.