OCRBench V2
We show this table for reference; we do not rank on it.
A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.
Benchmark score on OCRBench V2 — September 27, 2026
We compile the OCRBench V2 rows from provider self-reports. Qwen3.8 Max leads the table at 74.2%, followed by Qwen3.7 Plus (70.7%) and Interfaze Beta (70.7%). We do not use these results to rank models overall.
Qwen3.8 Max
Alibaba
Qwen3.7 Plus
Alibaba
Interfaze Beta
Interfaze
5 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (5 models)
ScoreAmong the reported OCRBench V2 rows, Qwen3.8 Max is first at 74.2%. The third row is 3.5 points behind. The broader top-10 range is 26.7 points, so the table still separates the published systems.
5 models have been evaluated on OCRBench V2. The benchmark falls in the Multimodal & Grounded category. OCRBench V2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About OCRBench V2
Year
2025
Tasks
Image OCR tasks
Format
Accuracy
Difficulty
Native visual text understanding
OCRBench V2 evaluates whether multimodal models can extract visual text directly from images before downstream reasoning or structure extraction. BenchLM stores Interfaze's reported score as a display-only OCR row.
Freshness and provenance
Version
OCRBench V2 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does OCRBench V2 measure?
A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.
Which model scores highest on OCRBench V2?
Qwen3.8 Max by Alibaba currently leads with a score of 74.2% on OCRBench V2.
How many models are evaluated on OCRBench V2?
5 AI models have been evaluated on OCRBench V2 on BenchLM.
Compare top models on OCRBench V2
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.