Skip to main content
BenchLM

OCRBench V2

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.

Benchmark score on OCRBench V2 — September 27, 2026

We compile the OCRBench V2 rows from provider self-reports. Qwen3.8 Max leads the table at 74.2%, followed by Qwen3.7 Plus (70.7%) and Interfaze Beta (70.7%). We do not use these results to rank models overall.

5 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (5 models)

Score
1
Qwen3.8 MaxAlibaba · Open weight
74.2%
2
Qwen3.7 PlusAlibaba · Closed
70.7%
3
Interfaze BetaInterfaze · Closed
70.7%
4
Ternary Bonsai 2 27BPrism ML · Open weight
56.9%
5
LFM2.5-VL-3BLiquidAI · Open weight
47.5%

Among the reported OCRBench V2 rows, Qwen3.8 Max is first at 74.2%. The third row is 3.5 points behind. The broader top-10 range is 26.7 points, so the table still separates the published systems.

5 models have been evaluated on OCRBench V2. The benchmark falls in the Multimodal & Grounded category. OCRBench V2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About OCRBench V2

Year

2025

Tasks

Image OCR tasks

Format

Accuracy

Difficulty

Native visual text understanding

OCRBench V2 evaluates whether multimodal models can extract visual text directly from images before downstream reasoning or structure extraction. BenchLM stores Interfaze's reported score as a display-only OCR row.

Freshness and provenance

Version

OCRBench V2 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does OCRBench V2 measure?

A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.

Which model scores highest on OCRBench V2?

Qwen3.8 Max by Alibaba currently leads with a score of 74.2% on OCRBench V2.

How many models are evaluated on OCRBench V2?

5 AI models have been evaluated on OCRBench V2 on BenchLM.

Last updated: September 27, 2026 · BenchLM version OCRBench V2 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.