Skip to main content
BenchLM
Data

ERQA

We show this table for reference; we do not rank on it.

Qwen3.8 Max has the highest reported ERQA score at 77.8%, ahead of Qwen3.8-Flash-Next (72.3%) and Seed 2.1 Pro (72.0%) among 13 models. We compile the table from provider self-reports and secondary reports and keep it for display only; it does not affect overall rankings.

Data verified 38 confirmed releases in the last 30 daysFollow model changes

A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

Benchmark score on ERQA — October 7, 2026

We compile the ERQA rows from provider self-reports and secondary reports. Qwen3.8 Max leads the table at 77.8%, followed by Qwen3.8-Flash-Next (72.3%) and Seed 2.1 Pro (72.0%). We do not use these results to rank models overall.

13 modelsMultimodal & GroundedCurrentDisplay onlyUpdated October 7, 2026

Benchmark score results for ERQA
RankModel / configurationScoreParameters (B)Open / closed
1Qwen3.8 MaxAlibaba
77.8%
Not reportedOpen
2Qwen3.8-Flash-NextAlibaba
72.3%
Not reportedOpen
3Seed 2.1 ProByteDance
72.0%
Not reportedClosed
4Seed 2.1 TurboByteDance
71.3%
Not reportedClosed
5Qwen3.8-Omni-FlashAlibaba
71.0%
Not reportedClosed
6Qwen3.7 PlusAlibaba
69.8%
Not reportedClosed
7Gemini 3.1 ProGoogle
69.4%
Not reportedClosed
8Qwen3.8-27BAlibaba
65.5%
Not reportedOpen
9GPT-5.4OpenAI
65.4%
Not reportedClosed
10Muse SparkMeta
64.7%
Not reportedClosed
11Qwen3.6-27BAlibaba
62.5%
Not reportedOpen
12Grok 4.20xAI
54.1%
Not reportedClosed
13Claude Opus 4.6Anthropic
51.6%
Not reportedClosed

Among the reported ERQA rows, Qwen3.8 Max is first at 77.8%. The third row is 5.8 points behind. The broader top-10 range is 13.1 points, so the table still separates the published systems.

13 models have been evaluated on ERQA. The benchmark falls in the Multimodal & Grounded category. ERQA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ERQA

Year

2026

Tasks

Evidence-based visual QA

Format

Grounded image reasoning

Difficulty

Grounded multimodal reasoning

ERQA is useful as a grounded reasoning check because it emphasizes answer correctness tied to visual evidence rather than fluent but ungrounded descriptions.

Freshness and provenance

Version

ERQA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ERQA measure?

A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

Which model scores highest on ERQA?

Qwen3.8 Max by Alibaba currently leads with a score of 77.8% on ERQA.

How many models are evaluated on ERQA?

13 AI models have published results on ERQA in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version ERQA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.