ERQA
We show this table for reference; we do not rank on it.
Qwen3.8 Max has the highest reported ERQA score at 77.8%, ahead of Qwen3.8-Flash-Next (72.3%) and Seed 2.1 Pro (72.0%) among 13 models. We compile the table from provider self-reports and secondary reports and keep it for display only; it does not affect overall rankings.
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
Benchmark score on ERQA — October 7, 2026
We compile the ERQA rows from provider self-reports and secondary reports. Qwen3.8 Max leads the table at 77.8%, followed by Qwen3.8-Flash-Next (72.3%) and Seed 2.1 Pro (72.0%). We do not use these results to rank models overall.
Qwen3.8 Max
Alibaba
Qwen3.8-Flash-Next
Alibaba
Seed 2.1 Pro
ByteDance
13 modelsMultimodal & GroundedCurrentDisplay onlyUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Qwen3.8 MaxAlibaba | 77.8% | Not reported | Open |
| 2 | Qwen3.8-Flash-NextAlibaba | 72.3% | Not reported | Open |
| 3 | Seed 2.1 ProByteDance | 72.0% | Not reported | Closed |
| 4 | Seed 2.1 TurboByteDance | 71.3% | Not reported | Closed |
| 5 | Qwen3.8-Omni-FlashAlibaba | 71.0% | Not reported | Closed |
| 6 | Qwen3.7 PlusAlibaba | 69.8% | Not reported | Closed |
| 7 | Gemini 3.1 ProGoogle | 69.4% | Not reported | Closed |
| 8 | Qwen3.8-27BAlibaba | 65.5% | Not reported | Open |
| 9 | GPT-5.4OpenAI | 65.4% | Not reported | Closed |
| 10 | Muse SparkMeta | 64.7% | Not reported | Closed |
| 11 | Qwen3.6-27BAlibaba | 62.5% | Not reported | Open |
| 12 | Grok 4.20xAI | 54.1% | Not reported | Closed |
| 13 | Claude Opus 4.6Anthropic | 51.6% | Not reported | Closed |
Among the reported ERQA rows, Qwen3.8 Max is first at 77.8%. The third row is 5.8 points behind. The broader top-10 range is 13.1 points, so the table still separates the published systems.
13 models have been evaluated on ERQA. The benchmark falls in the Multimodal & Grounded category. ERQA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ERQA
Year
2026
Tasks
Evidence-based visual QA
Format
Grounded image reasoning
Difficulty
Grounded multimodal reasoning
ERQA is useful as a grounded reasoning check because it emphasizes answer correctness tied to visual evidence rather than fluent but ungrounded descriptions.
Freshness and provenance
Version
ERQA 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ERQA measure?
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
Which model scores highest on ERQA?
Qwen3.8 Max by Alibaba currently leads with a score of 77.8% on ERQA.
How many models are evaluated on ERQA?
13 AI models have published results on ERQA in the BenchLM catalog.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.