# ERQA

> A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

Canonical page: https://benchlm.ai/benchmarks/erqa

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: October 7, 2026

## About ERQA

- Year: 2026
- Tasks: Evidence-based visual QA
- Format: Grounded image reasoning
- Difficulty: Grounded multimodal reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

ERQA is useful as a grounded reasoning check because it emphasizes answer correctness tied to visual evidence rather than fluent but ungrounded descriptions.

ERQA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (13 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 77.8% |
| 2 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 72.3% |
| 3 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 72.0% |
| 4 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 71.3% |
| 5 | [Qwen3.8-Omni-Flash](/models/qwen3-8-omni-flash) | Alibaba | 71.0% |
| 6 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 69.8% |
| 7 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 69.4% |
| 8 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 65.5% |
| 9 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 65.4% |
| 10 | [Muse Spark](/models/muse-spark) | Meta | 64.7% |
| 11 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 62.5% |
| 12 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 54.1% |
| 13 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 51.6% |

## FAQ

### What does ERQA measure?

A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

### Which model scores highest on ERQA?

Qwen3.8 Max by Alibaba currently leads with a score of 77.8% on ERQA.

### How many models are evaluated on ERQA?

13 AI models have published results on ERQA in the BenchLM catalog.

### Does ERQA affect BenchLM's overall score?

Not directly. ERQA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on ERQA

- [Qwen3.8 Max vs Qwen3.8-Flash-Next](/compare/qwen3-8-flash-next-vs-qwen3-8-max)
- [Qwen3.8-Flash-Next vs Seed 2.1 Pro](/compare/qwen3-8-flash-next-vs-seed-2-1-pro)
- [Seed 2.1 Pro vs Seed 2.1 Turbo](/compare/seed-2-1-pro-vs-seed-2-1-turbo)
- [Seed 2.1 Turbo vs Qwen3.8-Omni-Flash](/compare/qwen3-8-omni-flash-vs-seed-2-1-turbo)
