# AI2D test split (AI2D_TEST)

> A diagram understanding benchmark focused on scientific and educational visual question answering.

Canonical page: https://benchlm.ai/benchmarks/ai2dtest

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About AI2D_TEST

- Year: 2026
- Tasks: Diagram understanding
- Format: Diagram-grounded QA
- Difficulty: Structured visual reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

AI2D-style tasks matter because diagrams compress structure differently from photos or office documents. They test whether a model can parse arrows, labels, and spatial relations in technical illustrations.

AI2D_TEST is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 92.7% |
| 2 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 88.5% |
| 3 | [ZAYA1-VL-8B](/models/zaya1-vl-8b) | Zyphra | 87.5% |

## FAQ

### What does AI2D_TEST measure?

A diagram understanding benchmark focused on scientific and educational visual question answering.

### Which model scores highest on AI2D_TEST?

Qwen3.6-35B-A3B by Alibaba currently leads with a score of 92.7% on AI2D_TEST.

### How many models are evaluated on AI2D_TEST?

3 AI models have been evaluated on AI2D_TEST on BenchLM.

### Does AI2D_TEST affect BenchLM's overall score?

Not directly. AI2D_TEST is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AI2D_TEST

- [Qwen3.6-35B-A3B vs Nemotron 3 Nano Omni 30B A3B](/compare/nemotron-3-nano-omni-30b-a3b-vs-qwen3-6-35b-a3b)
- [Nemotron 3 Nano Omni 30B A3B vs ZAYA1-VL-8B](/compare/nemotron-3-nano-omni-30b-a3b-vs-zaya1-vl-8b)
