# MathVision

> A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.

Canonical page: https://benchlm.ai/benchmarks/mathvision

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About MathVision

- Year: 2026
- Tasks: Visually grounded math problems
- Format: Image + math reasoning
- Difficulty: Advanced multimodal mathematics
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

MathVision matters because text-only math ability does not guarantee strong performance when the relevant information is embedded in images, geometry diagrams, or formatted equations.

MathVision is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (19 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 95.2% |
| 2 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 94.3% |
| 3 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 92.6% |
| 4 | [Qwen3.8-Omni-Flash](/models/qwen3-8-omni-flash) | Alibaba | 91.8% |
| 5 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 90.6% |
| 6 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 90.3% |
| 7 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 90.1% |
| 8 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 90.0% |
| 9 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 88.6% |
| 10 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 88.0% |
| 11 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 87.7% |
| 12 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 87.4% |
| 13 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 86.6% |
| 14 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 86.2% |
| 15 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 86.0% |
| 16 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 83.9% |
| 17 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 83.0% |
| 18 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 79.7% |
| 19 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 74.3% |

## FAQ

### What does MathVision measure?

A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.

### Which model scores highest on MathVision?

Qwen3.8 Max by Alibaba currently leads with a score of 95.2% on MathVision.

### How many models are evaluated on MathVision?

19 AI models have been evaluated on MathVision on BenchLM.

### Does MathVision affect BenchLM's overall score?

Not directly. MathVision is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MathVision

- [Qwen3.8 Max vs Kimi K3](/compare/kimi-k3-vs-qwen3-8-max)
- [Kimi K3 vs Seed 2.1 Pro](/compare/kimi-k3-vs-seed-2-1-pro)
- [Seed 2.1 Pro vs Qwen3.8-Omni-Flash](/compare/qwen3-8-omni-flash-vs-seed-2-1-pro)
- [Qwen3.8-Omni-Flash vs Qwen3.8-Flash-Next](/compare/qwen3-8-flash-next-vs-qwen3-8-omni-flash)
