# MMAnswerBench

> A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.

Canonical page: https://benchlm.ai/benchmarks/mmanswerbench

- Category: [Mathematics](/math)
- Last updated: September 18, 2026

## About MMAnswerBench

- Year: 2026
- Tasks: Multimodal math questions
- Format: Visual and structured mathematical QA
- Difficulty: Advanced mathematical reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

MMAnswerBench matters because text-only math ability does not guarantee strong performance when the relevant information is embedded in diagrams, tables, or other visual inputs. It acts as a multimodal math transfer check.

MMAnswerBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GLM-5.2](/models/glm-5-2) | Z.AI | 91.0% |
| 2 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 86.0% |
| 3 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 84.0% |
| 4 | [GLM-5.1](/models/glm-5-1) | Z.AI | 83.8% |
| 5 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 83.8% |
| 6 | [GLM-5](/models/glm-5) | Z.AI | 82.5% |
| 7 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 81.8% |
| 8 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 80.9% |
| 9 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 80.8% |
| 10 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 78.9% |

## FAQ

### What does MMAnswerBench measure?

A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.

### Which model scores highest on MMAnswerBench?

GLM-5.2 by Z.AI currently leads with a score of 91.0% on MMAnswerBench.

### How many models are evaluated on MMAnswerBench?

10 AI models have been evaluated on MMAnswerBench on BenchLM.

### Does MMAnswerBench affect BenchLM's overall score?

Not directly. MMAnswerBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MMAnswerBench

- [GLM-5.2 vs Kimi K2.6](/compare/glm-5-2-vs-kimi-2-6)
- [Kimi K2.6 vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-kimi-2-6)
- [Claude Opus 4.5 vs GLM-5.1](/compare/claude-opus-4-5-vs-glm-5-1)
- [GLM-5.1 vs Qwen3.6 Plus](/compare/glm-5-1-vs-qwen3-6-plus)
