Massive Multi-discipline Multimodal Understanding Pro (MMMU-Pro)
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
Data verified 33 confirmed releases in the last 30 daysSee provider release alertsTop models on MMMU-Pro — September 15, 2026
As of September 15, 2026, Gemini 3.1 Pro leads the MMMU-Pro leaderboard with 83.9% , followed by Gemini 3.5 Flash (83.6%) and GPT-5.6 Sol (83%).
Gemini 3.1 Pro
Gemini 3.5 Flash
GPT-5.6 Sol
OpenAI
40 modelsMultimodal & Grounded40% of category scoreRefreshingUpdated September 15, 2026
Leaderboard (40 models)
ScoreAccording to BenchLM.ai, Gemini 3.1 Pro leads the MMMU-Pro benchmark with a score of 83.9%, followed by Gemini 3.5 Flash (83.6%) and GPT-5.6 Sol (83%). The top models are clustered within 0.9 points, suggesting this benchmark is nearing saturation for frontier models.
40 models have been evaluated on MMMU-Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMMU-Pro contributes 40% of the category score, so strong performance here directly affects a model's overall ranking.
About MMMU-Pro
Year
2024
Tasks
Multimodal academic reasoning
Format
Image + text question answering
Difficulty
Frontier multimodal
MMMU-Pro extends the original MMMU setup with more difficult multimodal questions and stronger separation at the top end of the model market.
BenchLM freshness & provenance
Version
MMMU-Pro 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MMMU-Pro measure?
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
Which model scores highest on MMMU-Pro?
Gemini 3.1 Pro by Google currently leads with a score of 83.9% on MMMU-Pro.
How many models are evaluated on MMMU-Pro?
40 AI models have been evaluated on MMMU-Pro on BenchLM.
Compare Top Models on MMMU-Pro
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.