# MedXpertQA Multimodal (MedXpertQA (MM))

> A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

Canonical page: https://benchlm.ai/benchmarks/medxpertqamm

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About MedXpertQA (MM)

- Year: 2026
- Tasks: 2,000 multimodal medical questions
- Format: Medical visual MCQ
- Difficulty: Clinical multimodal reasoning
- Paper: [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology)

Meta describes the multimodal MedXpertQA variant as 2,000 clinically grounded medical questions with five answer choices. BenchLM stores it as a display-only health and multimodal reference.

MedXpertQA (MM) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 81.3% |
| 2 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 80.4% |
| 3 | [Muse Spark](/models/muse-spark) | Meta | 78.4% |
| 4 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 77.1% |
| 5 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 71.0% |
| 6 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 65.8% |
| 7 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 64.8% |
| 8 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 48.7% |

## FAQ

### What does MedXpertQA (MM) measure?

A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

### Which model scores highest on MedXpertQA (MM)?

Gemini 3.1 Pro by Google currently leads with a score of 81.3% on MedXpertQA (MM).

### How many models are evaluated on MedXpertQA (MM)?

8 AI models have been evaluated on MedXpertQA (MM) on BenchLM.

### Does MedXpertQA (MM) affect BenchLM's overall score?

Not directly. MedXpertQA (MM) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MedXpertQA (MM)

- [Gemini 3.1 Pro vs Qwen3.8 Max](/compare/gemini-3-1-pro-vs-qwen3-8-max)
- [Qwen3.8 Max vs Muse Spark](/compare/muse-spark-vs-qwen3-8-max)
- [Muse Spark vs GPT-5.4](/compare/gpt-5-4-vs-muse-spark)
- [GPT-5.4 vs Qwen3.7 Plus](/compare/gpt-5-4-vs-qwen3-7-plus)
