# MStar

> A general visual question-answering benchmark used in provider tables for real-image reasoning quality.

Canonical page: https://benchlm.ai/benchmarks/mstar

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About MStar

- Year: 2026
- Tasks: Real-image visual QA
- Format: Image-grounded QA
- Difficulty: General visual reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

MStar sits between broad multimodal reasoning and grounded VQA. It is useful for checking whether a model can answer real-image questions without the stronger domain structure of office or academic benchmarks.

MStar is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 81.4% |

## FAQ

### What does MStar measure?

A general visual question-answering benchmark used in provider tables for real-image reasoning quality.

### Which model scores highest on MStar?

Qwen3.6-27B by Alibaba currently leads with a score of 81.4% on MStar.

### How many models are evaluated on MStar?

1 AI models have been evaluated on MStar on BenchLM.

### Does MStar affect BenchLM's overall score?

Not directly. MStar is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
