# Artificial Analysis MMLU-Pro (AA MMLU-Pro)

> An independently evaluated MMLU-Pro result from Artificial Analysis.

Canonical page: https://benchlm.ai/benchmarks/aammlupro

- Category: [Knowledge](/knowledge)
- Last updated: September 27, 2026

## About AA MMLU-Pro

- Year: 2026
- Tasks: Professional multi-subject questions
- Format: Accuracy
- Difficulty: Professional knowledge and reasoning
- Paper: [Artificial Analysis MMLU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmlu-pro)

Stored separately from the core MMLU-Pro lane because Artificial Analysis controls the evaluation configuration.

AA MMLU-Pro is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 89.8% |
| 2 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 89.5% |
| 3 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 88.9% |

## FAQ

### What does AA MMLU-Pro measure?

An independently evaluated MMLU-Pro result from Artificial Analysis.

### Which model scores highest on AA MMLU-Pro?

Gemini 3 Pro by Google currently leads with a score of 89.8% on AA MMLU-Pro.

### How many models are evaluated on AA MMLU-Pro?

3 AI models have been evaluated on AA MMLU-Pro on BenchLM.

### Does AA MMLU-Pro affect BenchLM's overall score?

Not directly. AA MMLU-Pro is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AA MMLU-Pro

- [Gemini 3 Pro vs Claude Opus 4.5 Thinking](/compare/claude-opus-4-5-thinking-vs-gemini-3-pro)
- [Claude Opus 4.5 Thinking vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-claude-opus-4-5-thinking)
