# MMLU-Pro first-party comparison snapshot (MMLU-Pro (Arcee))

> A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.

Canonical page: https://benchlm.ai/benchmarks/mmluproarcee

- Category: [Knowledge](/knowledge)
- Last updated: September 18, 2026

## About MMLU-Pro (Arcee)

- Year: 2026
- Tasks: Professional academic QA
- Format: 10-way multiple choice
- Difficulty: Professional level
- Paper: [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking)

BenchLM stores this chart-specific MMLU-Pro row separately so it does not overwrite the standardized weighted MMLU-Pro benchmark values.

MMLU-Pro (Arcee) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (6 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 89.1% |
| 2 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 87.1% |
| 3 | [GLM-5](/models/glm-5) | Z.AI | 85.8% |
| 4 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 83.4% |
| 5 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 80.8% |
| 6 | [Trinity-Large-Preview](/models/trinity-large-preview) | Arcee AI | 75.2% |

## FAQ

### What does MMLU-Pro (Arcee) measure?

A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.

### Which model scores highest on MMLU-Pro (Arcee)?

Claude Opus 4.6 by Anthropic currently leads with a score of 89.1% on MMLU-Pro (Arcee).

### How many models are evaluated on MMLU-Pro (Arcee)?

6 AI models have been evaluated on MMLU-Pro (Arcee) on BenchLM.

### Does MMLU-Pro (Arcee) affect BenchLM's overall score?

Not directly. MMLU-Pro (Arcee) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MMLU-Pro (Arcee)

- [Claude Opus 4.6 vs Kimi K2.5](/compare/claude-opus-4-6-vs-kimi-k2-5)
- [Kimi K2.5 vs GLM-5](/compare/glm-5-vs-kimi-k2-5)
- [GLM-5 vs Trinity-Large-Thinking](/compare/glm-5-vs-trinity-large-thinking)
- [Trinity-Large-Thinking vs MiniMax M2.7](/compare/minimax-m2-7-vs-trinity-large-thinking)
