# Medical Long Context Reasoning (MLCR-AA) (MLCR-AA)

> An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.

Canonical page: https://benchlm.ai/benchmarks/aamlcr

- Category: [Reasoning](/reasoning)
- Last updated: September 27, 2026

## About MLCR-AA

- Year: 2026
- Tasks: Long, fragmented medical-record reasoning
- Format: Accuracy
- Difficulty: Long-context medical reasoning
- Paper: [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa)

Independently benchmarked by Artificial Analysis. Display-only; not yet admitted as independent-run evidence.

MLCR-AA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (16 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 71.1% |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 64.4% |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 55.6% |
| 4 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 55.0% |
| 5 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 51.1% |
| 6 | [GLM-5.3](/models/glm-5-3) | Z.AI | 48.3% |
| 7 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 43.3% |
| 8 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 38.3% |
| 9 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 35.0% |
| 10 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 26.1% |
| 11 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 22.8% |
| 12 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 21.7% |
| 13 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 21.7% |
| 14 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 20.0% |
| 15 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 19.4% |
| 16 | [MiniMax M3](/models/minimax-m3) | MiniMax | 17.2% |

## FAQ

### What does MLCR-AA measure?

An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.

### Which model scores highest on MLCR-AA?

Claude Fable 5.1 by Anthropic currently leads with a score of 71.1% on MLCR-AA.

### How many models are evaluated on MLCR-AA?

16 AI models have been evaluated on MLCR-AA on BenchLM.

### Does MLCR-AA affect BenchLM's overall score?

Not directly. MLCR-AA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MLCR-AA

- [Claude Fable 5.1 vs Claude Fable 5](/compare/claude-fable-vs-claude-fable-5-1)
- [Claude Fable 5 vs Claude Opus 5](/compare/claude-fable-vs-claude-opus-5)
- [Claude Opus 5 vs Claude Sonnet 5](/compare/claude-opus-5-vs-claude-sonnet-5)
- [Claude Sonnet 5 vs GLM-5.3-Flash](/compare/claude-sonnet-5-vs-glm-5-3-flash)
