Skip to main content
BenchLM

Medical Long Context Reasoning (MLCR-AA) (MLCR-AA)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.

Benchmark score on MLCR-AA — September 27, 2026

We compile the MLCR-AA rows from secondary reports. Claude Fable 5.1 leads the table at 71.1%, followed by Claude Fable 5 (64.4%) and Claude Opus 5 (55.6%). We do not use these results to rank models overall.

16 modelsReasoningCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (16 models)

Score
1
Claude Fable 5.1Anthropic · Closed
71.1%
2
Claude Fable 5Anthropic · Closed
64.4%
3
Claude Opus 5Anthropic · Closed
55.6%
4
Claude Sonnet 5Anthropic · Closed
55.0%
5
GLM-5.3-FlashZ.AI · Open weight
51.1%
6
GLM-5.3Z.AI · Open weight
48.3%
7
Muse Spark 1.3Meta · Closed
43.3%
8
Kimi K3Moonshot AI · Closed
38.3%
9
GPT-6 AstraOpenAI · Closed
35.0%
10
GPT-5.6 SolOpenAI · Closed
26.1%
11
DeepSeek V4.1 FlashDeepSeek · Open weight
22.8%
12
Gemini 3.8 FlashGoogle · Closed
21.7%
13
Qwen3.8-27BAlibaba · Open weight
21.7%
14
Muse Glimmer 30BMeta · Open weight
20.0%
15
GPT-5.6 LunaOpenAI · Closed
19.4%
16
MiniMax M3MiniMax · Open weight
17.2%

Among the reported MLCR-AA rows, Claude Fable 5.1 is first at 71.1%. The third row is 15.5 points behind. The broader top-10 range is 45.0 points, so the table still separates the published systems.

16 models have been evaluated on MLCR-AA. The benchmark falls in the Reasoning category. MLCR-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MLCR-AA

Year

2026

Tasks

Long, fragmented medical-record reasoning

Format

Accuracy

Difficulty

Long-context medical reasoning

Independently benchmarked by Artificial Analysis. Display-only; not yet admitted as independent-run evidence.

Freshness and provenance

Version

MLCR-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MLCR-AA measure?

An open benchmark from Wisedocs, run by Artificial Analysis, measuring how well models reason over long, fragmented medical records with multi-document synthesis.

Which model scores highest on MLCR-AA?

Claude Fable 5.1 by Anthropic currently leads with a score of 71.1% on MLCR-AA.

How many models are evaluated on MLCR-AA?

16 AI models have been evaluated on MLCR-AA on BenchLM.

Last updated: September 27, 2026 · BenchLM version MLCR-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.