Skip to main content
BenchLM

Vals MedScribe (MedScribe)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI healthcare benchmark for whether models can support doctors with administrative work.

MedScribe score on MedScribe — September 26, 2026

We mirror the published medscribe score view for MedScribe. Claude Opus 5.5 leads the public snapshot at 91.43%, followed by Claude Fable 5.1 (91.29%) and Claude Sonnet 5.5 (91.10%). We do not use these results to rank models overall.

104 modelsAgenticCurrentDisplay onlyUpdated September 26, 2026

MedScribe score table (104 models)

Score
1
Claude Opus 5.5Anthropic · Closed
91.43%
2
Claude Fable 5.1Anthropic · Closed
91.29%
3
Claude Sonnet 5.5Anthropic · Closed
91.10%
4
Claude Opus 5Anthropic · Closed
90.98%
5
Muse Spark 1.2Meta · Closed
90.06%
6
Grok 4.7xAI · Closed
89.38%
7
GLM-5.3-FlashZ.AI · Open weight
88.94%
8
Muse Spark 1.1Meta · Closed
88.89%
9
GLM-5.3Z.AI · Open weight
88.81%
10
Claude Fable 5Anthropic · Closed
88.52%
11
MiMo-V2.6-ProXiaomi · Open weight
88.31%
12
GPT-5.1OpenAI · Closed
88.09%
13
Kimi K3Moonshot AI · Closed
87.96%
14
GPT-6 AstraOpenAI · Closed
87.91%
15
MiniMax M3MiniMax · Open weight
87.25%
16
Grok 4.5xAI · Closed
86.88%
17
GPT-5.5OpenAI · Closed
86.87%
18
Claude Opus 4.6Anthropic · Closed
86.74%
19
Grok 4.6xAI · Closed
86.53%
20
Claude Opus 4.6 (Adaptive)Anthropic · Closed
86.13%
21
Muse SparkMeta · Closed
85.90%
22
Claude Opus 4.8Anthropic · Closed
85.75%
23
DeepSeek V4.1 FlashDeepSeek · Open weight
85.50%
24
InklingThinking Machines Lab · Open weight
85.41%
25
Claude Opus 4.5 ThinkingAnthropic · Closed
85.32%
26
MiMo-V2.6-FlashXiaomi · Open weight
85.28%
27
GPT-5.6 SolOpenAI · Closed
85.23%
28
Claude Haiku 4.5 ThinkingAnthropic · Closed
85.23%
29
Qwen3.8 MaxAlibaba · Open weight
84.95%
30
Claude Sonnet 4.5Anthropic · Closed
84.52%
31
Gemini 3.8 FlashGoogle · Closed
84.50%
32
GPT-5.6 LunaOpenAI · Closed
84.39%
33
GPT-5.2OpenAI · Closed
84.39%
34
Inkling-SmallThinking Machines Lab · Open weight
84.11%
35
Claude Sonnet 4.5 ThinkingAnthropic · Closed
84.10%
36
Gemini 3.7 FlashGoogle · Closed
83.94%
37
Qwen3.8-27BAlibaba · Open weight
83.85%
38
MiMo-V2.5-ProXiaomi · Closed
83.73%
39
GPT-6 LunaOpenAI · Closed
83.71%
40
GPT-5OpenAI
83.65%
41
Hy4 previewTencent · Open weight
83.60%
42
GLM-5.2Z.AI · Open weight
83.53%
43
Claude Opus 4.5Anthropic · Closed
83.25%
45
Claude Opus 4.7Anthropic · Closed
82.95%
46
Gemini 2.5 FlashGoogle · Closed
82.87%
47
GPT-5.6 TerraOpenAI · Closed
82.87%
48
GPT-6 SolOpenAI · Closed
82.03%
49
81.63%
51
80.78%
52
GPT-5 miniOpenAI · Closed
80.58%
53
DeepSeek V4 Flash 0731DeepSeek · Open weight
80.36%
54
DeepSeek V4 Pro 0813DeepSeek · Open weight
80.17%
55
MiniMax M2.7MiniMax · Open weight
79.87%
57
Gemini 3.6 FlashGoogle · Closed
79.66%
58
Qwen3.7 MaxAlibaba · Closed
79.40%
59
78.73%
61
78.15%
62
Kimi K2.6Moonshot AI · Open weight
78.15%
64
GPT-5.4OpenAI · Closed
77.55%
66
77.13%
67
GPT-5.4 nanoOpenAI · Closed
77.09%
68
Qwen3.6 PlusAlibaba · Closed
76.96%
69
o3OpenAI · Closed
76.65%
70
Gemini 3.5 FlashGoogle · Closed
76.57%
71
76.44%
73
Claude Sonnet 5Anthropic · Closed
76.05%
76
DeepSeek V4 Pro 0813DeepSeek · Open weight
75.14%
77
Grok 4.3xAI · Closed
74.40%
79
Gemini 2.5 ProGoogle · Closed
73.55%
80
GPT-5 nanoOpenAI · Closed
72.86%
81
72.83%
82
Qwen3 MaxAlibaba · Closed
72.71%
83
72.41%
84
GLM-5.1Z.AI · Open weight
72.27%
85
MiMo-V2.5Xiaomi · Closed
72.15%
86
72.04%
87
71.75%
88
Gemini 3.5 Flash-LiteGoogle · Closed
70.89%
89
Qwen3.5 FlashAlibaba · Closed
70.62%
92
O4 MiniOpenAI
69.14%
93
GLM-4.7Z.AI · Open weight
68.63%
94
67.73%
96
Laguna M.1Poolside · Closed
65.91%
99
Laguna XS.2Poolside · Open weight
61.43%
100
55.68%
101
Mercury 2.5Inception · Closed
55.09%

How MedScribe is shown here

BenchLM mirrors the public Vals AI MedScribe leaderboard captured from https://www.vals.ai/benchmarks/medscribe and updated by Vals on September 26, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

MedScribe is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

104 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

The published MedScribe snapshot places Claude Opus 5.5 first at 91.43%. The third row is 0.33 points behind. The broader top-10 range is 2.91 points, so many of the published results sit in a relatively narrow band.

104 models have been evaluated on MedScribe. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. MedScribe is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MedScribe

Year

2026

Tasks

Medical administrative support tasks

Format

Accuracy score

Difficulty

Professional healthcare administration

BenchLM mirrors the public Vals MedScribe leaderboard as display-only healthcare evidence.

Freshness and provenance

Version

MedScribe 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MedScribe measure?

Vals AI healthcare benchmark for whether models can support doctors with administrative work.

Which model leads the published MedScribe snapshot?

Claude Opus 5.5 currently leads the published MedScribe snapshot with 91.43% medscribe score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MedScribe?

The September 26, 2026 snapshot contains 104 AI models.

Last updated: September 26, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.