Skip to main content
BenchLM

Vals MedCode (MedCode)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI healthcare benchmark for whether models can support the medical billing process.

MedCode score on MedCode — September 26, 2026

We mirror the published medcode score view for MedCode. Claude Opus 5 leads the public snapshot at 63.57%, followed by Gemini 3.1 Pro Preview (59.06%) and Claude Fable 5 (56.07%). We do not use these results to rank models overall.

102 modelsAgenticCurrentDisplay onlyUpdated September 26, 2026

MedCode score table (102 models)

Score
1
Claude Opus 5Anthropic · Closed
63.57%
3
Claude Fable 5Anthropic · Closed
56.07%
5
Gemini 3.5 FlashGoogle · Closed
55.83%
6
Claude Opus 4.7Anthropic · Closed
54.86%
7
Claude Fable 5.1Anthropic · Closed
53.51%
8
Gemini 3.7 FlashGoogle · Closed
53.39%
9
Claude Opus 4.8Anthropic · Closed
53.22%
10
Gemini 3.6 FlashGoogle · Closed
53.15%
11
Claude Sonnet 5.5Anthropic · Closed
52.92%
12
GPT-5.1OpenAI · Closed
52.73%
13
52.20%
14
Muse SparkMeta · Closed
51.31%
15
Gemini 2.5 ProGoogle · Closed
50.59%
16
Claude Opus 5.5Anthropic · Closed
49.80%
17
GPT-5.2OpenAI · Closed
49.75%
18
GPT-5OpenAI
49.63%
19
Grok 4.7xAI · Closed
49.55%
20
Muse Spark 1.2Meta · Closed
49.35%
21
Claude Opus 4.5 ThinkingAnthropic · Closed
49.16%
22
Claude Opus 4.6 (Adaptive)Anthropic · Closed
49.13%
23
GPT-5.5OpenAI · Closed
49.10%
24
Kimi K3Moonshot AI · Closed
48.88%
25
GPT-6 AstraOpenAI · Closed
48.49%
26
Claude Opus 4.6Anthropic · Closed
48.24%
27
Gemini 3.8 FlashGoogle · Closed
48.13%
29
Claude Sonnet 5Anthropic · Closed
47.54%
30
o3OpenAI · Closed
47.29%
32
GPT-6 SolOpenAI · Closed
47.07%
33
MiniMax M3MiniMax · Open weight
46.29%
34
Claude Opus 4.5Anthropic · Closed
45.17%
35
MiMo-V2.6-ProXiaomi · Open weight
44.97%
36
Grok 4.6xAI · Closed
44.71%
37
GPT-6 LunaOpenAI · Closed
44.69%
38
Claude Sonnet 4.5 ThinkingAnthropic · Closed
44.13%
39
GPT-5.6 SolOpenAI · Closed
43.97%
40
Gemini 3.5 Flash-LiteGoogle · Closed
43.49%
41
GPT-5.6 TerraOpenAI · Closed
43.41%
42
Grok 4.5xAI · Closed
43.29%
43
Hy4 previewTencent · Open weight
43.25%
44
GPT-5 miniOpenAI · Closed
43.05%
45
GLM-5.3Z.AI · Open weight
42.86%
46
DeepSeek V4 Pro 0813DeepSeek · Open weight
42.47%
47
GPT-5.6 LunaOpenAI · Closed
42.39%
48
GLM-5.1Z.AI · Open weight
41.60%
49
DeepSeek V4 Flash 0731DeepSeek · Open weight
41.41%
50
41.37%
51
GPT-5.4OpenAI · Closed
41.29%
52
InklingThinking Machines Lab · Open weight
41.19%
53
DeepSeek V4.1 FlashDeepSeek · Open weight
41.17%
54
MiMo-V2.6-FlashXiaomi · Open weight
41.06%
55
GPT-5.4 nanoOpenAI · Closed
41.03%
56
GLM-5.2Z.AI · Open weight
40.77%
57
Qwen3.8 MaxAlibaba · Open weight
40.67%
58
Claude Sonnet 4.5Anthropic · Closed
40.57%
60
DeepSeek V4 Pro 0813DeepSeek · Open weight
40.45%
63
Kimi K2.6Moonshot AI · Open weight
40.14%
64
39.32%
65
Qwen3.7 MaxAlibaba · Closed
38.75%
67
Gemini 2.5 FlashGoogle · Closed
38.42%
68
38.08%
69
Grok 4.3xAI · Closed
38.07%
70
Inkling-SmallThinking Machines Lab · Open weight
37.89%
71
37.38%
72
Qwen3.6 PlusAlibaba · Closed
36.89%
75
MiniMax M2.7MiniMax · Open weight
34.44%
77
34.08%
78
33.94%
79
O4 MiniOpenAI
33.79%
80
33.75%
81
Qwen3.5 FlashAlibaba · Closed
33.00%
82
GLM-4.7Z.AI · Open weight
32.77%
83
Claude Haiku 4.5 ThinkingAnthropic · Closed
32.68%
84
MiMo-V2.5-ProXiaomi · Closed
32.48%
87
MiMo-V2.5Xiaomi · Closed
31.89%
88
31.65%
89
Qwen3 MaxAlibaba · Closed
31.37%
90
Mercury 2.5Inception · Closed
31.33%
91
GPT-5 nanoOpenAI · Closed
30.44%
94
Qwen3.8-27BAlibaba · Open weight
28.70%
96
28.08%
97
27.11%
100
Laguna M.1Poolside · Closed
23.11%
101
Laguna XS.2Poolside · Open weight
21.25%
102
19.72%

How MedCode is shown here

BenchLM mirrors the public Vals AI MedCode leaderboard captured from https://www.vals.ai/benchmarks/medcode and updated by Vals on September 26, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

MedCode is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

102 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

The published MedCode snapshot places Claude Opus 5 first at 63.57%. The third row is 7.50 points behind. The broader top-10 range is 10.42 points, so the table still separates the published systems.

102 models have been evaluated on MedCode. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. MedCode is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MedCode

Year

2026

Tasks

Medical billing support tasks

Format

Accuracy score

Difficulty

Professional healthcare administration

BenchLM mirrors the public Vals MedCode leaderboard as display-only healthcare evidence.

Freshness and provenance

Version

MedCode 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MedCode measure?

Vals AI healthcare benchmark for whether models can support the medical billing process.

Which model leads the published MedCode snapshot?

Claude Opus 5 currently leads the published MedCode snapshot with 63.57% medcode score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MedCode?

The September 26, 2026 snapshot contains 102 AI models.

Last updated: September 26, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.