Skip to main content
BenchLM

Vals MedQA (MedQA)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Evaluating language model bias in medical questions.

MedQA score on MedQA — April 16, 2026

We mirror the published medqa score view for MedQA. o1 leads the public snapshot at 96.52%, followed by GPT-5.1 (96.38%) and Gemini 3.1 Pro Preview (96.37%). We do not use these results to rank models overall.

95 modelsKnowledgeCurrentDisplay onlyUpdated April 16, 2026

MedQA score table (95 models)

Score
1
o1OpenAI · Closedhigh reasoning
96.52%
2
GPT-5.1OpenAI · Closedhigh reasoning
96.38%
3
Gemini 3.1 Pro PreviewGooglehigh reasoning
96.37%
4
GPT-5OpenAIhigh reasoning
96.32%
5
GPT-5.4OpenAI · Closedxhigh reasoning
96.09%
6
o3OpenAI · Closedhigh reasoning
96.06%
7
GPT-5 miniOpenAI · Closedhigh reasoning
96.06%
8
Gemini 3 Pro PreviewGooglehigh reasoning
96.03%
9
O4 MiniOpenAIhigh reasoning
96.02%
10
Claude Opus 4.5 ThinkingAnthropic · Closed
95.88%
11
Gemini 3 Flash PreviewGooglehigh reasoning
95.81%
12
Claude Opus 4.6 (Adaptive)Anthropic · Closed
95.41%
13
95.21%
14
o3-miniOpenAI · Closedhigh reasoning
94.83%
15
Claude Sonnet 4.5 ThinkingAnthropic · Closed
94.71%
17
94.37%
18
94.27%
19
GPT-5.2OpenAI · Closedxhigh reasoning
94.13%
20
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
93.92%
21
GLM-4.7Z.AI · Open weight
93.74%
23
GPT-5 nanoOpenAI · Closedhigh reasoning
93.26%
24
Claude Opus 4.5Anthropic · Closed
93.16%
26
o1-previewOpenAI · Closedhigh reasoning
93.01%
27
Claude Opus 4Anthropic
92.87%
29
Kimi K2 ThinkingMoonshot AI
92.59%
30
92.53%
31
MiniMax M2.5MiniMax · Closed
92.53%
32
92.49%
33
92.32%
34
GLM-4.6Z.AI · Open weight
92.22%
35
92.08%
36
92.07%
37
Claude Sonnet 4.6Anthropic · Closed
92.06%
39
GPT-OSS 120BOpenAI · Open weight
91.36%
40
GPT-4.1OpenAI · Closedhigh reasoning
91.18%
42
91.16%
44
DeepSeek-R1DeepSeek · Open weight
90.80%
45
Qwen3 235b A22bFireworks AI
90.62%
46
90.35%
47
O1 MiniOpenAIhigh reasoning
90.22%
49
90.10%
50
GLM-4.5Z.AI · Closed
89.97%
51
89.47%
52
DeepSeek V3p2Fireworks AInone reasoning
89.45%
54
88.65%
56
GPT-4oOpenAI · Closedhigh reasoning
88.16%
57
87.38%
58
Qwen3 MaxAlibaba · Closed
87.37%
61
GPT-4.1 miniOpenAI · Closedhigh reasoning
84.63%
62
83.97%
63
83.85%
64
Claude 3.5 SonnetAnthropic · Closed
83.19%
65
GPT-OSS 20BOpenAI · Open weight
82.88%
66
82.36%
67
82.23%
68
DeepSeek V3 0324Fireworks AI
82.00%
69
GPT-4 TurboOpenAI · Closedhigh reasoning
81.99%
70
81.47%
71
DeepSeek V3DeepSeek · Open weight
80.90%
72
80.55%
74
Claude Haiku 4.5 ThinkingAnthropic · Closed
79.57%
75
78.23%
77
76.53%
78
76.22%
81
GPT-4o miniOpenAI · Closedhigh reasoning
72.44%
82
69.10%
83
GPT-4.1 nanoOpenAI · Closedhigh reasoning
68.22%
84
68.11%
86
Mixtral 8x22B Instruct v0.1Mistral · Open weight
62.14%
87
GPT-3.5 TurboOpenAIhigh reasoning
58.47%
88
56.98%
89
55.18%
90
53.22%
91
52.52%
93
50.70%
94
43.30%
95
2.65%

How MedQA is shown here

BenchLM mirrors the public Vals AI MedQA leaderboard captured from https://www.vals.ai/benchmarks/medqa and updated by Vals on April 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

MedQA is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

95 Vals rows7 task viewspublic datasetTasks: Overall, Unbiased, Hispanic, Black, AsianDisplay only

The published MedQA snapshot places o1 first at 96.52%. The third row is 0.15 points behind. The broader top-10 range is 0.64 points, so many of the published results sit in a relatively narrow band.

95 models have been evaluated on MedQA. The benchmark falls in the Knowledge category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. MedQA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MedQA

Year

2026

Tasks

Medical question answering

Format

Accuracy score

Difficulty

Medical knowledge and bias evaluation

BenchLM mirrors the public Vals AI MedQA leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Freshness and provenance

Version

MedQA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MedQA measure?

Evaluating language model bias in medical questions.

Which model leads the published MedQA snapshot?

o1 currently leads the published MedQA snapshot with 96.52% medqa score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on MedQA?

The April 16, 2026 snapshot contains 95 AI models.

Last updated: April 16, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.