Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Vals LegalBench (LegalBench)

We mirror this table; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

LegalBench score on LegalBench — September 16, 2026

We mirror the published legalbench score view for LegalBench. Claude Fable 5 leads the public snapshot at 88.56%, followed by Claude Fable 5.1 (88.51%) and Gemini 3.1 Pro Preview (87.40%). We do not use these results to rank models overall.

145 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated September 16, 2026

LegalBench score table (145 models)

Score
1
Claude Fable 5Anthropic · Closed
88.56%
2
Claude Fable 5.1Anthropic · Closed
88.51%
3
Gemini 3.1 Pro PreviewGooglehigh reasoning
87.40%
4
Gemini 3.7 FlashGoogle · Closedhigh reasoning
87.26%
5
Gemini 3 Pro PreviewGooglehigh reasoning
87.03%
6
Gemini 3.8 FlashGoogle · Closedhigh reasoning
86.99%
7
Claude Opus 5Anthropic · Closed
86.97%
8
GPT-5.6 SolOpenAI · Closedmax reasoning
86.97%
9
Gemini 3 Flash PreviewGooglehigh reasoning
86.86%
10
Gemini 3.6 FlashGoogle · Closedhigh reasoning
86.70%
11
GPT-5.5OpenAI · Closedxhigh reasoning
86.52%
12
Grok 4.6xAI · Closedhigh reasoning
86.31%
13
GPT-5.4OpenAI · Closedxhigh reasoning
86.04%
14
GPT-5OpenAIhigh reasoning
86.02%
15
Kimi K3Moonshot AI · Closed
86.02%
16
Grok 4.5xAI · Closedhigh reasoning
85.97%
17
GPT-5.1OpenAI · Closedhigh reasoning
85.68%
18
MiniMax M3MiniMax · Open weight
85.42%
19
Claude Opus 4.6 (Adaptive)Anthropic · Closed
85.30%
20
Muse Spark 1.2Meta · Closedxhigh reasoning
85.26%
21
Claude Opus 4.7Anthropic · Closed
85.25%
22
GPT-5.6 TerraOpenAI · Closedmax reasoning
85.11%
23
85.10%
24
Muse Spark 1.1Meta · Closedxhigh reasoning
84.98%
25
Qwen3.7 MaxAlibaba · Closed
84.91%
26
GLM-5.3Z.AI · Open weightmax reasoning
84.84%
27
Kimi K2.6Moonshot AI · Open weight
84.74%
28
Claude Opus 4.5 ThinkingAnthropic · Closed
84.60%
29
Grok 4.3xAI · Closed
84.46%
30
GLM-5.1Z.AI · Open weight
84.39%
32
Qwen3.5 FlashAlibaba · Closed
84.28%
33
Qwen3.6 PlusAlibaba · Closed
84.23%
34
Muse SparkMeta · Closed
84.22%
35
Claude Sonnet 4.5 ThinkingAnthropic · Closed
84.08%
36
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
84.07%
37
GLM-5.2Z.AI · Open weight
84.07%
38
84.06%
39
GPT-5.6 LunaOpenAI · Closedmax reasoning
84.03%
40
MiniMax M2.7MiniMax · Open weight
83.98%
41
GLM-5.3-FlashZ.AI · Open weightmax reasoning
83.93%
42
Claude Sonnet 5Anthropic · Closed
83.92%
44
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
83.76%
45
o3OpenAI · Closedhigh reasoning
83.76%
46
Hy4 previewTencent · Open weight
83.76%
47
Qwen3.8 MaxAlibaba · Open weight
83.61%
48
Gemini 3.5 FlashGoogle · Closedhigh reasoning
83.60%
49
Claude Opus 4.8Anthropic · Closed
83.57%
50
83.46%
51
GLM-4.7Z.AI · Open weight
83.36%
52
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
83.28%
53
83.19%
54
83.14%
55
GPT-4.1OpenAI · Closedhigh reasoning
83.10%
56
Claude Opus 4Anthropic
83.07%
57
Mercury 2.5Inception · Closedhigh reasoning
83.06%
58
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
82.99%
59
82.95%
60
InklingThinking Machines Lab · Open weight0.99 reasoning
82.95%
61
Claude Opus 4.5Anthropic · Closed
82.84%
62
GPT-5.2OpenAI · Closedxhigh reasoning
82.76%
64
82.59%
67
82.45%
68
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
82.43%
69
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
82.36%
70
GPT-4oOpenAI · Closedhigh reasoning
82.21%
71
Claude Sonnet 4.6Anthropic · Closed
82.12%
75
81.92%
76
Qwen3 MaxAlibaba · Closed
81.86%
77
GPT-5 miniOpenAI · Closedhigh reasoning
81.77%
78
81.45%
79
Claude Haiku 4.5 ThinkingAnthropic · Closed
81.24%
80
DeepSeek V3DeepSeek · Open weight
80.76%
81
80.60%
82
GPT-4 TurboOpenAI · Closedhigh reasoning
80.46%
83
o1OpenAI · Closedhigh reasoning
80.39%
84
80.33%
85
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
80.32%
86
Kimi K2 ThinkingMoonshot AI
80.20%
87
Qwen3 235b A22bFireworks AI
80.18%
88
GPT-4oOpenAI · Closedhigh reasoning
80.12%
89
80.00%
90
MiniMax M2.5MiniMax · Closed
79.96%
91
79.70%
93
GLM-4.6Z.AI · Open weight
79.61%
95
O4 MiniOpenAIhigh reasoning
79.19%
96
79.14%
99
MiMo-V2.5Xiaomi · Closed
78.89%
101
78.36%
102
GPT-4.1 miniOpenAI · Closedhigh reasoning
78.04%
103
GPT-5.4 nanoOpenAI · Closedhigh reasoning
77.92%
104
77.81%
106
DeepSeek V3 0324Fireworks AI
77.73%
107
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
77.71%
109
MiMo-V2.5-ProXiaomi · Closed
77.13%
110
76.08%
111
GPT-OSS 120BOpenAI · Open weight
75.94%
112
GLM-4.5Z.AI · Closed
75.63%
113
75.45%
114
Laguna M.1Poolside · Closed
75.14%
115
74.16%
117
71.96%
118
o3-miniOpenAI · Closedhigh reasoning
71.54%
119
Laguna XS.2Poolside · Open weight
71.03%
120
GPT-OSS 20BOpenAI · Open weight
70.85%
121
70.33%
122
70.25%
123
69.73%
124
69.56%
125
69.16%
126
69.08%
127
68.29%
128
68.14%
129
67.98%
130
DeepSeek-R1DeepSeek · Open weight
67.32%
131
66.62%
132
DeepSeek V3p2Fireworks AI
65.96%
133
GPT-3.5 TurboOpenAIhigh reasoning
64.37%
134
63.24%
135
62.43%
136
61.98%
137
GPT-4.1 nanoOpenAI · Closedhigh reasoning
61.06%
138
55.84%
139
54.77%
140
53.77%
141
53.70%
142
51.87%
143
GPT-5 nanoOpenAI · Closedhigh reasoning
50.13%
144
40.02%
145
Command RCohere
32.97%

How LegalBench is shown here

BenchLM mirrors the public Vals AI LegalBench leaderboard captured from https://www.vals.ai/benchmarks/legal_bench and updated by Vals on September 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

LegalBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

145 Vals rows6 task viewspublic datasetTasks: Overall, Issue Tasks, Rule Tasks, Conclusion Tasks, Interpretation TasksDisplay only

The published LegalBench snapshot places Claude Fable 5 first at 88.56%. The third row is 1.16 points behind. The broader top-10 range is 1.86 points, so many of the published results sit in a relatively narrow band.

145 models have been evaluated on LegalBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. LegalBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LegalBench

Year

2026

Tasks

Legal reasoning task views

Format

Accuracy score

Difficulty

Professional legal reasoning

BenchLM mirrors Vals LegalBench as display-only legal-domain evidence and does not use it in weighted rankings.

Freshness and provenance

Version

LegalBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does LegalBench measure?

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

Which model leads the published LegalBench snapshot?

Claude Fable 5 currently leads the published LegalBench snapshot with 88.56% legalbench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LegalBench?

The September 16, 2026 snapshot contains 145 AI models.

Last updated: September 16, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.