Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Vals LegalBench (LegalBench)

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

Data verified 32 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows LegalBench

BenchLM mirrors the public Vals AI LegalBench leaderboard captured from https://www.vals.ai/benchmarks/legal_bench and updated by Vals on September 10, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

LegalBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

143 Vals rows6 task viewspublic datasetTasks: Overall, Issue Tasks, Rule Tasks, Conclusion Tasks, Interpretation TasksDisplay only

LegalBench score on LegalBench — September 10, 2026

We mirror the published legalbench score view for LegalBench. Claude Fable 5 leads the public snapshot at 88.56%, followed by Claude Fable 5.1 (88.51%) and Gemini 3.1 Pro Preview (87.40%). We do not use these results to rank models overall.

143 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated September 10, 2026

LegalBench score table (143 models)

Score
1
Claude Fable 5Anthropic · Closed
88.56%
2
Claude Fable 5.1Anthropic · Closed
88.51%
3
Gemini 3.1 Pro PreviewGooglehigh reasoning
87.40%
4
Gemini 3.7 FlashGoogle · Closedhigh reasoning
87.26%
5
Gemini 3 Pro PreviewGooglehigh reasoning
87.03%
6
Gemini 3.8 FlashGoogle · Closedhigh reasoning
86.99%
7
Claude Opus 5Anthropic · Closed
86.97%
8
GPT-5.6 SolOpenAI · Closedmax reasoning
86.97%
9
Gemini 3 Flash PreviewGooglehigh reasoning
86.86%
10
Gemini 3.6 FlashGoogle · Closedhigh reasoning
86.70%
11
GPT-5.5OpenAI · Closedxhigh reasoning
86.52%
12
Grok 4.6xAI · Closedhigh reasoning
86.31%
13
GPT-5.4OpenAI · Closedxhigh reasoning
86.04%
14
GPT-5OpenAIhigh reasoning
86.02%
15
Kimi K3Moonshot AI · Closed
86.02%
16
Grok 4.5xAI · Closedhigh reasoning
85.97%
17
GPT-5.1OpenAI · Closedhigh reasoning
85.68%
18
MiniMax M3MiniMax · Open weight
85.42%
19
Claude Opus 4.6 (Adaptive)Anthropic · Closed
85.30%
20
Muse Spark 1.2Meta · Closedxhigh reasoning
85.26%
21
Claude Opus 4.7Anthropic · Closed
85.25%
22
GPT-5.6 TerraOpenAI · Closedmax reasoning
85.11%
23
85.10%
24
Muse Spark 1.1Meta · Closedxhigh reasoning
84.98%
25
Qwen3.7 MaxAlibaba · Closed
84.91%
26
GLM-5.3Z.AI · Open weightmax reasoning
84.84%
27
Kimi K2.6Moonshot AI · Open weight
84.74%
28
Claude Opus 4.5 ThinkingAnthropic · Closed
84.60%
29
Grok 4.3xAI · Closed
84.46%
30
GLM-5.1Z.AI · Open weight
84.39%
32
Qwen3.5 FlashAlibaba · Closed
84.28%
33
Qwen3.6 PlusAlibaba · Closed
84.23%
34
Muse SparkMeta · Closed
84.22%
35
Claude Sonnet 4.5 ThinkingAnthropic · Closed
84.08%
36
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
84.07%
37
GLM-5.2Z.AI · Open weight
84.07%
38
84.06%
39
GPT-5.6 LunaOpenAI · Closedmax reasoning
84.03%
40
MiniMax M2.7MiniMax · Open weight
83.98%
41
GLM-5.3-FlashZ.AI · Open weightmax reasoning
83.93%
42
Claude Sonnet 5Anthropic · Closed
83.92%
44
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
83.76%
45
o3OpenAI · Closedhigh reasoning
83.76%
46
Qwen3.8 MaxAlibaba · Open weight
83.61%
47
Gemini 3.5 FlashGoogle · Closedhigh reasoning
83.60%
48
Claude Opus 4.8Anthropic · Closed
83.57%
49
83.46%
50
GLM-4.7Z.AI · Open weight
83.36%
51
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
83.28%
52
83.19%
53
83.14%
54
GPT-4.1OpenAI · Closedhigh reasoning
83.10%
55
Claude Opus 4Anthropic
83.07%
56
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
82.99%
57
82.95%
58
InklingThinking Machines Lab · Open weight0.99 reasoning
82.95%
59
Claude Opus 4.5Anthropic · Closed
82.84%
60
GPT-5.2OpenAI · Closedxhigh reasoning
82.76%
62
82.59%
65
82.45%
66
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
82.43%
67
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
82.36%
68
GPT-4oOpenAI · Closedhigh reasoning
82.21%
69
Claude Sonnet 4.6Anthropic · Closed
82.12%
73
81.92%
74
Qwen3 MaxAlibaba · Closed
81.86%
75
GPT-5 miniOpenAI · Closedhigh reasoning
81.77%
76
81.45%
77
Claude Haiku 4.5 ThinkingAnthropic · Closed
81.24%
78
DeepSeek V3DeepSeek · Open weight
80.76%
79
80.60%
80
GPT-4 TurboOpenAI · Closedhigh reasoning
80.46%
81
o1OpenAI · Closedhigh reasoning
80.39%
82
80.33%
83
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
80.32%
84
Kimi K2 ThinkingMoonshot AI
80.20%
85
Qwen3 235b A22bFireworks AI
80.18%
86
GPT-4oOpenAI · Closedhigh reasoning
80.12%
87
80.00%
88
MiniMax M2.5MiniMax · Closed
79.96%
89
79.70%
91
GLM-4.6Z.AI · Open weight
79.61%
93
O4 MiniOpenAIhigh reasoning
79.19%
94
79.14%
97
MiMo-V2.5Xiaomi · Closed
78.89%
99
78.36%
100
GPT-4.1 miniOpenAI · Closedhigh reasoning
78.04%
101
GPT-5.4 nanoOpenAI · Closedhigh reasoning
77.92%
102
77.81%
104
DeepSeek V3 0324Fireworks AI
77.73%
105
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
77.71%
107
MiMo-V2.5-ProXiaomi · Closed
77.13%
108
76.08%
109
GPT-OSS 120BOpenAI · Open weight
75.94%
110
GLM-4.5Z.AI · Closed
75.63%
111
75.45%
112
Laguna M.1Poolside · Closed
75.14%
113
74.16%
115
71.96%
116
o3-miniOpenAI · Closedhigh reasoning
71.54%
117
Laguna XS.2Poolside · Open weight
71.03%
118
GPT-OSS 20BOpenAI · Open weight
70.85%
119
70.33%
120
70.25%
121
69.73%
122
69.56%
123
69.16%
124
69.08%
125
68.29%
126
68.14%
127
67.98%
128
DeepSeek-R1DeepSeek · Open weight
67.32%
129
66.62%
130
DeepSeek V3p2Fireworks AI
65.96%
131
GPT-3.5 TurboOpenAIhigh reasoning
64.37%
132
63.24%
133
62.43%
134
61.98%
135
GPT-4.1 nanoOpenAI · Closedhigh reasoning
61.06%
136
55.84%
137
54.77%
138
53.77%
139
53.70%
140
51.87%
141
GPT-5 nanoOpenAI · Closedhigh reasoning
50.13%
142
40.02%
143
Command RCohere
32.97%

The published LegalBench snapshot places Claude Fable 5 first at 88.56%. The third row is 1.16 points behind. The broader top-10 range is 1.86 points, so many of the published results sit in a relatively narrow band.

143 models have been evaluated on LegalBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. LegalBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About LegalBench

Year

2026

Tasks

Legal reasoning task views

Format

Accuracy score

Difficulty

Professional legal reasoning

BenchLM mirrors Vals LegalBench as display-only legal-domain evidence and does not use it in weighted rankings.

BenchLM freshness & provenance

Version

LegalBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does LegalBench measure?

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

Which model leads the published LegalBench snapshot?

Claude Fable 5 currently leads the published LegalBench snapshot with 88.56% legalbench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on LegalBench?

The September 10, 2026 snapshot contains 143 AI models.

Last updated: September 10, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.