Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Vals Legal Research Bench (Legal Research Bench)

Evaluating agents on legal research tasks across diverse areas of US law

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Legal Research Bench

BenchLM mirrors the public Vals AI Legal Research Bench leaderboard captured from https://www.vals.ai/benchmarks/legal_research and updated by Vals on September 10, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Legal Research Bench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

59 Vals rows18 task viewsprivate datasetTasks: Overall, Administrative / Regulatory, Business & Commercial, Civil Litigation, Constitutional / Civil RightsDisplay only

Legal Research Bench score on Legal Research Bench — September 10, 2026

We mirror the published legal research bench score view for Legal Research Bench. Muse Spark 1.3 Max leads the public snapshot at 55.29%, followed by Claude Opus 5 (55.29%) and Claude Fable 5.1 (55.29%). We do not use these results to rank models overall.

59 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated September 10, 2026

Legal Research Bench score table (59 models)

Score
1
Muse Spark 1.3 MaxMetamax reasoning
55.29%
2
Claude Opus 5Anthropic · Closed
55.29%
3
Claude Fable 5.1Anthropic · Closed
55.29%
4
Claude Fable 5Anthropic · Closed
49.52%
5
GLM-5.3Z.AI · Open weightmax reasoning
49.04%
6
Grok 4.6xAI · Closedhigh reasoning
48.08%
7
GPT-5.6 SolOpenAI · Closedmax reasoning
48.08%
8
Qwen3.8 MaxAlibaba · Open weight
47.60%
9
GLM-5.3-FlashZ.AI · Open weightmax reasoning
45.19%
10
Kimi K3Moonshot AI · Closed
44.23%
11
Muse Spark 1.2Meta · Closedxhigh reasoning
43.75%
12
Claude Opus 4.8Anthropic · Closed
43.75%
13
Claude Sonnet 5Anthropic · Closed
41.83%
14
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
41.35%
15
GPT-5.6 TerraOpenAI · Closedmax reasoning
41.35%
16
Muse Spark 1.3Meta · Closedxhigh reasoning
40.87%
17
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
40.87%
18
GPT-5.5OpenAI · Closedxhigh reasoning
40.38%
19
GPT-6 AstraOpenAI · Closedmax reasoning
39.42%
20
Gemini 3.8 FlashGoogle · Closedhigh reasoning
38.94%
21
Claude Opus 4.7Anthropic · Closed
38.46%
22
Claude Sonnet 4.6Anthropic · Closed
38.46%
23
Muse Spark 1.1Meta · Closedxhigh reasoning
37.98%
24
Grok 4.5xAI · Closedhigh reasoning
37.98%
25
GPT-5.6 LunaOpenAI · Closedmax reasoning
36.54%
26
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
36.06%
27
Gemini 3.7 FlashGoogle · Closedhigh reasoning
34.62%
28
GLM-5.2Z.AI · Open weightmax reasoning
31.25%
29
Gemini 3.5 FlashGoogle · Closedhigh reasoning
30.77%
30
DeepSeek V4 Flash 0731DeepSeek · Closedhigh reasoning
30.29%
31
MiniMax M3MiniMax · Open weight
29.81%
32
InklingThinking Machines Lab · Open weight0.99 reasoning
28.36%
33
GLM-5.1Z.AI · Open weight
27.89%
34
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
25.48%
35
Qwen3.7 MaxAlibaba · Closed
25.48%
36
Gemini 3.6 FlashGoogle · Closedhigh reasoning
25.00%
37
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
23.08%
38
Gemini 3.1 Pro PreviewGooglehigh reasoning
20.67%
39
Gemini 3 Flash PreviewGooglehigh reasoning
18.27%
40
Qwen3.7 PlusAlibaba · Closed
16.35%
41
MiMo-V2.5-ProXiaomi · Closed
15.87%
42
15.87%
43
Kimi K2.6Moonshot AI · Open weight
15.87%
44
Grok 4.3xAI · Closedhigh reasoning
15.38%
46
Qwen3.6 PlusAlibaba · Closed
14.90%
47
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
13.94%
49
GPT-5.4 miniOpenAI · Closedxhigh reasoning
12.50%
50
Claude Haiku 4.5 ThinkingAnthropic · Closed
10.58%
51
MiniMax M2.7MiniMax · Open weight
10.58%
52
MiMo-V2.5Xiaomi · Closed
9.13%
53
Mistral Medium 3.5Mistral AIhigh reasoning
9.13%
54
GPT-5.4 nanoOpenAI · Closedhigh reasoning
6.25%
57
Laguna M.1Poolside · Closed
2.40%
58
Laguna XS.2Poolside · Open weight
0.96%

The published Legal Research Bench snapshot places Muse Spark 1.3 Max first at 55.29%. The third row is 0.00 points behind. The broader top-10 range is 11.06 points, so the table still separates the published systems.

59 models have been evaluated on Legal Research Bench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Legal Research Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Legal Research Bench

Year

2026

Tasks

US-law legal research tasks

Format

Accuracy score

Difficulty

Professional legal research

BenchLM mirrors the public Vals AI Legal Research Bench leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

BenchLM freshness & provenance

Version

Legal Research Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Legal Research Bench measure?

Evaluating agents on legal research tasks across diverse areas of US law

Which model leads the published Legal Research Bench snapshot?

Muse Spark 1.3 Max currently leads the published Legal Research Bench snapshot with 55.29% legal research bench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Legal Research Bench?

The September 10, 2026 snapshot contains 59 AI models.

Last updated: September 10, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.