Skip to main content
BenchLM

Vals CaseLaw v2 (CaseLaw v2)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI private question-answer benchmark over Canadian court cases.

CaseLaw v2 score on CaseLaw v2 — May 4, 2026

We mirror the published caselaw v2 score view for CaseLaw v2. Grok 4.3 leads the public snapshot at 79.31%, followed by GPT-5.1 (73.42%) and GPT-4.1 (69.88%). We do not use these results to rank models overall.

54 modelsKnowledgeCurrentDisplay onlyUpdated May 4, 2026

CaseLaw v2 score table (54 models)

Score
1
Grok 4.3xAI · Closed
79.31%
2
GPT-5.1OpenAI · Closed
73.42%
3
GPT-4.1OpenAI · Closed
69.88%
4
GPT-5 miniOpenAI · Closed
68.49%
5
Claude Opus 4.7Anthropic · Closed
68.38%
6
GPT-5OpenAI
66.45%
7
GPT-5.5OpenAI · Closed
66.24%
8
GPT-5.2OpenAI · Closed
66.02%
9
65.81%
10
65.70%
11
Kimi K2 ThinkingMoonshot AI
65.70%
13
64.52%
14
Claude Sonnet 4.6Anthropic · Closed
63.99%
15
Gemini 2.5 ProGoogle · Closed
63.88%
16
GPT-5.4OpenAI · Closed
63.77%
17
Muse SparkMeta · Closed
63.13%
18
Claude Opus 4.5 ThinkingAnthropic · Closed
62.59%
19
Claude Sonnet 4.5 ThinkingAnthropic · Closed
62.16%
20
Claude Opus 4.6 (Adaptive)Anthropic · Closed
62.06%
21
61.41%
22
Kimi K2.6Moonshot AI · Open weight
61.20%
23
MiniMax M2.7MiniMax · Open weight
60.88%
24
60.45%
25
GPT-4oOpenAI · Closed
59.70%
26
59.70%
27
DeepSeek V4 Pro 0813DeepSeek · Open weight
59.38%
28
58.73%
29
Trinity-Large-ThinkingArcee AI · Open weight
57.88%
30
Claude Haiku 4.5 ThinkingAnthropic · Closed
56.48%
31
Qwen3.5 FlashAlibaba · Closed
55.95%
32
55.84%
34
55.41%
36
Qwen3 MaxAlibaba · Closed
54.98%
37
GLM-4.7Z.AI · Open weight
54.88%
39
DeepSeek V3p1Fireworks
53.91%
40
MiniMax M2.5MiniMax · Closed
53.48%
41
Qwen3.6-27BAlibaba · Open weight
53.16%
42
53.05%
43
52.63%
44
GPT-5 nanoOpenAI · Closed
52.63%
45
52.52%
46
GPT-5.4 nanoOpenAI · Closed
51.88%
47
GPT-5.4 miniOpenAI · Closed
51.66%
48
GLM-5.1Z.AI · Open weight
51.55%
49
Qwen3.6 PlusAlibaba · Closed
51.45%
50
GPT-OSS 120BOpenAI · Open weight
48.77%
51
Qwen 3.6 Max (preview)Alibaba · Closed
47.91%
52
Qwen3 MaxAlibaba · Closed
47.48%
53
44.16%
54
GPT-OSS 20BOpenAI · Open weight
43.84%

How CaseLaw v2 is shown here

BenchLM mirrors the public Vals AI CaseLaw v2 leaderboard captured from https://www.vals.ai/benchmarks/case_law_v2 and updated by Vals on May 4, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

CaseLaw v2 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

54 Vals rows1 task viewsprivate datasetTasks: OverallDisplay only

The published CaseLaw v2 snapshot places Grok 4.3 first at 79.31%. The third row is 9.43 points behind. The broader top-10 range is 13.61 points, so the table still separates the published systems.

54 models have been evaluated on CaseLaw v2. The benchmark falls in the Knowledge category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. CaseLaw v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CaseLaw v2

Year

2026

Tasks

Canadian case-law question answering

Format

Accuracy score

Difficulty

Professional legal retrieval and reasoning

Vals marks CaseLaw v2 as archived. BenchLM mirrors the public leaderboard as display-only historical legal-domain context.

Freshness and provenance

Version

CaseLaw v2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CaseLaw v2 measure?

Vals AI private question-answer benchmark over Canadian court cases.

Which model leads the published CaseLaw v2 snapshot?

Grok 4.3 currently leads the published CaseLaw v2 snapshot with 79.31% caselaw v2 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on CaseLaw v2?

The May 4, 2026 snapshot contains 54 AI models.

Last updated: May 4, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.