Skip to main content
BenchLM

Artificial Analysis Harvey LAB-AA (AA Harvey LAB)

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An independently evaluated legal-agent benchmark from Artificial Analysis.

Top models on AA Harvey LAB — September 27, 2026

As of September 27, 2026, Kimi K3 leads the AA Harvey LAB leaderboard with 94.6% , followed by Claude Fable 5 (93.6%) and Claude Opus 5 (93.5%).

12 modelsAgentic1% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (12 models)

Score
1
Kimi K3Moonshot AI · Closed
94.6%
2
Claude Fable 5Anthropic · Closed
93.6%
3
Claude Opus 5Anthropic · Closed
93.5%
4
Claude Fable 5.1Anthropic · Closed
93.0%
5
Claude Opus 5.5Anthropic · Closed
91.2%
6
Gemini 3.7 FlashGoogle · Closed
90.7%
7
MiniMax M3MiniMax · Open weight
88.4%
8
GPT-5.6 LunaOpenAI · Closed
87.9%
9
GPT-5.6 SolOpenAI · Closed
87.2%
10
Nemotron 3 UltraNVIDIA · Open weight
81.7%
11
Mistral Medium 3.5 128BMistral · Open weight
69.1%
12
Grok 4.7xAI · Closed
19.6%

According to BenchLM.ai, Kimi K3 leads the AA Harvey LAB benchmark with a score of 94.6%, followed by Claude Fable 5 (93.6%) and Claude Opus 5 (93.5%). The top models are clustered within 1.1 points, suggesting this benchmark is nearing saturation for frontier models.

12 models have been evaluated on AA Harvey LAB. The benchmark falls in the Agentic category. BenchAlign v5.7 gives AA Harvey LAB 1% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About AA Harvey LAB

Year

2026

Tasks

Legal agent tasks

Format

Task success rate

Difficulty

Professional legal work

Stored as its own agentic row.

Freshness and provenance

Version

AA Harvey LAB 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA Harvey LAB measure?

An independently evaluated legal-agent benchmark from Artificial Analysis.

Which model scores highest on AA Harvey LAB?

Kimi K3 by Moonshot AI currently leads with a score of 94.6% on AA Harvey LAB.

How many models are evaluated on AA Harvey LAB?

12 AI models have been evaluated on AA Harvey LAB on BenchLM.

Last updated: September 27, 2026 · BenchLM version AA Harvey LAB 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.