Skip to main content

Benchmark profile

Artificial Analysis Briefcase (AA Briefcase)

An independently evaluated professional-work benchmark reported as Elo.

Data verified

Benchmark score on AA Briefcase — July 29, 2026

BenchLM mirrors the published score view for AA Briefcase. Claude Opus 5 leads the public snapshot at 1720 , followed by Claude Fable 5 (1574) and Kimi K3 (1540). BenchLM does not use these results to rank models overall.

19 modelsAgenticCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (19 models)

Score
1
Claude Opus 5Anthropic · Closed
1720
2
Claude Fable 5Anthropic · Closed
1574
3
Kimi K3Moonshot AI · Closed
1540
4
GPT-5.6 SolOpenAI · Closed
1503
5
Claude Sonnet 5Anthropic · Closed
1386
6
Claude Opus 4.8Anthropic · Closed
1346
7
Grok 4.5xAI · Closed
1317
8
GLM-5.2Z.AI · Open weight
1254
9
MiniMax M3MiniMax · Open weight
1110
10
Gemini 3.6 FlashGoogle · Closed
964
11
DeepSeek V4 Pro (Max)DeepSeek · Open weight
930
12
Qwen3.7 MaxAlibaba · Closed
912
13
MiMo-V2.5-ProXiaomi · Closed
878
14
Nemotron 3 UltraNVIDIA · Open weight
873
15
Muse Spark 1.1Meta · Closed
868
16
InklingThinking Machines Lab · Open weight
839
17
DeepSeek V4 Flash (Max)DeepSeek · Open weight
833
18
Gemini 3.5 Flash-LiteGoogle · Closed
636
19
Mistral Medium 3.5 128BMistral · Open weight
516

The published AA Briefcase snapshot places Claude Opus 5 first at 1720. The third row is 180 score units behind. The broader top-10 range is 756 score units, so the table still separates the published systems.

19 models have been evaluated on AA Briefcase. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. AA Briefcase is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA Briefcase

Year

2026

Tasks

Professional knowledge-work tasks

Format

Elo

Difficulty

Professional work

Stored as a display-only Elo rating.

BenchLM freshness & provenance

Version

AA Briefcase 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA Briefcase measure?

An independently evaluated professional-work benchmark reported as Elo.

Which model scores highest on AA Briefcase?

Claude Opus 5 by Anthropic currently leads with a score of 1720 on AA Briefcase.

How many models are evaluated on AA Briefcase?

19 AI models have been evaluated on AA Briefcase on BenchLM.

Last updated: July 29, 2026 · BenchLM version AA Briefcase 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.