Skip to main content
BenchLM

APEX-Agents-AA

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

Benchmark score on APEX-Agents-AA — September 27, 2026

We compile the APEX-Agents-AA rows from secondary reports. Gemini 3.5 Flash leads the table at 47.1%, followed by Kimi K3 (41.3%) and GPT-5.6 Terra (38.9%). We do not use these results to rank models overall.

27 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (27 models)

Score
1
Gemini 3.5 FlashGoogle · Closed
47.1%
2
Kimi K3Moonshot AI · Closed
41.3%
3
GPT-5.6 TerraOpenAI · Closed
38.9%
4
GPT-5.5OpenAI · Closed
37.7%
5
GPT-5.6 LunaOpenAI · Closed
35.8%
6
GLM-5.2Z.AI · Open weight
33.7%
7
GPT-5.4OpenAI · Closed
33.3%
8
Claude Opus 4.6 (Adaptive)Anthropic · Closed
33.0%
9
Gemini 3.1 ProGoogle · Closed
32.0%
10
Apodex 1.1Apodex · Closed
31.2%
11
Apodex 1.1 MiniApodex · Open weight
31.2%
12
Kimi K2.6Moonshot AI · Open weight
28.5%
13
GPT-5.4 miniOpenAI · Closed
28.2%
14
GPT-5.4 nanoOpenAI · Closed
24.9%
15
DeepSeek V4 Pro 0813DeepSeek · Open weight
24.3%
16
Qwen3.7 PlusAlibaba · Closed
22.4%
17
Grok 4.3xAI · Closed
17.0%
18
Step 3.7 FlashStepFun · Open weight
14.8%
19
GLM-5Z.AI · Open weight
14.5%
20
Kimi K2.5Moonshot AI · Open weight
11.5%
21
Kimi K2.5 (Reasoning)Moonshot AI · Closed
11.5%
22
MiniMax M2.7MiniMax · Open weight
10.6%
23
GPT-OSS 120BOpenAI · Open weight
3.1%
24
MiMo-V2.5-ProXiaomi · Closed
2.4%
25
Nemotron 3 Super 100BNVIDIA · Open weight
1.8%
26
Nemotron 3 Super 120B A12BNVIDIA · Open weight
1.8%
27
GPT-OSS 20BOpenAI · Open weight
0.7%

Among the reported APEX-Agents-AA rows, Gemini 3.5 Flash is first at 47.1%. The third row is 8.2 points behind. The broader top-10 range is 15.9 points, so the table still separates the published systems.

27 models have been evaluated on APEX-Agents-AA. The benchmark falls in the Agentic category. APEX-Agents-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About APEX-Agents-AA

Year

2026

Tasks

452 professional-services agent tasks

Format

Pass@1

Difficulty

Long-horizon workplace agent tasks

BenchLM stores APEX-Agents-AA as a display-only agentic row. Artificial Analysis reports pass@1 over 452 public APEX-Agents tasks spanning investment banking, management consulting, and corporate law.

Freshness and provenance

Version

APEX-Agents-AA 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does APEX-Agents-AA measure?

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

Which model scores highest on APEX-Agents-AA?

Gemini 3.5 Flash by Google currently leads with a score of 47.1% on APEX-Agents-AA.

How many models are evaluated on APEX-Agents-AA?

27 AI models have been evaluated on APEX-Agents-AA on BenchLM.

Last updated: September 27, 2026 · BenchLM version APEX-Agents-AA 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.