APEX-Agents-AA
We show this table for reference; we do not rank on it.
Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.
Benchmark score on APEX-Agents-AA — September 27, 2026
We compile the APEX-Agents-AA rows from secondary reports. Gemini 3.5 Flash leads the table at 47.1%, followed by Kimi K3 (41.3%) and GPT-5.6 Terra (38.9%). We do not use these results to rank models overall.
Gemini 3.5 Flash
Kimi K3
Moonshot AI
GPT-5.6 Terra
OpenAI
27 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (27 models)
ScoreAmong the reported APEX-Agents-AA rows, Gemini 3.5 Flash is first at 47.1%. The third row is 8.2 points behind. The broader top-10 range is 15.9 points, so the table still separates the published systems.
27 models have been evaluated on APEX-Agents-AA. The benchmark falls in the Agentic category. APEX-Agents-AA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About APEX-Agents-AA
Year
2026
Tasks
452 professional-services agent tasks
Format
Pass@1
Difficulty
Long-horizon workplace agent tasks
BenchLM stores APEX-Agents-AA as a display-only agentic row. Artificial Analysis reports pass@1 over 452 public APEX-Agents tasks spanning investment banking, management consulting, and corporate law.
Freshness and provenance
Version
APEX-Agents-AA 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does APEX-Agents-AA measure?
Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.
Which model scores highest on APEX-Agents-AA?
Gemini 3.5 Flash by Google currently leads with a score of 47.1% on APEX-Agents-AA.
How many models are evaluated on APEX-Agents-AA?
27 AI models have been evaluated on APEX-Agents-AA on BenchLM.
Compare top models on APEX-Agents-AA
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.