Skip to main content
BenchLM

Artificial Analysis EnterpriseOps-Gym (AA EnterpriseOps-Gym)

Data verified 34 confirmed releases in the last 30 daysFollow model changes

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

Top models on AA EnterpriseOps-Gym — September 27, 2026

As of September 27, 2026, Claude Fable 5 leads the AA EnterpriseOps-Gym leaderboard with 51.1% , followed by Gemini 3.5 Flash (50.1%) and DeepSeek V4 Pro 0813 (49.6%).

17 modelsAgentic2% of Agentic reference weightCurrentUpdated September 27, 2026

Leaderboard (17 models)

Score
1
Claude Fable 5Anthropic · Closed
51.1%
2
Gemini 3.5 FlashGoogle · Closed
50.1%
3
DeepSeek V4 Pro 0813DeepSeek · Open weight
49.6%
4
Grok 4.6xAI · Closed
48.3%
5
Claude Opus 5Anthropic · Closed
47.5%
6
Kimi K3Moonshot AI · Closed
45.3%
7
Qwen3.8-27BAlibaba · Open weight
44.2%
8
GPT-5.6 SolOpenAI · Closed
42.9%
9
Gemini 3.5 Flash-LiteGoogle · Closed
42.3%
10
GPT-5.6 LunaOpenAI · Closed
40.8%
11
InklingThinking Machines Lab · Open weight
38.0%
12
GLM-5.3Z.AI · Open weight
36.4%
13
Muse Glimmer 30BMeta · Open weight
34.7%
14
Mistral Medium 3.5 128BMistral · Open weight
33.7%
15
GLM-5.3-FlashZ.AI · Open weight
33.2%
16
MiniMax M3MiniMax · Open weight
32.1%
17
Nemotron 3 UltraNVIDIA · Open weight
28.9%

According to BenchLM.ai, Claude Fable 5 leads the AA EnterpriseOps-Gym benchmark with a score of 51.1%, followed by Gemini 3.5 Flash (50.1%) and DeepSeek V4 Pro 0813 (49.6%). The top models are clustered within 1.5 points, suggesting this benchmark is nearing saturation for frontier models.

17 models have been evaluated on AA EnterpriseOps-Gym. The benchmark falls in the Agentic category. BenchAlign v5.7 gives AA EnterpriseOps-Gym 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About AA EnterpriseOps-Gym

Year

2026

Tasks

Enterprise operations workflows

Format

Task success rate

Difficulty

Enterprise agent operations

Stored as its own agentic row.

Freshness and provenance

Version

AA EnterpriseOps-Gym 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA EnterpriseOps-Gym measure?

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

Which model scores highest on AA EnterpriseOps-Gym?

Claude Fable 5 by Anthropic currently leads with a score of 51.1% on AA EnterpriseOps-Gym.

How many models are evaluated on AA EnterpriseOps-Gym?

17 AI models have been evaluated on AA EnterpriseOps-Gym on BenchLM.

Last updated: September 27, 2026 · BenchLM version AA EnterpriseOps-Gym 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.