# Artificial Analysis EnterpriseOps-Gym (AA EnterpriseOps-Gym)

> An independently evaluated enterprise-operations benchmark from Artificial Analysis.

Canonical page: https://benchlm.ai/benchmarks/aaenterpriseopsgym

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About AA EnterpriseOps-Gym

- Year: 2026
- Tasks: Enterprise operations workflows
- Format: Task success rate
- Difficulty: Enterprise agent operations
- Paper: [Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa)

Stored as its own agentic row.

BenchAlign v5.7 gives AA EnterpriseOps-Gym 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Leaderboard (17 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | Anthropic | 51.1% |
| 2 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 50.1% |
| 3 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 49.6% |
| 4 | [Grok 4.6](/models/grok-4-6) | xAI | 48.3% |
| 5 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 47.5% |
| 6 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 45.3% |
| 7 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 44.2% |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 42.9% |
| 9 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 42.3% |
| 10 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 40.8% |
| 11 | [Inkling](/models/inkling) | Thinking Machines Lab | 38.0% |
| 12 | [GLM-5.3](/models/glm-5-3) | Z.AI | 36.4% |
| 13 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 34.7% |
| 14 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 33.7% |
| 15 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 33.2% |
| 16 | [MiniMax M3](/models/minimax-m3) | MiniMax | 32.1% |
| 17 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 28.9% |

## FAQ

### What does AA EnterpriseOps-Gym measure?

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

### Which model scores highest on AA EnterpriseOps-Gym?

Claude Fable 5 by Anthropic currently leads with a score of 51.1% on AA EnterpriseOps-Gym.

### How many models are evaluated on AA EnterpriseOps-Gym?

17 AI models have been evaluated on AA EnterpriseOps-Gym on BenchLM.

### Does AA EnterpriseOps-Gym affect BenchLM's overall score?

Yes. BenchAlign v5.7 gives AA EnterpriseOps-Gym 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

## Compare Top Models on AA EnterpriseOps-Gym

- [Claude Fable 5 vs Gemini 3.5 Flash](/compare/claude-fable-vs-gemini-3-5-flash)
- [Gemini 3.5 Flash vs DeepSeek V4 Pro 0813](/compare/deepseek-v4-pro-0813-vs-gemini-3-5-flash)
- [DeepSeek V4 Pro 0813 vs Grok 4.6](/compare/deepseek-v4-pro-0813-vs-grok-4-6)
- [Grok 4.6 vs Claude Opus 5](/compare/claude-opus-5-vs-grok-4-6)
