Artificial Analysis AutomationBench (AA AutomationBench)
An independently evaluated automation benchmark from Artificial Analysis.
Top models on AA AutomationBench — October 7, 2026
As of October 7, 2026, Gemini 4 Argon leads the AA AutomationBench leaderboard with 77.5% , followed by Claude Sonnet 5.5 (71.8%) and Claude Opus 5.5 (69.5%).
Gemini 4 Argon
Claude Sonnet 5.5
Anthropic
Claude Opus 5.5
Anthropic
25 modelsAgentic2% of Agentic reference weightCurrentUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Gemini 4 ArgonGoogle | 77.5% | Not reported | Closed |
| 2 | Claude Sonnet 5.5Anthropic | 71.8% | Not reported | Closed |
| 3 | Claude Opus 5.5Anthropic | 69.5% | Not reported | Closed |
| 4 | DeepSeek V4.1 FlashDeepSeek | 68.9% | Not reported | Open |
| 5 | GPT-6 AstraOpenAI | 68.5% | Not reported | Closed |
| 6 | Grok 4.6xAI | 66.7% | Not reported | Closed |
| 7 | Grok 4.7xAI | 65.6% | Not reported | Closed |
| 8 | GPT-6.1 SolOpenAI | 64.9% | Not reported | Closed |
| 9 | GLM-5.3Z.AI | 62.2% | Not reported | Open |
| 10 | GPT-6 SolOpenAI | 61.6% | Not reported | Closed |
| 11 | GLM-5.3-FlashZ.AI | 60.4% | Not reported | Open |
| 12 | Gemini 3.8 FlashGoogle | 59.9% | Not reported | Closed |
| 13 | Mistral Large 4Mistral | 59.9% | Not reported | Pending |
| 14 | Claude Fable 5.1Anthropic | 59.4% | Not reported | Closed |
| 15 | MiMo-V2.6-ProXiaomi | 58.6% | Not reported | Open |
| 16 | Kimi K3Moonshot AI | 58.3% | Not reported | Pending |
| 17 | Muse Spark 1.3Meta | 57.9% | Not reported | Closed |
| 18 | GPT-6 LunaOpenAI | 53.2% | Not reported | Closed |
| 19 | Step 5 PreviewStepFun | 51.0% | Not reported | Pending |
| 20 | Qwen3.8-27BAlibaba | 48.2% | Not reported | Open |
| 21 | Claude Haiku 5.5Anthropic | 35.4% | Not reported | Closed |
| 22 | MiniMax M3MiniMax | 21.3% | Not reported | Open |
| 23 | Muse Glimmer 30BMeta | 6.8% | Not reported | Open |
| 24 | InklingThinking Machines Lab | 5.0% | Not reported | Open |
| 25 | Nemotron 3 UltraNVIDIA | 3.0% | Not reported | Open |
According to BenchLM.ai, Gemini 4 Argon leads the AA AutomationBench benchmark with a score of 77.5%, followed by Claude Sonnet 5.5 (71.8%) and Claude Opus 5.5 (69.5%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.
25 models have been evaluated on AA AutomationBench. The benchmark falls in the Agentic category. BenchAlign v5.8 gives AA AutomationBench 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.
About AA AutomationBench
Year
2026
Tasks
Business-process automation tasks
Format
Task success rate
Difficulty
Agentic automation
Stored as its own agentic row.
Freshness and provenance
Version
AA AutomationBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does AA AutomationBench measure?
An independently evaluated automation benchmark from Artificial Analysis.
Which model scores highest on AA AutomationBench?
Gemini 4 Argon by Google currently leads with a score of 77.5% on AA AutomationBench.
How many models are evaluated on AA AutomationBench?
25 AI models have published results on AA AutomationBench in the BenchLM catalog.
Compare top models on AA AutomationBench
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.