Skip to main content
BenchLM
Data

Artificial Analysis AutomationBench (AA AutomationBench)

Data verified 38 confirmed releases in the last 30 daysFollow model changes

An independently evaluated automation benchmark from Artificial Analysis.

Top models on AA AutomationBench — October 7, 2026

As of October 7, 2026, Gemini 4 Argon leads the AA AutomationBench leaderboard with 77.5% , followed by Claude Sonnet 5.5 (71.8%) and Claude Opus 5.5 (69.5%).

25 modelsAgentic2% of Agentic reference weightCurrentUpdated October 7, 2026

AA AutomationBench leaderboard
RankModel / configurationScoreParameters (B)Open / closed
1Gemini 4 ArgonGoogle
77.5%
Not reportedClosed
2Claude Sonnet 5.5Anthropic
71.8%
Not reportedClosed
3Claude Opus 5.5Anthropic
69.5%
Not reportedClosed
4DeepSeek V4.1 FlashDeepSeek
68.9%
Not reportedOpen
5GPT-6 AstraOpenAI
68.5%
Not reportedClosed
6Grok 4.6xAI
66.7%
Not reportedClosed
7Grok 4.7xAI
65.6%
Not reportedClosed
8GPT-6.1 SolOpenAI
64.9%
Not reportedClosed
9GLM-5.3Z.AI
62.2%
Not reportedOpen
10GPT-6 SolOpenAI
61.6%
Not reportedClosed
11GLM-5.3-FlashZ.AI
60.4%
Not reportedOpen
12Gemini 3.8 FlashGoogle
59.9%
Not reportedClosed
13Mistral Large 4Mistral
59.9%
Not reportedPending
14Claude Fable 5.1Anthropic
59.4%
Not reportedClosed
15MiMo-V2.6-ProXiaomi
58.6%
Not reportedOpen
16Kimi K3Moonshot AI
58.3%
Not reportedPending
17Muse Spark 1.3Meta
57.9%
Not reportedClosed
18GPT-6 LunaOpenAI
53.2%
Not reportedClosed
19Step 5 PreviewStepFun
51.0%
Not reportedPending
20Qwen3.8-27BAlibaba
48.2%
Not reportedOpen
21Claude Haiku 5.5Anthropic
35.4%
Not reportedClosed
22MiniMax M3MiniMax
21.3%
Not reportedOpen
23Muse Glimmer 30BMeta
6.8%
Not reportedOpen
24InklingThinking Machines Lab
5.0%
Not reportedOpen
25Nemotron 3 UltraNVIDIA
3.0%
Not reportedOpen

According to BenchLM.ai, Gemini 4 Argon leads the AA AutomationBench benchmark with a score of 77.5%, followed by Claude Sonnet 5.5 (71.8%) and Claude Opus 5.5 (69.5%). The scores show moderate spread, with meaningful differences between the top tier and mid-tier models.

25 models have been evaluated on AA AutomationBench. The benchmark falls in the Agentic category. BenchAlign v5.8 gives AA AutomationBench 2% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About AA AutomationBench

Year

2026

Tasks

Business-process automation tasks

Format

Task success rate

Difficulty

Agentic automation

Stored as its own agentic row.

Freshness and provenance

Version

AA AutomationBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AA AutomationBench measure?

An independently evaluated automation benchmark from Artificial Analysis.

Which model scores highest on AA AutomationBench?

Gemini 4 Argon by Google currently leads with a score of 77.5% on AA AutomationBench.

How many models are evaluated on AA AutomationBench?

25 AI models have published results on AA AutomationBench in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version AA AutomationBench 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.