AutomationBench
An agent benchmark for completing automation workflows in reproducible task environments.
Top models on AutomationBench — September 18, 2026
As of September 18, 2026, DeepSeek V4.1 Flash leads the AutomationBench leaderboard with 54.8% , followed by Atria Dawn Preview (53.8%) and Muse Spark 1.3 (49.4%).
DeepSeek V4.1 Flash
DeepSeek
Atria Dawn Preview
Shanghai Artificial Intelligence Laboratory
Muse Spark 1.3
Meta
14 modelsAgentic10% of category scoreCurrentUpdated September 18, 2026
Leaderboard (14 models)
ScoreAccording to BenchLM.ai, DeepSeek V4.1 Flash leads the AutomationBench benchmark with a score of 54.8%, followed by Atria Dawn Preview (53.8%) and Muse Spark 1.3 (49.4%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
14 models have been evaluated on AutomationBench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Within that category, AutomationBench contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.
About AutomationBench
Year
2026
Tasks
600 public automation tasks
Format
Agent task-completion score
Difficulty
Long-horizon automation
Moonshot evaluates the 600-task public subset while following the official GitHub setup. BenchLM stores this provider-run result separately from Artificial Analysis AutomationBench-AA.
Freshness and provenance
Version
AutomationBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does AutomationBench measure?
An agent benchmark for completing automation workflows in reproducible task environments.
Which model scores highest on AutomationBench?
DeepSeek V4.1 Flash by DeepSeek currently leads with a score of 54.8% on AutomationBench.
How many models are evaluated on AutomationBench?
14 AI models have been evaluated on AutomationBench on BenchLM.
Compare top models on AutomationBench
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.