# AutomationBench

> An agent benchmark for completing automation workflows in reproducible task environments.

Canonical page: https://benchlm.ai/benchmarks/automationbench

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026

## About AutomationBench

- Year: 2026
- Tasks: 600 public automation tasks
- Format: Agent task-completion score
- Difficulty: Long-horizon automation
- Paper: [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3)

Moonshot evaluates the 600-task public subset while following the official GitHub setup. BenchLM stores this provider-run result separately from Artificial Analysis AutomationBench-AA.

AutomationBench is currently weighted in BenchLM's scoring formula. The Agentic category carries 22% of the overall score, and AutomationBench contributes 10% of that category score.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 54.8% |
| 2 | [Atria Dawn Preview](/models/atria-dawn-preview) | Shanghai Artificial Intelligence Laboratory | 53.8% |
| 3 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 49.4% |
| 4 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 48.8% |
| 5 | [GLM-5.3](/models/glm-5-3) | Z.AI | 48.2% |
| 6 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 41.4% |
| 7 | [Hy4 preview](/models/hy4-preview) | Tencent | 32.1% |
| 8 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 31.8% |
| 9 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 31.4% |
| 10 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 30.8% |
| 11 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 30.4% |
| 12 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 27.3% |
| 13 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 26.0% |
| 14 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 25.1% |

## FAQ

### What does AutomationBench measure?

An agent benchmark for completing automation workflows in reproducible task environments.

### Which model scores highest on AutomationBench?

DeepSeek V4.1 Flash by DeepSeek currently leads with a score of 54.8% on AutomationBench.

### How many models are evaluated on AutomationBench?

14 AI models have been evaluated on AutomationBench on BenchLM.

### Does AutomationBench affect BenchLM's overall score?

Yes. AutomationBench is a weighted benchmark inside the Agentic category, which carries 22% of BenchLM's overall score. AutomationBench itself contributes 10% of that category score.

## Compare Top Models on AutomationBench

- [DeepSeek V4.1 Flash vs Atria Dawn Preview](/compare/atria-dawn-preview-vs-deepseek-v4-1-flash)
- [Atria Dawn Preview vs Muse Spark 1.3](/compare/atria-dawn-preview-vs-muse-spark-1-3)
- [Muse Spark 1.3 vs GLM-5.3-Flash](/compare/glm-5-3-flash-vs-muse-spark-1-3)
- [GLM-5.3-Flash vs GLM-5.3](/compare/glm-5-3-vs-glm-5-3-flash)
