# Artificial Analysis AutomationBench (AA AutomationBench)

> An independently evaluated automation benchmark from Artificial Analysis.

Canonical page: https://benchlm.ai/benchmarks/aaautomationbench

- Category: [Agentic](/agentic)
- Last updated: September 15, 2026

## About AA AutomationBench

- Year: 2026
- Tasks: Business-process automation tasks
- Format: Task success rate
- Difficulty: Agentic automation
- Paper: [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa)

Stored as a display-only agentic row.

AA AutomationBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 68.9% |
| 2 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 68.5% |
| 3 | [Grok 4.6](/models/grok-4-6) | xAI | 66.7% |
| 4 | [GLM-5.3](/models/glm-5-3) | Z.AI | 62.2% |
| 5 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 60.4% |
| 6 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 60.1% |
| 7 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 59.9% |
| 8 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 59.6% |
| 9 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 59.4% |
| 10 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 58.3% |
| 11 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 57.9% |
| 12 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 56.7% |
| 13 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 56.6% |
| 14 | [Claude Fable 5](/models/claude-fable) | Anthropic | 54.1% |

## FAQ

### What does AA AutomationBench measure?

An independently evaluated automation benchmark from Artificial Analysis.

### Which model scores highest on AA AutomationBench?

DeepSeek V4.1 Flash by DeepSeek currently leads with a score of 68.9% on AA AutomationBench.

### How many models are evaluated on AA AutomationBench?

14 AI models have been evaluated on AA AutomationBench on BenchLM.

### Does AA AutomationBench affect BenchLM's overall score?

Not directly. AA AutomationBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AA AutomationBench

- [DeepSeek V4.1 Flash vs GPT-6 Astra](/compare/deepseek-v4-1-flash-vs-gpt-6-astra)
- [GPT-6 Astra vs Grok 4.6](/compare/gpt-6-astra-vs-grok-4-6)
- [Grok 4.6 vs GLM-5.3](/compare/glm-5-3-vs-grok-4-6)
- [GLM-5.3 vs GLM-5.3-Flash](/compare/glm-5-3-vs-glm-5-3-flash)
