# AutomationBench, Zapier private evaluation set release 1.0.6 (AutomationBench (Zapier 1.0.6))

> An agent benchmark for completing business workflows across simulated applications and API endpoints.

Canonical page: https://benchlm.ai/benchmarks/automationbenchzapier106

- Category: [Agentic](/agentic)
- Last updated: September 29, 2026

## About AutomationBench (Zapier 1.0.6)

- Year: 2026
- Tasks: Simulated business workflows across 47 applications
- Format: Task pass rate
- Difficulty: Long-horizon automation
- Paper: [Claude Sonnet 5.5 System Card](https://www-cdn.anthropic.com/870c8f525702625d2c62fc6dd04c857e3250bec1/Claude%20Sonnet%205.5%20System%20Card.pdf)

Section 8.14.6 reports Zapier's private release 1.0.6 evaluation set. The Claude Sonnet 5.5 run used API default fallbacks. This lane stays separate from the 600-task public subset and Artificial Analysis AutomationBench-AA.

AutomationBench (Zapier 1.0.6) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Anthropic | 44.7% |

## FAQ

### What does AutomationBench (Zapier 1.0.6) measure?

An agent benchmark for completing business workflows across simulated applications and API endpoints.

### Which model scores highest on AutomationBench (Zapier 1.0.6)?

Claude Sonnet 5.5 by Anthropic currently leads with a score of 44.7% on AutomationBench (Zapier 1.0.6).

### How many models are evaluated on AutomationBench (Zapier 1.0.6)?

1 AI models have been evaluated on AutomationBench (Zapier 1.0.6) on BenchLM.

### Does AutomationBench (Zapier 1.0.6) affect BenchLM's overall score?

Not directly. AutomationBench (Zapier 1.0.6) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
