# Claw-Eval

> A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

Canonical page: https://benchlm.ai/benchmarks/claw-eval

- Category: [Agentic](/agentic)
- Last updated: 2026-05-09 snapshot

## About Claw-Eval

- Year: 2026
- Tasks: 300 tasks, 2,159 rubrics
- Format: End-to-end autonomous-agent evaluation with Pass^3 scoring
- Difficulty: Real-world general, multi-turn, and native multimodal agent execution
- Paper: [Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents](https://arxiv.org/abs/2604.06132)

Claw-Eval v1.1.0 evaluates autonomous agents on full-trajectory tasks audited for completion, safety, and robustness. Its primary Pass^3 metric requires a task to pass in all three independent trials, reducing lucky-run effects. BenchLM mirrors the official leaderboard as display-only because rows reflect benchmark harness execution as well as model capability.

Claw-Eval is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (26 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 70.4% |
| 2 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 68.3% |
| 3 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 67.8% |
| 4 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 63.8% |
| 5 | [Muse Spark](/models/muse-spark) | Meta | 63.8% |
| 6 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 62.3% |
| 7 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 62.3% |
| 8 | [GLM-5.1](/models/glm-5-1) | Z.AI | 62.3% |
| 9 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 60.3% |
| 10 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 59.8% |
| 11 | [SenseNova 6.7 Flash-Lite](https://claw-eval.github.io/model/sensenova_67_flash_lite) | SenseTime | 58.8% |
| 12 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 58.8% |
| 13 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 57.8% |
| 14 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 57.8% |
| 15 | [MiMo-V2-Pro](/models/mimo-v2-pro) | Xiaomi | 57.8% |
| 16 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 56.8% |
| 17 | [GLM-5-Turbo](/models/glm-5-turbo) | Z.AI | 55.8% |
| 18 | [GLM-5V-Turbo](/models/glm-5v-turbo) | Z.AI | 53.8% |
| 19 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 52.3% |
| 20 | [Agnes-2.0-flash](https://claw-eval.github.io/model/agnes_20_flash) | SapiensAI | 51.8% |
| 21 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 49.2% |
| 22 | [Mach-Mind-4-pro](https://claw-eval.github.io/model/mach_mind_4_pro) | Li | 49.2% |
| 23 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 48.7% |
| 24 | [MiMo-V2-Omni](/models/mimo-v2-omni) | Xiaomi | 45.2% |
| 25 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 40.2% |
| 26 | [Nemotron 3 Super 100B](/models/nemotron-3-super-100b) | NVIDIA | 5.5% |

## FAQ

### What does Claw-Eval measure?

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

### Which model leads the published Claw-Eval snapshot?

Claude Opus 4.6 currently leads the published Claw-Eval snapshot with a score of 70.4%.

### How many models are evaluated on Claw-Eval?

The 2026-05-09 snapshot contains 26 AI models.

### Does Claw-Eval affect BenchLM's overall score?

Not directly. Claw-Eval is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
