Claw-Eval
We show this table for reference; we do not rank on it.
A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.
Pass^3 on Claw-Eval — 2026-05-09 snapshot
We mirror the published pass^3 view for Claw-Eval. Claude Opus 4.6 leads the public snapshot at 70.4%, followed by Step 3.7 Flash (68.3%) and Claude Sonnet 4.6 (67.8%). We do not use these results to rank models overall.
Claude Opus 4.6
Anthropic
Step 3.7 Flash
StepFun
Claude Sonnet 4.6
Anthropic
26 modelsAgenticCurrentDisplay onlyUpdated 2026-05-09 snapshot
Pass^3 table (26 models)
ScoreHow Claw-Eval is shown here
BenchLM mirrors the official Claw-Eval 2026-05-09 leaderboard snapshot. The source benchmark contains 300 human-verified tasks, 2,159 rubric items, and uses Pass^3 as the primary metric across 3 independent trials.
The public Claw-Eval site separates the 199-task general plus multi-turn agent table from the 101-task native multimodal table. BenchLM sorts this page by the primary general plus multi-turn Pass^3 table and preserves native multimodal split scores in the mirrored snapshot metadata.
Claw-Eval is display only on BenchLM. It is strong evidence about agent reliability, but the public rows are benchmark-harness results rather than normalized model-only rankings, so they are excluded from BenchLM overall and category scores.
Snapshot
The published Claw-Eval snapshot places Claude Opus 4.6 first at 70.4%. The third row is 2.6 points behind. The broader top-10 range is 10.6 points, so the table still separates the published systems.
26 models have been evaluated on Claw-Eval. The benchmark falls in the Agentic category. Claw-Eval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Claw-Eval
Year
2026
Tasks
300 tasks, 2,159 rubrics
Format
End-to-end autonomous-agent evaluation with Pass^3 scoring
Difficulty
Real-world general, multi-turn, and native multimodal agent execution
Claw-Eval v1.1.0 evaluates autonomous agents on full-trajectory tasks audited for completion, safety, and robustness. Its primary Pass^3 metric requires a task to pass in all three independent trials, reducing lucky-run effects. BenchLM mirrors the official leaderboard as display-only because rows reflect benchmark harness execution as well as model capability.
Freshness and provenance
Version
Claw-Eval 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Claw-Eval measure?
A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.
Which model leads the published Claw-Eval snapshot?
Claude Opus 4.6 currently leads the published Claw-Eval snapshot with 70.4% pass^3. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Claw-Eval?
The 2026-05-09 snapshot snapshot contains 26 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.