# Best LLMs for Agentic — September 2026 Leaderboard

> As of September 2026, Claude Fable 5.1 leads BenchLM's agentic leaderboard with a weighted score of 80.2.

- **Last verified:** September 10, 2026
- Canonical page: https://benchlm.ai/agentic
- **Ranking coverage:** 152 category-ranked models from 483 tracked models
- **Category weight:** 22% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 80.2 | 20 | 33 total |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 78.1 | 31 | 84 total |
| 3 | [Claude Fable 5](/models/claude-fable) | Anthropic | 74.6 | 18 | 31 total |
| 4 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 71.9 | 25 | 52 total |
| 5 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 70.4 | 19 | 34 total |
| 6 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 70.1 | 23 | 47 total |
| 7 | [Grok 4.6](/models/grok-4-6) | xAI | 69.5 | 14 | 23 total |
| 8 | [GLM-5.3](/models/glm-5-3) | Z.AI | 68.4 | 18 | 32 total |
| 9 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 67.2 | 15 | 55 total |
| 10 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 66.3 | 13 | 29 total |
| 11 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 65.8 | 10 | 32 total |
| 12 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 64 | 11 | 30 total |
| 13 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 63.4 | 16 | 42 total |
| 14 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 62.5 | 17 | 44 total |
| 15 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 61.6 | 5 | 14 total |
| 16 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 61.1 | 6 | 16 total |
| 17 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 61 | 21 | 44 total |
| 18 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 60.8 | 7 | 18 total |
| 19 | [Grok 4.5](/models/grok-4-5) | xAI | 60.5 | 8 | 24 total |
| 20 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 60.2 | 23 | 44 total |
| 21 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 60.1 | 13 | 25 total |
| 22 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 59.7 | 4 | 10 total |
| 23 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 59.4 | 17 | 32 total |
| 24 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 59.4 | 4 | 11 total |
| 25 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 59.3 | 9 | 30 total |
| 26 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 58.5 | 3 | 18 total |
| 27 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 58.5 | 9 | 29 total |
| 28 | [GLM-5.2](/models/glm-5-2) | Z.AI | 58.5 | 12 | 30 total |
| 29 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | OpenAI | 58.4 | 1 | 7 total |
| 30 | [Hy4 preview](/models/hy4-preview) | Tencent | 58.2 | 14 | 24 total |
| 31 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 57.4 | 12 | 20 total |
| 32 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 56.8 | 19 | 38 total |
| 33 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 56.5 | 20 | 36 total |
| 34 | [Muse Spark](/models/muse-spark) | Meta | 56.1 | 7 | 30 total |
| 35 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 55.9 | 1 | 11 total |
| 36 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 55.8 | 5 | 11 total |
| 37 | [GLM-5.1](/models/glm-5-1) | Z.AI | 54.3 | 13 | 31 total |
| 38 | [Apodex 1.1](/models/apodex-1-1) | Apodex | 53.3 | 7 | 18 total |
| 39 | [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) | Anthropic | 53 | 2 | 1 total |
| 40 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 52.5 | 16 | 40 total |
| 41 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 51.8 | 5 | 16 total |
| 42 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 51.8 | 3 | 12 total |
| 43 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 51.7 | 16 | 38 total |
| 44 | [Holo3-35B-A3B](/models/holo3-35b-a3b) | H Company | 51.3 | 1 | 1 total |
| 45 | [GLM-5](/models/glm-5) | Z.AI | 51 | 13 | 40 total |
| 46 | [SWE-1.7](/models/swe-1-7) | Cognition | 51 | 1 | 4 total |
| 47 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 50.6 | 3 | 9 total |
| 48 | [Ornith-1.0-397B](/models/ornith-1-0-397b) | DeepReinforce AI | 50.4 | 2 | 7 total |
| 49 | [Holo3-122B-A10B](/models/holo3-122b-a10b) | H Company | 50.4 | 1 | 1 total |
| 50 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 50.3 | 1 | 13 total |
| 51 | [Agents-A1](/models/agents-a1) | InternScience | 50.1 | 3 | 6 total |
| 52 | [Agents-A1-F16-GGUF](/models/agents-a1-f16-gguf) | InternScience | 50.1 | 0 | 0 total |
| 53 | [Agents-A1-FP8](/models/agents-a1-fp8) | InternScience | 50.1 | 0 | 0 total |
| 54 | [Agents-A1-Q4_K_M-GGUF](/models/agents-a1-q4-k-m-gguf) | InternScience | 50.1 | 0 | 0 total |
| 55 | [Agents-A1-Q8_0-GGUF](/models/agents-a1-q8-0-gguf) | InternScience | 50.1 | 0 | 0 total |
| 56 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 50.1 | 10 | 37 total |
| 57 | [BTL-4](/models/btl-4) | Bad Theory Labs | 49.9 | 1 | 2 total |
| 58 | [Composer 2.5](/models/composer-2-5) | Cursor | 49.5 | 1 | 5 total |
| 59 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 49.4 | 6 | 13 total |
| 60 | [Pokee-Isaac 28B](/models/pokee-isaac-28b) | Pokee AI | 49.2 | 5 | 7 total |
| 61 | [GPT-5 (high)](/models/gpt-5-high) | OpenAI | 48.8 | 4 | 2 total |
| 62 | [Composer 2](/models/composer-2) | Cursor | 48.8 | 1 | 5 total |
| 63 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | OpenAI | 48.6 | 3 | 9 total |
| 64 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 48.6 | 11 | 19 total |
| 65 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 48.2 | 7 | 8 total |
| 66 | [Qwen3.5 Plus](/models/qwen3-5-plus) | Alibaba | 48 | 1 | 4 total |
| 67 | [Hy3](/models/hy3) | Tencent | 47.8 | 3 | 7 total |
| 68 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 47.7 | 5 | 14 total |
| 69 | [Hy3 Preview](/models/hy3-preview) | Tencent | 47.7 | 5 | 12 total |
| 70 | [MiMo-V2-Flash](/models/mimo-v2-flash) | Xiaomi | 46.9 | 3 | 8 total |
| 71 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 46.8 | 14 | 35 total |
| 72 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | Moonshot AI | 46.8 | 7 | 11 total |
| 73 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 46.7 | 5 | 20 total |
| 74 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 46.7 | 17 | 34 total |
| 75 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 46.7 | 7 | 18 total |
| 76 | [Ornith-1.0-35B](/models/ornith-1-0-35b) | DeepReinforce AI | 46.6 | 2 | 7 total |
| 77 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 46.4 | 18 | 48 total |
| 78 | [Laguna S 2.1](/models/laguna-s-2-1) | Poolside | 46.2 | 2 | 6 total |
| 79 | [Mellum2-12B-A2.5B-Thinking](/models/mellum2-12b-a2-5b-thinking) | JetBrains | 45.8 | 1 | 5 total |
| 80 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 45.7 | 1 | 10 total |
| 81 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 45.5 | 5 | 15 total |
| 82 | [GLM-4.7](/models/glm-4-7) | Z.AI | 45.3 | 7 | 16 total |
| 83 | [Mellum2-12B-A2.5B-Instruct](/models/mellum2-12b-a2-5b-instruct) | JetBrains | 45.3 | 1 | 5 total |
| 84 | [Grok Build 0.1](/models/grok-build-0-1) | xAI | 45.1 | 1 | 0 total |
| 85 | [K-Exaone](/models/k-exaone) | LG AI Research | 45.1 | 3 | 7 total |
| 86 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 45 | 5 | 20 total |
| 87 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 44.9 | 3 | 19 total |
| 88 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 44.4 | 7 | 16 total |
| 89 | [Ornith-1.0-9B](/models/ornith-1-0-9b) | DeepReinforce AI | 44.3 | 2 | 7 total |
| 90 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 44.3 | 4 | 22 total |
| 91 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 44.2 | 5 | 9 total |
| 92 | [Mercury 2.5](/models/mercury-2-5) | Inception | 44 | 2 | 7 total |
| 93 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 44 | 9 | 30 total |
| 94 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 43.6 | 7 | 22 total |
| 95 | [Gemma 4 E4B](/models/gemma-4-e4b) | Google | 43.4 | 3 | 8 total |
| 96 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 43.3 | 2 | 12 total |
| 97 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 43.2 | 16 | 49 total |
| 98 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 43.1 | 1 | 13 total |
| 99 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 43 | 8 | 19 total |
| 100 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 43 | 8 | 23 total |
| 101 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | Google | 42.9 | 2 | 7 total |
| 102 | [Claude 4.1 Opus](/models/claude-4-1-opus) | Anthropic | 42.9 | 1 | 2 total |
| 103 | [Gemma 4 E2B](/models/gemma-4-e2b) | Google | 42.9 | 3 | 8 total |
| 104 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 42.9 | 5 | 20 total |
| 105 | [LFM2.5-VL-450M](/models/lfm2-5-vl-450m) | LiquidAI | 42.2 | 1 | 7 total |
| 106 | [LFM2.5-230M](/models/lfm2-5-230m) | LiquidAI | 42.2 | 1 | 6 total |
| 107 | [MiniMax M3](/models/minimax-m3) | MiniMax | 42.1 | 17 | 32 total |
| 108 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 41.9 | 12 | 31 total |
| 109 | [Ling 2.6 Flash](/models/ling-2-6-flash) | InclusionAI | 41.8 | 3 | 10 total |
| 110 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 41.2 | 16 | 47 total |
| 111 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 41.2 | 12 | 30 total |
| 112 | [MiniMax M2.5](/models/minimax-m2-5) | MiniMax | 41.1 | 0 | 1 total |
| 113 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 40.8 | 9 | 10 total |
| 114 | [Command A+](/models/command-a-plus) | Cohere | 40.6 | 4 | 12 total |
| 115 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 40.1 | 15 | 49 total |
| 116 | [Claude 4 Sonnet](/models/claude-4-sonnet) | Anthropic | 40 | 3 | 9 total |
| 117 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 39.8 | 10 | 20 total |
| 118 | [Mistral Large 3](/models/mistral-large-3) | Mistral | 39.8 | 4 | 9 total |
| 119 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 39.1 | 10 | 26 total |
| 120 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 39 | 0 | 1 total |
| 121 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 38.9 | 10 | 25 total |
| 122 | [Inkling](/models/inkling) | Thinking Machines Lab | 38.3 | 10 | 31 total |
| 123 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 38 | 5 | 12 total |
| 124 | [Mercury 2](/models/mercury-2) | Inception | 38 | 0 | 0 total |
| 125 | [Mistral Small 4](/models/mistral-small-4) | Mistral | 37.8 | 4 | 9 total |
| 126 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 37.3 | 17 | 61 total |
| 127 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 37 | 3 | 13 total |
| 128 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 36.9 | 8 | 32 total |
| 129 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 36.8 | 6 | 13 total |
| 130 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 36.7 | 4 | 9 total |
| 131 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 36.3 | 5 | 9 total |
| 132 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 36 | 16 | 52 total |
| 133 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 34.6 | 10 | 25 total |
| 134 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 34 | 5 | 17 total |
| 135 | [GPT-4o mini](/models/gpt-4o-mini) | OpenAI | 33.7 | 2 | 5 total |
| 136 | [GPT-4.1 nano](/models/gpt-4-1-nano) | OpenAI | 33.4 | 3 | 12 total |
| 137 | [Gemma 3 27B](/models/gemma-3-27b) | Google | 33 | 4 | 9 total |
| 138 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 31.8 | 6 | 20 total |
| 139 | [Llama 4 Scout](/models/llama-4-scout) | Meta | 31.4 | 4 | 10 total |
| 140 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 30.9 | 10 | 46 total |
| 141 | [Ministral 3 14B](/models/ministral-3-14b) | Mistral | 29.3 | 0 | 0 total |
| 142 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 27 | 2 | 10 total |
| 143 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 27 | 13 | 26 total |
| 144 | [Grok 4.3](/models/grok-4-3) | xAI | 26.9 | 8 | 20 total |
| 145 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 26.7 | 4 | 22 total |
| 146 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 26.1 | 7 | 14 total |
| 147 | [Laguna M.1](/models/laguna-m-1) | Poolside | 23.9 | 2 | 10 total |
| 148 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 23.8 | 4 | 10 total |
| 149 | [Laguna XS.2](/models/laguna-xs-2) | Poolside | 22.4 | 2 | 10 total |
| 150 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 21.9 | 11 | 17 total |
| 151 | [Ministral 3 8B](/models/ministral-3-8b) | Mistral | 21 | 0 | 0 total |
| 152 | [Ministral 3 3B](/models/ministral-3-3b) | Mistral | 19.8 | 0 | 0 total |

## Decision-ready shortlist

- #1 [Claude Fable 5.1](/models/claude-fable-5-1) — 80.2 weighted score, Proprietary, 1M context.
- #2 [Claude Opus 5](/models/claude-opus-5) — 78.1 weighted score, Proprietary, null context.
- #3 [Claude Fable 5](/models/claude-fable) — 74.6 weighted score, Proprietary, 1M+ context.
- #4 [Kimi K3](/models/kimi-k3) — 71.9 weighted score, Pending, 1.05M context.
- #5 [GPT-6 Astra](/models/gpt-6-astra) — 70.4 weighted score, Proprietary, 1.05M context.

## Benchmarks in this category

### [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) (Terminal-Bench 2.0)

A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.

- Ranking status: Weighted (30% of this category)
- Year: 2026
- Format: Interactive CLI agent evaluation
- Difficulty: Professional software engineering

### [BrowseComp](/benchmarks/browsecomp) (BrowseComp)

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

- Ranking status: Weighted (25% of this category)
- Year: 2025
- Format: Web search and evidence synthesis
- Difficulty: Hard web research

### [OSWorld-Verified](/benchmarks/osworld-verified) (OSWorld-Verified)

OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

- Ranking status: Weighted (25% of this category)
- Year: 2025
- Format: Execution-based interactive task success
- Difficulty: Multi-step desktop and cross-application workflows

### [OSWorld 2.0](/benchmarks/osworld2) (OSWorld 2.0)

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

- Ranking status: Weighted (10% of this category)
- Year: 2026
- Format: Interactive computer-use evaluation
- Difficulty: Long-horizon professional workflows

### [CyberGym](/benchmarks/cybergym) (CyberGym)

A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.

- Ranking status: Display only
- Year: 2026
- Format: Vulnerability reproduction and PoC generation
- Difficulty: Real-world cybersecurity

### [CWE-Bench](/benchmarks/cwebench) (CWE-Bench)

An external benchmark for evaluating whether coding agents can produce correct patches for real-world software vulnerabilities.

- Ranking status: Display only
- Year: 2026
- Format: Pass@1
- Difficulty: Automated vulnerability remediation

### [Cybench](/benchmarks/cybench) (Cybench)

A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.

- Ranking status: Display only
- Year: 2025
- Format: Cybersecurity agent task completion
- Difficulty: Professional cybersecurity

### [ExploitGym](/benchmarks/exploitgym) (ExploitGym)

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

- Ranking status: Display only
- Year: 2026
- Format: Working exploit generation
- Difficulty: Advanced cybersecurity exploitation

### [JobBench](/benchmarks/jobbench) (JobBench)

An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.

- Ranking status: Display only
- Year: 2026
- Format: Agentic workplace deliverables
- Difficulty: Professional multi-source workflows

### [BrowseComp-VL](/benchmarks/browsecompvl) (BrowseComp-VL)

A vision-language browsing benchmark for multimodal web research and tool-use workflows.

- Ranking status: Display only
- Year: 2026
- Format: Vision-language web research evaluation
- Difficulty: Multimodal browser-agent

### [OSWorld](/benchmarks/osworld) (OSWorld)

A computer-use benchmark for GUI task completion across the broader OSWorld task suite.

- Ranking status: Display only
- Year: 2026
- Format: Interactive GUI evaluation
- Difficulty: Broad computer-use suite

### [AndroidWorld](/benchmarks/androidworld) (AndroidWorld)

A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.

- Ranking status: Display only
- Year: 2026
- Format: Interactive mobile-agent evaluation
- Difficulty: Complex mobile task completion

### [WebVoyager](/benchmarks/webvoyager) (WebVoyager)

A browser-agent benchmark for completing multi-step workflows on live websites.

- Ranking status: Display only
- Year: 2026
- Format: Interactive browser-agent evaluation
- Difficulty: Multi-step web navigation

### [MCP Atlas](/benchmarks/mcpatlas) (MCP Atlas)

A benchmark for tool-calling over Model Context Protocol integrations and external tools.

- Ranking status: Display only
- Year: 2026
- Format: Interactive tool-calling evaluation
- Difficulty: Advanced tool use

### [Toolathlon](/benchmarks/toolathlon) (Toolathlon)

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

- Ranking status: Display only
- Year: 2026
- Format: Interactive tool-calling evaluation
- Difficulty: Advanced tool use

### [Finance Agent v2](/benchmarks/financeagentv2) (Finance Agent v2)

Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.

- Ranking status: Display only
- Year: 2026
- Format: Mean score across repeated runs
- Difficulty: Professional expert-task agent workflow

### [GDPval-AA](/benchmarks/gdpvalaa) (GDPval-AA)

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

- Ranking status: Display only
- Year: 2026
- Format: Elo
- Difficulty: Professional agentic workflows

### [ZClawBench](/benchmarks/zclawbench) (ZClawBench)

A Z.AI benchmark for OpenClaw-style agent workflows spanning information search, office work, data analysis, development and operations, automation, and security.

- Ranking status: Display only
- Year: 2026
- Format: End-to-end agent benchmark
- Difficulty: Broad productivity and operations workflows

### [τ²-bench results](/benchmarks/tau2-bench) (τ²-Bench Tool-Agent-User Evaluation)

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

- Ranking status: Display only
- Year: 2025
- Format: Published domain success or pass^k results
- Difficulty: Dual-control customer-service workflows

### [DeepSearchQA](/benchmarks/deepsearchqa) (DeepSearchQA)

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

- Ranking status: Display only
- Year: 2026
- Format: Search / open / find browser-agent evaluation
- Difficulty: Agentic web research

### [τ²-bench Airline](/benchmarks/tau2airline) (τ²-Bench Airline Domain)

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

- Ranking status: Display only
- Year: 2025
- Format: Domain success under a published trial policy
- Difficulty: Policy-constrained airline support workflows

### [PinchBench](/benchmarks/pinchbench) (PinchBench)

An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

- Ranking status: Display only
- Year: 2026
- Format: Average success rate from official runs
- Difficulty: Long-horizon agent workflows

### [OpenHands Index](/benchmarks/openhandsindex) (OpenHands Index)

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

- Ranking status: Display only
- Year: 2025
- Format: Macro-average across five coding-agent categories
- Difficulty: Real-world software engineering agent tasks

### [SWE-Atlas Refactoring](/benchmarks/sweatlasrefactoring) (SWE-Atlas Refactoring)

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

- Ranking status: Display only
- Year: 2026
- Format: Refactoring score with confidence intervals
- Difficulty: Real-world software-engineering agent tasks

### [SWE Refactor Bench](/benchmarks/swe-refactor-bench) (SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?)

Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.

- Ranking status: Display only
- Year: 2026
- Format: Migration audit, frozen behavioral checks, and agentic verification
- Difficulty: 6- to 30-hour autonomous repository migrations

### [AI4AI-Bench](/benchmarks/ai4ai-bench) (AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement)

Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.

- Ranking status: Display only
- Year: 2026
- Format: Four-hour code rewrite followed by sealed training and held-out evaluation
- Difficulty: End-to-end AI research and algorithm design

### [BFCL v4](/benchmarks/bfcl-v4) (Berkeley Function Calling Leaderboard v4)

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

- Ranking status: Display only
- Year: 2026
- Format: Tool invocation and schema evaluation
- Difficulty: Advanced tool use

### [MLE-Bench Lite](/benchmarks/mle-bench-lite) (MLE-Bench Lite)

A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.

- Ranking status: Display only
- Year: 2026
- Format: Autonomous iterative ML optimization
- Difficulty: Agentic machine learning

### [MM-ClawBench](/benchmarks/mmclawbench) (MM-ClawBench)

An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.

- Ranking status: Display only
- Year: 2026
- Format: Agent workflow evaluation
- Difficulty: Broad real-world agentic execution

### [Gert Labs](/benchmarks/gertlabs) (Gert Labs Composite Game Benchmark)

A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.

- Ranking status: Display only
- Year: 2026
- Format: Composite game leaderboard
- Difficulty: Agentic coding and decision-making
