Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Agent & Tool-Use Benchmarks

Which AI models handle function calling, MCP tool use, browsing, and multi-step agent workflows best? Verified-ranked results across 24 agentic benchmarks.

Agentic carries 22% weightin BenchLM.ai's overall score — the single biggest category.

This page now shows only core agentic benchmark rows with an attached exact source record. Source-unverified manual rows are excluded from the displayed agentic score and table cells.

Best Agentic Model

GPT-5.6 Sol

Verified score: 92 · OpenAI

Best Open-Weight Agent

Qwen3.8 Max

Verified score: 86.1 · Alibaba

Benchmarks Tracked

26 benchmarks

Terminal, browsing, tool-use, and computer-use

Benchmark Categories

Core Weighted (3)

These 3 benchmarks determine agentic rankings

Tool Calling & MCP (6)

Function calling, MCP tool use, and structured workflows

Computer & Browser Use (5)

Desktop GUI, mobile, and browser navigation tasks

Specialized (4)

Domain-specific agentic tasks across ML, research, and airline

Top 15 Models by Weighted Agentic Score

OpenAIAnthropicGoogleMetaDeepSeekMistralxAIAlibaba

78 models match these filters.

CSVJSON
RankModelVerified score
1
GPT-5.6 SolOpenAI · proprietaryTB 91.9BC 92.2OS
92agentic / 100
2
Claude Opus 5Anthropic · proprietaryTB BC 90.8OS
90.8agentic / 100
3
GPT-5.5 ProOpenAI · proprietaryTB BC 90.1OS
90.1agentic / 100
4
Kimi K3Moonshot AI · proprietaryTB 88.3BC 91.2OS
89.5agentic / 100
5
GPT-5.4 ProOpenAI · proprietaryTB BC 89.3OS
89.3agentic / 100
6
GPT-5.6 TerraOpenAI · proprietaryTB 87.4BC 87.5OS
87.4agentic / 100
8
Qwen3.8 MaxAlibaba · open weightTB BC OS 86.1
86.1agentic / 100
9
Claude Fable 5Anthropic · proprietaryTB 84.3BC OS 85
84.6agentic / 100
10
Qwen3.8-27BAlibaba · open weightTB BC OS 84.3
84.3agentic / 100
11
GPT-5.6 LunaOpenAI · proprietaryTB 84.7BC 83.3OS
84.1agentic / 100
13
Grok 4.5xAI · proprietaryTB 83.3BC OS
83.3agentic / 100
15
Holo3-35B-A3BH Company · open weightTB BC OS 82.56
82.6agentic / 100
17
Claude Sonnet 5Anthropic · proprietaryTB 80.4BC 84.7OS 81.2
81.9agentic / 100
18
GPT-5.5OpenAI · proprietaryTB 82BC 84.4OS 78.7
81.6agentic / 100
19
SWE-1.7Cognition · proprietaryTB 81.5BC OS
81.5agentic / 100
20
GLM-5.2Z.AI · open weightTB 81BC OS
81agentic / 100
21
Muse Spark 1.1Meta · proprietaryTB 80BC OS 80.8
80.4agentic / 100
22
Claude Opus 4.8Anthropic · proprietaryTB 74.6BC 84.3OS 83.4
80.3agentic / 100
23
Sakana FuguSakana AI · proprietaryTB 80.2BC OS
80.2agentic / 100
24
Holo3-122B-A10BH Company · proprietaryTB BC OS 78.85
78.9agentic / 100
25
Ornith-1.0-397BDeepReinforce AI · open weightTB 77.5BC OS
77.5agentic / 100
27
GPT-5.4OpenAI · proprietaryTB 75.1BC 82.7OS 75
77.2agentic / 100
28
Agents-A1InternScience · open weightTB BC 75.51OS
75.5agentic / 100
31
Kimi K2.6Moonshot AI · open weightTB 66.7BC 83.2OS 73.1
73.5agentic / 100
32
Claude Opus 4.6Anthropic · proprietaryTB 65.4BC 83.7OS 72.7
73agentic / 100
33
MiniMax M3MiniMax · open weightTB 66BC 83.52OS 70.06
72.3agentic / 100
34
Ling 3.0 FlashInclusionAI · open weightTB BC 72.2OS
72.2agentic / 100
35
Qwen3.7 PlusAlibaba · proprietaryTB 70.3BC OS 73.3
71.7agentic / 100
36
GPT-5.3 CodexOpenAI · proprietaryTB 77.3BC OS 64.7
71.4agentic / 100
37
Laguna S 2.1Poolside · open weightTB 70.2BC OS
70.2agentic / 100
38
Inkling-SmallThinking Machines Lab · open weightTB 64.7BC 77.4OS
70.1agentic / 100
39
Qwen3.7 MaxAlibaba · proprietaryTB 69.7BC OS
69.7agentic / 100
40
InklingThinking Machines Lab · open weightTB 63.8BC 77.1OS
69.4agentic / 100
41
Composer 2.5Cursor · proprietaryTB 69.3BC OS
69.3agentic / 100
42
MiMo-V2.5-ProXiaomi · proprietaryTB 68.4BC OS
68.4agentic / 100
43
Step 3.7 FlashStepFun · open weightTB 59.5BC 75.82OS
66.4agentic / 100
45
MiMo-V2.5Xiaomi · proprietaryTB 65.8BC OS
65.8agentic / 100
46
GPT-5.4 miniOpenAI · proprietaryTB 60BC OS 72.1
65.7agentic / 100
47
GLM-5.1Z.AI · open weightTB 63.5BC 68OS
65.4agentic / 100
50
Ornith-1.0-35BDeepReinforce AI · open weightTB 64.2BC OS
64.2agentic / 100
53
Claude Opus 4.5Anthropic · proprietaryTB 59.3BC OS 66.3
62.6agentic / 100
54
Composer 2Cursor · proprietaryTB 61.7BC OS
61.7agentic / 100
55
Qwen3.6 PlusAlibaba · proprietaryTB 61.6BC OS
61.6agentic / 100
56
Qwen3.6-27BAlibaba · open weightTB 59.3BC OS
59.3agentic / 100
57
Muse SparkMeta · proprietaryTB 59BC OS
59agentic / 100
58
MiniMax M2.7MiniMax · open weightTB 57BC OS
57agentic / 100
59
Qwen3.5 397BAlibaba · open weightTB 52.5BC 62OS
56.5agentic / 100
61
GLM-5Z.AI · open weightTB 56.2BC OS
56.2agentic / 100
62
GPT-5.2OpenAI · proprietaryTB BC 65.8OS 47.3
55.7agentic / 100
64
Kimi K2.5Moonshot AI · open weightTB 50.8BC 60.6OS
55agentic / 100
66
Hy3 PreviewTencent · open weightTB 54.4BC OS
54.4agentic / 100
67
Qwen3.5-27BAlibaba · open weightTB 41.6BC 61OS 56.2
52agentic / 100
68
Qwen3.6-35B-A3BAlibaba · open weightTB 51.5BC OS
51.5agentic / 100
71
Grok 4.20xAI · proprietaryTB 47.1BC OS
47.1agentic / 100
72
MAI-Thinking-1Microsoft · proprietaryTB 46BC OS
46agentic / 100
73
Laguna M.1Poolside · proprietaryTB 45.8BC OS
45.8agentic / 100
74
GLM-4.7Z.AI · open weightTB 41BC 52OS
45.7agentic / 100
75
Ornith-1.0-9BDeepReinforce AI · open weightTB 43.1BC OS
43.1agentic / 100
76
GPT-5.4 nanoOpenAI · proprietaryTB 46.3BC OS 39
42.9agentic / 100
77
Laguna XS.2Poolside · open weightTB 35.7BC OS
35.7agentic / 100

Agentic score = weighted average of Terminal-Bench 2.0 (40%), OSWorld-Verified (35%), and BrowseComp (25%), normalized by available weights. This page intentionally stays on BenchLM's verified ranking lane and only includes exact-source rows. Display-only benchmarks (MCP Atlas, Toolathlon, etc.) are tracked but do not affect rankings.

Frequently Asked Questions

What are LLM agent benchmarks?

Agent benchmarks test whether AI models can go beyond answering questions and actually complete multi-step tasks: browsing the web, writing and running code in a terminal, calling external APIs via function calling, and operating desktop or mobile interfaces. They measure real-world usefulness for autonomous workflows.

What is function calling and why does it matter?

Function calling (or tool use) lets an LLM invoke external tools, APIs, or databases as part of its response. This is critical for building AI agents that can search the web, query databases, send emails, or control other software. Benchmarks like BFCL v4 and Toolathlon specifically measure how reliably models select the right function and pass correct arguments.

What is MCP (Model Context Protocol)?

MCP is an open standard for connecting LLMs to external tools and data sources. MCP Atlas and MCP-Tasks benchmark how well models work with MCP-backed integrations. Strong MCP performance means a model integrates well into tool-rich agent architectures.

Why does agentic carry the most weight in BenchLM scores?

Agentic carries 22% of BenchLM's overall score because the ability to use tools, browse, and complete multi-step tasks is the strongest differentiator between models in production use. A model that scores well on knowledge but cannot reliably call functions or navigate software has limited real-world utility for agent workflows.

Which models are best for building AI agents?

Currently, GPT-5.6 Sol by OpenAI leads BenchLM's verified agentic rankings with a score of 92. The best open-weight agent model is Qwen3.8 Max (86.1). Check the leaderboard above for the full verified ranking.

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.