# Best Tool Use & Function Calling Models in 2026

> AI models ranked on sourced tool-use benchmarks including BFCL v4, MCP Atlas, Toolathlon, Tau Bench, and related tool-calling evaluations.

This reporting page focuses on structured output, tool routing, function calling, and MCP-style task completion. It is narrower than the general agentic leaderboard and is better aligned to developers choosing models for tool-heavy applications.

Canonical page: https://benchlm.ai/best/tool-use

Last updated: October 5, 2026

## Rankings

| Rank | Model | Creator | Type | Context | Score |
|------|-------|---------|------|---------|-------|
| 1 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | Proprietary | 1M | 72 |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | Proprietary | 1M | 70.6 |
| 3 | [GLM-5.1](/models/glm-5-1) | Z.AI | Open Weight | 203K | 70.1 |
| 4 | [MiniMax M3](/models/minimax-m3) | MiniMax | Open Weight | 1M | 70.1 |
| 5 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | Proprietary | 1M | 69.8 |
| 6 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | Proprietary | 1M | 68.8 |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | Proprietary | 1M | 67.8 |
| 8 | [GLM-5.2](/models/glm-5-2) | Z.AI | Open Weight | 1M | 67 |
| 9 | [GPT-5.4](/models/gpt-5-4) | OpenAI | Proprietary | 1.05M | 65.5 |
| 10 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | Open Weight | 200K | 62.9 |
| 11 | [LLaDA2.2-flash](/models/llada2-2-flash) | InclusionAI | Open Weight | 128K | 62.7 |
| 12 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | Open Weight | 256K | 60.5 |
| 13 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | Open Weight | 262K | 60.1 |
| 14 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | Proprietary | 1M | 60 |
| 15 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | Proprietary | 200K | 58.7 |
| 16 | [GLM-5](/models/glm-5) | Z.AI | Open Weight | 200K | 58.3 |
| 17 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | Open Weight | 128K | 56 |
| 18 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | Open Weight | 256K | 52 |

## Key Takeaways

- Top model: [Qwen3.7 Plus](/models/qwen3-7-plus) with a score of 72
- Best open-weight option: [GLM-5.1](/models/glm-5-1) at #3
- Models included: 18

## Compare the Leaders

- [Qwen3.7 Plus vs Claude Opus 4.8](/compare/claude-opus-4-8-vs-qwen3-7-plus)
