# MCP-Tasks

> A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.

Canonical page: https://benchlm.ai/benchmarks/mcptasks

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026

## About MCP-Tasks

- Year: 2026
- Tasks: MCP-integrated tool tasks
- Format: Interactive tool-use evaluation
- Difficulty: Advanced MCP workflows
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

MCP-Tasks is distinct from MCP Atlas in the Qwen3.6 comparisons. BenchLM keeps it separate until a fuller public benchmark specification is available because it appears to represent a different task protocol and score scale.

MCP-Tasks is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (5 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 74.2% |
| 2 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 74.1% |
| 3 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 71.8% |
| 4 | [GLM-5](/models/glm-5) | Z.AI | 60.8% |
| 5 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 59.1% |

## FAQ

### What does MCP-Tasks measure?

A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.

### Which model scores highest on MCP-Tasks?

Qwen3.5 397B by Alibaba currently leads with a score of 74.2% on MCP-Tasks.

### How many models are evaluated on MCP-Tasks?

5 AI models have been evaluated on MCP-Tasks on BenchLM.

### Does MCP-Tasks affect BenchLM's overall score?

Not directly. MCP-Tasks is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MCP-Tasks

- [Qwen3.5 397B vs Qwen3.6 Plus](/compare/qwen3-5-397b-vs-qwen3-6-plus)
- [Qwen3.6 Plus vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-qwen3-6-plus)
- [Claude Opus 4.5 vs GLM-5](/compare/claude-opus-4-5-vs-glm-5)
- [GLM-5 vs Kimi K2.5](/compare/glm-5-vs-kimi-k2-5)
