BenchLM recommendation
Best Alibaba Qwen Models in 2026
As of July 20, 2026, the top model in best alibaba qwen models on the BenchLM leaderboard is Qwen3.7 Max with a score of 72.8.
Last verified: July 20, 2026
All Alibaba Qwen models ranked by benchmark performance.
Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Bottom line: Alibaba's Qwen series has improved significantly. Qwen3.5 397B (Reasoning) leads with strong math (92) and coding (85). The MoE variants offer efficient alternatives.
Qwen3.7 Max leads this ranking with a score of 72.8, followed by Qwen3.7 Plus (67.2) and Qwen3.6 Plus (65.2). There is meaningful separation between the top models, suggesting genuine performance differences.
The best open-weight option is Qwen3.5-27B (ranked #4 with a score of 60.7). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.
This ranking is based on provisional overall weighted scores across BenchLM.ai's scoring formula tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Qwen3.5 397B (Reasoning) leads Alibaba's lineup — best math (92) and coding (85).
Qwen3.6 Plus newest generation with 1M context and strong instruction following (90).
Qwen3.5-122B-A10B efficient MoE variant — good performance per compute.
How to choose
Full Rankings (20 models)
Key Takeaways
The top model is Qwen3.7 Max by Alibaba with a BenchAlign v5 score of 72.8 and Supported evidence.
The best open-weight model is Qwen3.5-27B at position #4.
20 models are included in this ranking.
Score in Context
What these scores mean
Models are ranked by the same overall BenchLM score used across all leaderboards. Comparing within Alibaba's lineup helps identify which model fits your use case and budget.
Known limitations
This page only shows Alibaba models. Cross-provider comparison requires the overall or category-specific leaderboards. Newer models may have limited benchmark coverage initially.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.