BenchLM recommendation
Best Mistral Models in 2026
As of July 20, 2026, the top model in best mistral models on the BenchLM leaderboard is Mistral Large 3 with a score of 50.4.
Last verified: July 20, 2026
All Mistral AI models ranked by benchmark performance — Mistral Large, Mixtral, and more.
Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Bottom line: Mistral Large 3 is the only competitive model in the lineup — strongest on multimodal (67) but trailing the frontier. Mistral is Europe's leading AI lab.
Mistral Large 3 leads this ranking with a score of 50.4, followed by Ministral 3 14B (Reasoning) (49.3) and Mixtral 8x22B Instruct v0.1 (48.5). The top three are separated by just a few points — any of them would perform well for this use case.
The best open-weight option is Ministral 3 14B (Reasoning) (ranked #2 with a score of 49.3). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.
This ranking is based on provisional overall weighted scores across BenchLM.ai's scoring formula tracked by BenchLM.ai. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Mistral Large 3 leads Mistral's lineup — strongest on multimodal (67) and math (65).
Mistral Large 2 previous generation, now significantly behind Large 3.
How to choose
Full Rankings (14 models)
Key Takeaways
The top model is Mistral Large 3 by Mistral with a BenchAlign v5 score of 50.4 and Supported evidence.
The best open-weight model is Ministral 3 14B (Reasoning) at position #2.
14 models are included in this ranking.
Score in Context
What these scores mean
Models are ranked by the same overall BenchLM score used across all leaderboards. Comparing within Mistral's lineup helps identify which model fits your use case and budget.
Known limitations
This page only shows Mistral models. Cross-provider comparison requires the overall or category-specific leaderboards. Newer models may have limited benchmark coverage initially.
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.