Best Google AI Models in 2026
As of September 4, 2026, the top model in best google ai models on the BenchLM leaderboard is Gemini 3.8 Flash with a score of 78.4.
Bottom line: Gemini 3.1 Pro is Google's best — the top non-reasoning model on BenchLM. Gemini 3 Pro Deep Think adds reasoning capability. Flash variants offer strong value.
About this ranking
Last verified: September 4, 2026
All Google Gemini and Gemma models ranked by benchmark performance.
Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
Gemini 3.8 Flash leads this ranking with a score of 78.4, followed by Gemini 3.7 Flash (70.2) and Gemini 3.6 Flash (70.1). There is a significant gap between the leading models and the rest of the field.
The best open-weight option is Gemma 4 31B (ranked #10 with a score of 58.7). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.
This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
Gemini 3.1 Pro leads Google's lineup — best non-reasoning model on the leaderboard.
Gemini 3 Pro Deep Think reasoning variant with perfect multimodal and strong math (96).
Gemini 3 Pro solid all-rounder with good multimodal (86).
How to choose
Full Rankings (20 models)
Key Takeaways
The top model is Gemini 3.8 Flash by Google with a BenchAlign v5 score of 78.4 and Supported evidence.
The best open-weight model is Gemma 4 31B at position #10.
20 models are included in this ranking.
Score in Context
What these scores mean
Models are ranked by the same overall BenchLM score used across all leaderboards. Comparing within Google's lineup helps identify which model fits your use case and budget.
Known limitations
This page only shows Google models. Cross-provider comparison requires the overall or category-specific leaderboards. Newer models may have limited benchmark coverage initially.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.