Best OpenAI Models in 2026
As of September 3, 2026, the top model in best openai models on the BenchLM leaderboard is GPT-6 Astra with a score of 81.9.
Bottom line: GPT-5.4 is OpenAI's strongest model — it leads on knowledge and holds top 3 overall. GPT-5.4 Pro adds premium reasoning at higher cost. GPT-5.3 Codex is the coding specialist.
About this ranking
Last verified: September 3, 2026
All OpenAI models ranked by benchmark performance — GPT-5, GPT-4o, o1, o3, and more.
Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.
GPT-6 Astra leads this ranking with a score of 81.9, followed by GPT-5.6 Sol (81.8) and GPT-5.5 (73). There is a significant gap between the leading models and the rest of the field.
The best open-weight option is GPT-OSS 120B (ranked #25 with a score of 49.4). Proprietary models hold a clear advantage in this category, though open-weight options may suffice for less demanding use cases.
This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.
What changed
GPT-5.4 leads OpenAI's lineup with the highest overall score and best knowledge (98).
GPT-5.4 Pro premium tier with perfect multimodal and math scores.
GPT-5.3 Codex coding-focused with perfect math and multilingual.
Full Rankings (39 models)
Key Takeaways
The top model is GPT-6 Astra by OpenAI with a BenchAlign v5 score of 81.9 and Estimated evidence.
The best open-weight model is GPT-OSS 120B at position #25.
39 models are included in this ranking.
Score in Context
What these scores mean
Models are ranked by the same overall BenchLM score used across all leaderboards. Comparing within OpenAI's lineup helps identify which model fits your use case and budget.
Known limitations
This page only shows OpenAI models. Cross-provider comparison requires the overall or category-specific leaderboards. Newer models may have limited benchmark coverage initially.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.