BenchLM Leaderboard
Point estimates with Supported and Estimated evidence labels. Conditional ranges do not establish rank confidence.
#ModelCreatorEvidenceScore
1GPT-6 AstraOpenAISupported88.69
2Claude Opus 5.5AnthropicEstimated87.68
3Claude Sonnet 5.5AnthropicEstimated83.28
4Claude Fable 5.1AnthropicSupported82.82
5Claude Opus 5AnthropicSupported79.89
6Claude Fable 5AnthropicSupported79.41
7GPT-6 SolOpenAIEstimated79.16
8GPT-5.6 SolOpenAISupported78.91
9GPT-6.1 SolOpenAIEstimated77.55
10Gemini 4 ArgonGoogleEstimated77.18
LLM benchmark leaderboard by BenchLM.ai