# Best LLMs for Knowledge — September 2026 Leaderboard

> As of September 2026, Claude Fable 5.1 leads BenchLM's knowledge leaderboard with a weighted score of 86.7.

- **Last verified:** September 14, 2026
- Canonical page: https://benchlm.ai/knowledge
- **Ranking coverage:** 183 category-ranked models from 484 tracked models
- **Category weight:** 12% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 86.7 | 10 | 33 total |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 83.3 | 8 | 31 total |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 82.2 | 25 | 84 total |
| 4 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 82.2 | 10 | 34 total |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 80.5 | 14 | 47 total |
| 6 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 74.9 | 12 | 29 total |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 72.9 | 12 | 44 total |
| 8 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 72 | 12 | 52 total |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 71.4 | 12 | 44 total |
| 10 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 70.7 | 7 | 16 total |
| 11 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 70.6 | 11 | 32 total |
| 12 | [Grok 4.6](/models/grok-4-6) | xAI | 70.3 | 8 | 23 total |
| 13 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 69.7 | 12 | 30 total |
| 14 | [Grok 4.5](/models/grok-4-5) | xAI | 69.5 | 8 | 24 total |
| 15 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 69.2 | 13 | 40 total |
| 16 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 69.2 | 14 | 44 total |
| 17 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 68.8 | 6 | 55 total |
| 18 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 67.9 | 8 | 16 total |
| 19 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 67.8 | 11 | 38 total |
| 20 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 66.6 | 12 | 32 total |
| 21 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 66 | 12 | 25 total |
| 22 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 66 | 6 | 25 total |
| 23 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 65.7 | 6 | 14 total |
| 24 | [Muse Spark](/models/muse-spark) | Meta | 65.4 | 11 | 30 total |
| 25 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 64.7 | 7 | 18 total |
| 26 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 64.6 | 8 | 11 total |
| 27 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 64.1 | 12 | 38 total |
| 28 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 61.8 | 15 | 47 total |
| 29 | [GLM-5.3](/models/glm-5-3) | Z.AI | 61.8 | 8 | 32 total |
| 30 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 61.7 | 7 | 20 total |
| 31 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | Moonshot AI | 61.5 | 6 | 11 total |
| 32 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 61.4 | 11 | 34 total |
| 33 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 61.1 | 10 | 20 total |
| 34 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 61 | 15 | 37 total |
| 35 | [GLM-5.2](/models/glm-5-2) | Z.AI | 60.7 | 12 | 30 total |
| 36 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 60.3 | 7 | 10 total |
| 37 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 60.3 | 4 | 18 total |
| 38 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 59.8 | 14 | 36 total |
| 39 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | OpenAI | 59 | 2 | 7 total |
| 40 | [GPT-5.4 Pro](/models/gpt-5-4-pro) | OpenAI | 58.6 | 4 | 11 total |
| 41 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 58.1 | 6 | 11 total |
| 42 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 57.8 | 1 | 30 total |
| 43 | [Inkling](/models/inkling) | Thinking Machines Lab | 57.4 | 12 | 31 total |
| 44 | [Grok 4.3](/models/grok-4-3) | xAI | 57 | 10 | 20 total |
| 45 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 56.9 | 8 | 17 total |
| 46 | [Hy4 preview](/models/hy4-preview) | Tencent | 56.5 | 4 | 24 total |
| 47 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 56.5 | 12 | 32 total |
| 48 | [Hy3](/models/hy3) | Tencent | 56.3 | 6 | 7 total |
| 49 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 56.2 | 13 | 61 total |
| 50 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 56.2 | 6 | 9 total |
| 51 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 55.8 | 12 | 30 total |
| 52 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 55.6 | 11 | 26 total |
| 53 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 55.5 | 10 | 29 total |
| 54 | [GLM-5.1](/models/glm-5-1) | Z.AI | 54.8 | 10 | 31 total |
| 55 | [Apodex 1.1](/models/apodex-1-1) | Apodex | 54.6 | 9 | 18 total |
| 56 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | Anthropic | 54.3 | 7 | 8 total |
| 57 | [GLM-5](/models/glm-5) | Z.AI | 54.1 | 12 | 40 total |
| 58 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 54.1 | 12 | 42 total |
| 59 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 54 | 10 | 20 total |
| 60 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 54 | 13 | 49 total |
| 61 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 54 | 14 | 52 total |
| 62 | [GLM-5-Turbo](/models/glm-5-turbo) | Z.AI | 54 | 6 | 7 total |
| 63 | [MiMo-V2-Pro](/models/mimo-v2-pro) | Xiaomi | 53.8 | 6 | 7 total |
| 64 | [MiniMax M3](/models/minimax-m3) | MiniMax | 53.2 | 8 | 32 total |
| 65 | [Grok 4](/models/grok-4) | xAI | 52.6 | 6 | 10 total |
| 66 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 52.4 | 4 | 13 total |
| 67 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 52.2 | 2 | 13 total |
| 68 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 52.2 | 8 | 19 total |
| 69 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 52.1 | 12 | 48 total |
| 70 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | OpenAI | 51.6 | 6 | 9 total |
| 71 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | Google | 51.6 | 2 | 7 total |
| 72 | [GLM-5V-Turbo](/models/glm-5v-turbo) | Z.AI | 51.4 | 6 | 8 total |
| 73 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 51.2 | 1 | 9 total |
| 74 | [GPT-5 (medium)](/models/gpt-5-medium) | OpenAI | 50.9 | 6 | 7 total |
| 75 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 50.7 | 8 | 8 total |
| 76 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 50.6 | 8 | 14 total |
| 77 | [Agents-A1](/models/agents-a1) | InternScience | 50.2 | 1 | 6 total |
| 78 | [Agents-A1-F16-GGUF](/models/agents-a1-f16-gguf) | InternScience | 50.2 | 0 | 0 total |
| 79 | [Agents-A1-FP8](/models/agents-a1-fp8) | InternScience | 50.2 | 0 | 0 total |
| 80 | [Agents-A1-Q4_K_M-GGUF](/models/agents-a1-q4-k-m-gguf) | InternScience | 50.2 | 0 | 0 total |
| 81 | [Agents-A1-Q8_0-GGUF](/models/agents-a1-q8-0-gguf) | InternScience | 50.2 | 0 | 0 total |
| 82 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 50.2 | 12 | 35 total |
| 83 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | xAI | 49.5 | 6 | 8 total |
| 84 | [MiMo-V2-Omni](/models/mimo-v2-omni) | Xiaomi | 49.5 | 6 | 8 total |
| 85 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 49.3 | 6 | 22 total |
| 86 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 49 | 9 | 22 total |
| 87 | [o3](/models/o3) | OpenAI | 49 | 6 | 9 total |
| 88 | [Hy3 Preview](/models/hy3-preview) | Tencent | 49 | 9 | 12 total |
| 89 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 48.9 | 6 | 12 total |
| 90 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 48.7 | 10 | 30 total |
| 91 | [Interfaze Beta](/models/interfaze-beta) | Interfaze | 48.6 | 3 | 10 total |
| 92 | [o3-pro](/models/o3-pro) | OpenAI | 48.5 | 2 | 1 total |
| 93 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 48.5 | 6 | 19 total |
| 94 | [Mercury 2.5](/models/mercury-2-5) | Inception | 47.7 | 3 | 7 total |
| 95 | [DeepSeek-R1](/models/deepseek-r1) | DeepSeek | 47.6 | 6 | 6 total |
| 96 | [GLM-4.7](/models/glm-4-7) | Z.AI | 47.6 | 9 | 16 total |
| 97 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 47.5 | 11 | 25 total |
| 98 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 47.2 | 9 | 20 total |
| 99 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 47.2 | 6 | 23 total |
| 100 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | xAI | 47.1 | 6 | 8 total |
| 101 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 46.6 | 12 | 46 total |
| 102 | [DeepSeek V3.1 (Reasoning)](/models/deepseek-v3-1-reasoning) | DeepSeek | 46.6 | 6 | 6 total |
| 103 | [Mellum2-12B-A2.5B-Thinking](/models/mellum2-12b-a2-5b-thinking) | JetBrains | 46.5 | 3 | 5 total |
| 104 | [Mellum2-12B-A2.5B-Instruct](/models/mellum2-12b-a2-5b-instruct) | JetBrains | 45.8 | 3 | 5 total |
| 105 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 45.8 | 10 | 15 total |
| 106 | [Qwen3 235B 2507](/models/qwen3-235b-2507) | Alibaba | 45.8 | 3 | 4 total |
| 107 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 45.4 | 9 | 20 total |
| 108 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | Google | 45.3 | 6 | 9 total |
| 109 | [o1](/models/o1) | OpenAI | 45 | 8 | 10 total |
| 110 | [o1-preview](/models/o1-preview) | OpenAI | 44.9 | 2 | 1 total |
| 111 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 44.7 | 14 | 26 total |
| 112 | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Zyphra | 44.7 | 3 | 6 total |
| 113 | [DeepSeek V3.1](/models/deepseek-v3-1) | DeepSeek | 44.6 | 6 | 6 total |
| 114 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 44.6 | 11 | 31 total |
| 115 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 44.3 | 2 | 10 total |
| 116 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 44.2 | 11 | 49 total |
| 117 | [GLM-4.6](/models/glm-4-6) | Z.AI | 44 | 6 | 9 total |
| 118 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 43.5 | 3 | 10 total |
| 119 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 43.3 | 9 | 12 total |
| 120 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 42.9 | 4 | 16 total |
| 121 | [K-Exaone](/models/k-exaone) | LG AI Research | 42.3 | 6 | 7 total |
| 122 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 42.3 | 7 | 10 total |
| 123 | [MiMo-V2-Flash](/models/mimo-v2-flash) | Xiaomi | 42.2 | 8 | 8 total |
| 124 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 42.1 | 8 | 12 total |
| 125 | [Qwen3 Max](/models/qwen3-max) | Alibaba | 41.9 | 6 | 7 total |
| 126 | [Mistral Large 3](/models/mistral-large-3) | Mistral | 41.9 | 6 | 9 total |
| 127 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 41.3 | 4 | 18 total |
| 128 | [Sarvam 105B](/models/sarvam-105b) | Sarvam | 40.9 | 6 | 6 total |
| 129 | [Mistral Small 4](/models/mistral-small-4) | Mistral | 40.9 | 6 | 9 total |
| 130 | [Grok 4.1 Fast](/models/grok-4-1-fast) | xAI | 40.7 | 6 | 7 total |
| 131 | [Soofi S 30B-A3B](/models/soofi-s-30b-a3b) | Soofi Project | 40.6 | 4 | 8 total |
| 132 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 40.5 | 8 | 17 total |
| 133 | [GLM-4.5-Air](/models/glm-4-5-air) | Z.AI | 39.8 | 6 | 6 total |
| 134 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 39.8 | 8 | 13 total |
| 135 | [Ling 2.6 Flash](/models/ling-2-6-flash) | InclusionAI | 39.8 | 7 | 10 total |
| 136 | [o3-mini](/models/o3-mini) | OpenAI | 39.8 | 5 | 7 total |
| 137 | [Claude 4 Sonnet](/models/claude-4-sonnet) | Anthropic | 39.4 | 6 | 9 total |
| 138 | [Solar Pro 2](/models/solar-pro-2) | Upstage | 39.3 | 6 | 6 total |
| 139 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 39.3 | 8 | 9 total |
| 140 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 39.1 | 6 | 9 total |
| 141 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 39 | 11 | 19 total |
| 142 | [Qwen3-Omni-30B-A3B-Instruct](/models/qwen3-omni-30b-a3b-instruct) | Alibaba | 38.8 | 6 | 7 total |
| 143 | [Sarvam 30B](/models/sarvam-30b) | Sarvam | 38.8 | 6 | 6 total |
| 144 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 38.8 | 9 | 22 total |
| 145 | [Gemma 4 E4B](/models/gemma-4-e4b) | Google | 37.8 | 8 | 8 total |
| 146 | [Grok Code Fast 1](/models/grok-code-fast-1) | xAI | 37.2 | 6 | 6 total |
| 147 | [GPT-4o](/models/gpt-4o) | OpenAI | 37.1 | 6 | 7 total |
| 148 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 36.9 | 5 | 13 total |
| 149 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 36.8 | 6 | 13 total |
| 150 | [Command A+](/models/command-a-plus) | Cohere | 36.7 | 6 | 12 total |
| 151 | [Gemma 4 E2B](/models/gemma-4-e2b) | Google | 36.5 | 8 | 8 total |
| 152 | [Mistral Large 2](/models/mistral-large-2) | Mistral | 36.2 | 6 | 5 total |
| 153 | [Exaone 4.0 32B](/models/exaone-4-0-32b) | LG AI Research | 36.2 | 7 | 7 total |
| 154 | [Exaone 4.0 1.2B](/models/exaone-4-0-1-2b) | LG AI Research | 36 | 6 | 6 total |
| 155 | [Claude 4.1 Opus Thinking](/models/claude-4-1-opus-thinking) | Anthropic | 36 | 3 | 6 total |
| 156 | [Granite-4.0-1B](/models/granite-4-0-1b) | IBM | 35.8 | 6 | 6 total |
| 157 | [Granite-4.0-H-1B](/models/granite-4-0-h-1b) | IBM | 35.7 | 6 | 6 total |
| 158 | [Mistral Medium 3](/models/mistral-medium-3) | Mistral | 35.2 | 6 | 7 total |
| 159 | [Granite-4.0-H-350M](/models/granite-4-0-h-350m) | IBM | 35.1 | 6 | 6 total |
| 160 | [Granite-4.0-350M](/models/granite-4-0-350m) | IBM | 35.1 | 6 | 6 total |
| 161 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 34.7 | 8 | 20 total |
| 162 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 34.6 | 8 | 13 total |
| 163 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 34.5 | 6 | 12 total |
| 164 | [Claude 3 Opus](/models/claude-3-opus) | Anthropic | 34.1 | 3 | 2 total |
| 165 | [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) | DeepSeek | 34.1 | 3 | 4 total |
| 166 | [LFM2.5-VL-450M](/models/lfm2-5-vl-450m) | LiquidAI | 33.9 | 2 | 7 total |
| 167 | [LFM2.5-230M](/models/lfm2-5-230m) | LiquidAI | 33.7 | 3 | 6 total |
| 168 | [Gemma 3 27B](/models/gemma-3-27b) | Google | 33.6 | 6 | 9 total |
| 169 | [Gemini 1.5 Pro](/models/gemini-1-5-pro) | Google | 33.6 | 3 | 3 total |
| 170 | [Kimi K2](/models/kimi-k2) | Moonshot AI | 33.2 | 6 | 8 total |
| 171 | [Llama 4 Scout](/models/llama-4-scout) | Meta | 33.1 | 6 | 10 total |
| 172 | [Phi-4](/models/phi-4) | Microsoft | 33 | 6 | 6 total |
| 173 | [Qwen2.5 Coder 32B Instruct](/models/qwen2-5-coder-32b-instruct) | Alibaba | 32.6 | 3 | 2 total |
| 174 | [GPT-4o mini](/models/gpt-4o-mini) | OpenAI | 31.2 | 3 | 5 total |
| 175 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 30.5 | 6 | 10 total |
| 176 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 30.3 | 11 | 14 total |
| 177 | [GPT-4.1 nano](/models/gpt-4-1-nano) | OpenAI | 29.7 | 8 | 12 total |
| 178 | [GPT-4 Turbo](/models/gpt-4-turbo) | OpenAI | 28.5 | 2 | 1 total |
| 179 | [Nova Pro](/models/nova-pro) | Amazon | 26.8 | 6 | 7 total |
| 180 | [Claude 3 Haiku](/models/claude-3-haiku) | Anthropic | 25.1 | 6 | 7 total |
| 181 | [Gemini 1.0 Pro](/models/gemini-1-0-pro) | Google | 21.2 | 3 | 2 total |
| 182 | [Laguna M.1](/models/laguna-m-1) | Poolside | 18.6 | 2 | 10 total |
| 183 | [Laguna XS.2](/models/laguna-xs-2) | Poolside | 17.8 | 2 | 10 total |

## Decision-ready shortlist

- #1 [Claude Fable 5.1](/models/claude-fable-5-1) — 86.7 weighted score, Proprietary, 1M context.
- #2 [Claude Fable 5](/models/claude-fable) — 83.3 weighted score, Proprietary, 1M+ context.
- #3 [Claude Opus 5](/models/claude-opus-5) — 82.2 weighted score, Proprietary, null context.
- #4 [GPT-6 Astra](/models/gpt-6-astra) — 82.2 weighted score, Proprietary, 1.05M context.
- #5 [GPT-5.6 Sol](/models/gpt-5-6-sol) — 80.5 weighted score, Proprietary, 1.05M context.

## Benchmarks in this category

### [MMLU](/benchmarks/mmlu) (Massive Multitask Language Understanding)

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

- Ranking status: Display only
- Year: 2020
- Format: Multiple choice questions
- Difficulty: Elementary to professional level

### [GPQA](/benchmarks/gpqa) (Graduate-Level Google-Proof Q&A)

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

- Ranking status: Weighted (7% of this category)
- Year: 2023
- Format: Multiple choice questions
- Difficulty: Graduate level

### [GPQA-D](/benchmarks/gpqa-diamond) (GPQA Diamond)

A display-only GPQA Diamond reference from provider comparison charts.

- Ranking status: Display only
- Year: 2026
- Format: Multiple choice questions
- Difficulty: Graduate level

### [SuperGPQA](/benchmarks/supergpqa) (SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines)

An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.

- Ranking status: Weighted (7% of this category)
- Year: 2025
- Format: Multiple choice questions
- Difficulty: Graduate level

### [MMLU-Pro](/benchmarks/mmlu-pro) (Massive Multitask Language Understanding Professional)

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

- Ranking status: Weighted (20% of this category)
- Year: 2024
- Format: 10-way multiple choice
- Difficulty: Professional level

### [HLE](/benchmarks/hle) (Humanity's Last Exam)

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

- Ranking status: Weighted (35% of this category)
- Year: 2025
- Format: Open-ended and multiple choice
- Difficulty: Frontier expert level

### [FrontierScience](/benchmarks/frontierscience) (FrontierScience)

A benchmark for research-level scientific reasoning, designed to separate frontier models on difficult science tasks that mix domain knowledge with deep reasoning.

- Ranking status: Display only
- Year: 2026
- Format: Scientific reasoning benchmark
- Difficulty: Research frontier

### [HLE w/o tools](/benchmarks/hlenotools) (Humanity's Last Exam without tools)

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

- Ranking status: Weighted (10% of this category)
- Year: 2026
- Format: Tool-free expert QA
- Difficulty: Frontier expert level

### [SimpleQA](/benchmarks/simpleqa) (Measuring Short-Form Factuality in Large Language Models)

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

- Ranking status: Weighted (5% of this category)
- Year: 2024
- Format: Short-form Q&A
- Difficulty: Factual accuracy focused

### [HealthBench Hard](/benchmarks/healthbench-hard) (HealthBench Hard)

A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

- Ranking status: Display only
- Year: 2026
- Format: Open-ended health evaluation
- Difficulty: Advanced health reasoning

### [HealthBench Professional](/benchmarks/healthbenchprofessional) (HealthBench Professional)

An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.

- Ranking status: Display only
- Year: 2026
- Format: Rubric-graded open-ended responses
- Difficulty: Professional clinical workflows

### [MedXpertQA (Text)](/benchmarks/medxpertqatext) (MedXpertQA Text)

A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.

- Ranking status: Display only
- Year: 2026
- Format: Medical MCQ
- Difficulty: Professional medical knowledge

### [FrontierScience Research](/benchmarks/frontierscienceresearch) (FrontierScience Research)

A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.

- Ranking status: Display only
- Year: 2026
- Format: Research evaluation
- Difficulty: Frontier scientific research

### [MMLU-Pro (Arcee)](/benchmarks/mmluproarcee) (MMLU-Pro first-party comparison snapshot)

A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.

- Ranking status: Display only
- Year: 2026
- Format: 10-way multiple choice
- Difficulty: Professional level

### [MMMLU](/benchmarks/mmmlu) (MMMLU)

A multilingual MMLU-style benchmark reported in provider evaluation tables.

- Ranking status: Display only
- Year: 2026
- Format: Exact match
- Difficulty: Broad multilingual knowledge
