Best LLMs for Knowledge — September 2026 Leaderboard
Data refreshed:
General knowledge and factual understanding
As of September 2026, the top knowledge model on the BenchLM leaderboard is Claude Fable 5.1 with a BenchAlign knowledge score of 86.7.
Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.
- Data refreshed
- September 14, 2026
- Ranked
- 183 of 490 models
- Supported / Estimated
- 64 / 119
- Weighted evidence
- 6 of 15 benchmarks
15 tracked benchmarks
MMLU, GPQA, GPQA-D, SuperGPQA, MMLU-Pro, HLE, FrontierScience, HLE w/o tools, SimpleQA, HealthBench Hard, HealthBench Professional, MedXpertQA (Text), FrontierScience Research, MMLU-Pro (Arcee), MMMLU
Scope: Frontier science, Broad academic knowledge, Factuality
Evidence set: MMLU, GPQA, GPQA-D, SuperGPQA, MMLU-Pro, HLE, FrontierScience, HLE w/o tools, SimpleQA, HealthBench Hard, HealthBench Professional, MedXpertQA (Text), FrontierScience Research, MMLU-Pro (Arcee), MMMLU
Scope: Frontier science, Broad academic knowledge, Factuality
Best Knowledge picks
BenchLM summaries for knowledge plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Knowledge Leaderboard
Primary score: BenchAlign knowledge score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Supported positions have diverse direct evidence. Estimated positions remain ranked but carry wider uncertainty.
Filters
1 | 86.7% | 86.68 | — | — | — | — | — | 65% | — | 60.9% | — | — | — | — | — | — | — |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
2 | 83.3% | 83.28 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
3 | 82.2% | 82.23 | — | — | — | — | — | 64.7% | — | 56.3% | — | — | 59.8% | — | — | — | — |
4 | 82.2% | 82.16 | — | 96% | 96.0% | — | — | — | — | — | — | 36.3% | 63.4% | — | — | — | — |
5 | 80.5% | 80.49 | — | 94.6% | 94.6% | — | — | — | — | — | — | 33.1% | 60.5% | — | — | — | — |
6 | 74.9% | 74.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
7 | 72.9% | 72.86 | — | 93.6% | 93.6% | — | — | 52.2% | — | 41.4% | — | — | — | — | — | — | — |
8 | 72.0% | 72.01 | — | 93.5% | 93.5% | — | — | 56% | — | 43.5% | — | — | — | — | — | — | — |
9 | 71.4% | 71.41 | — | 93.6% | 93.6% | — | — | 57.9% | — | 49.8% | — | — | — | — | — | — | — |
10 | 70.7% | 70.74 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
11 | 70.6% | 70.64 | — | — | — | — | — | 62.1% | — | 52.2% | — | — | 59.3% | — | — | — | — |
12 | 70.3% | 70.32 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
13 | 69.7% | 69.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
14 | 69.5% | 69.46 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
15 | 69.2% | 69.19 | — | 92.8% | 92.8% | — | — | 52.1% | — | 39.8% | — | 40.1% | 48.1% | 59.6% | — | — | — |
16 | 69.2% | 69.16 | — | 92.9% | 92.9% | — | — | — | — | — | — | 32.7% | 57.7% | — | — | — | — |
17 | 68.8% | 68.78 | — | 92.6% | 92.6% | — | — | 43.6% | — | 43.6% | — | — | — | — | — | — | — |
18 | 67.9% | 67.91 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
19 | 67.8% | 67.81 | — | 92.2% | 92.7% | — | — | 40.2% | — | — | — | — | — | — | — | — | — |
20 | 66.6% | 66.58 | — | — | — | — | — | 57.4% | — | 43.2% | — | — | — | — | — | — | — |
21 | 66% | 66 | — | — | 94.3% | — | — | — | — | 45.4% | — | 20.6% | — | 71.5% | — | — | — |
22 | 66.0% | 65.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
23 | 65.7% | 65.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
24 | 65.4% | 65.4 | — | — | 89.5% | — | — | 50.4% | — | 42.8% | — | 42.8% | — | 52.6% | — | — | — |
25 | 64.7% | 64.69 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — |
Top AI Models for Knowledge — September 2026
As of September 2026, Claude Fable 5.1 leads the BenchAlign knowledge leaderboard with a score of 86.7, followed by Claude Fable 5 (83.3) and Claude Opus 5 (82.2). BenchLM is currently showing 64 Supported and 119 Estimated models in this category.
Ranks #1 on the current knowledge board with a Supported evidence label.
Ranks #2 on the current knowledge board with a Supported evidence label.
Ranks #3 on the current knowledge board with a Supported evidence label.
What changed
Claude Fable 5.1 ranks #1 at 86.7 with a Supported evidence label.
Claude Fable 5 ranks #2 at 83.3 with a Supported evidence label.
Claude Opus 5 ranks #3 at 82.2 with a Supported evidence label.
Top models by benchmark
Expert-level questions in biology, physics, and chemistry(7% of category score)
Score in Context
What these scores mean
BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.
Known limitations
Estimated rows have less diverse direct evidence and wider uncertainty. They remain ranked so a newly released model is not treated as weak merely because fewer benchmark publishers have evaluated it.
How we weight
This lens combines category-relevant external evidence with admitted benchmark protocols. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| MMLU | — | Display only | Tests knowledge across 57 academic subjects |
| GPQA | 7% | Weighted | Expert-level questions in biology, physics, and chemistry |
| GPQA-D | — | Display only | Provider-table reference for GPQA Diamond scores reported in first-party comparison charts. |
| SuperGPQA | 7% | Weighted | Enhanced version covering 285 disciplines |
| MMLU-Pro | 20% | Weighted | Harder version of MMLU with 10 answer choices and more reasoning-focused questions |
| HLE | 35% | Weighted | Extremely difficult questions contributed by domain experts worldwide to test frontier AI |
| FrontierScience | — | Display only | Research-level science and scientific reasoning benchmark |
| HLE w/o tools | 10% | Weighted | Tool-free variant of Humanity's Last Exam used to isolate raw frontier reasoning without external aids |
| SimpleQA | 5% | Weighted | Factual question answering benchmark |
| HealthBench Hard | — | Display only | A harder health reasoning benchmark subset used in first-party frontier model comparisons. |
| HealthBench Professional | — | Display only | An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks. |
| MedXpertQA (Text) | — | Display only | Medical multiple-choice benchmark covering many specialties with text-only questions. |
| FrontierScience Research | — | Display only | A research-oriented FrontierScience variant focused on scientific investigation and solution quality. |
| MMLU-Pro (Arcee) | — | Display only | Display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart. |
| MMMLU | — | Display only | A multilingual MMLU-style benchmark reported in provider evaluation tables. |
About Knowledge Benchmarks
Tests knowledge across 57 academic subjects
Common questions
What is the best LLM for knowledge tasks?
The top LLMs for knowledge tasks are ranked by benchmarks like MMLU and GPQA, which test factual accuracy and expert-level understanding across dozens of subjects.
What is MMLU and how does it measure LLM knowledge?
MMLU (Massive Multitask Language Understanding) tests LLMs across 57 subjects from STEM to humanities, measuring broad factual knowledge and reasoning at varying difficulty levels.
What benchmarks test knowledge in AI models?
Key knowledge benchmarks include MMLU, MMLU-Pro, GPQA, SuperGPQA, HLE, and FrontierScience, each evaluating different depths of factual and scientific understanding.
How do knowledge benchmarks differ from reasoning benchmarks?
Knowledge benchmarks focus on factual recall and domain expertise, while reasoning benchmarks test logical deduction and multi-step problem solving independent of specific facts.
Knowledge benchmark updates
Know which model knows the most — updated every week.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.
Related
AI Web Scraping Tools
Compare tools that collect web data for retrieval and monitoring.
Web Data Pipeline
See how collection, monitoring, and validation work together.
Best LLMs Overall
Top models ranked across all benchmark categories.
Reasoning Benchmarks
Multi-step inference and logical deduction leaderboard.
Multilingual Benchmarks
Cross-language knowledge and reasoning scores.
LLM Selector Quiz
Find the best model for knowledge-intensive work.