Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Knowledge benchmark report

Best LLMs for KnowledgeSeptember 2026 Leaderboard

Data refreshed:

General knowledge and factual understanding

As of September 2026, the top knowledge model on the BenchLM leaderboard is Claude Fable 5.1 with a BenchAlign knowledge score of 86.7.

Decision lens: the score determines position; Supported and Estimated labels describe the evidence behind that position without removing sparsely reported models.

Data refreshed
September 14, 2026
Ranked
183 of 490 models
Supported / Estimated
64 / 119
Weighted evidence
6 of 15 benchmarks
15 tracked benchmarks

MMLU, GPQA, GPQA-D, SuperGPQA, MMLU-Pro, HLE, FrontierScience, HLE w/o tools, SimpleQA, HealthBench Hard, HealthBench Professional, MedXpertQA (Text), FrontierScience Research, MMLU-Pro (Arcee), MMMLU

Scope: Frontier science, Broad academic knowledge, Factuality

Best Knowledge picks

BenchLM summaries for knowledge plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Knowledge Leaderboard

Primary score: BenchAlign knowledge score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Supported positions have diverse direct evidence. Estimated positions remain ranked but carry wider uncertainty.

Filters
Supported positions have diverse direct evidence. Estimated positions remain ranked with wider uncertainty.
Rank / modelWeighted Knowledge
1
Claude Fable 5.1Anthropic · ClosedSupported
86.7%
2
Claude Fable 5Anthropic · ClosedSupported
83.3%
3
Claude Opus 5Anthropic · ClosedSupported
82.2%
4
GPT-6 AstraOpenAI · ClosedSupported
82.2%
5
GPT-5.6 SolOpenAI · ClosedSupported
80.5%
6
Gemini 3.8 FlashGoogle · ClosedSupported
74.9%
7
GPT-5.5OpenAI · ClosedSupported
72.9%
8
Kimi K3Moonshot AI · ClosedSupported
72.0%
9
Claude Opus 4.8Anthropic · ClosedSupported
71.4%
10
Muse Spark 1.2Meta · ClosedEstimated
70.7%
11
Muse Spark 1.1Meta · ClosedSupported
70.6%
12
Grok 4.6xAI · ClosedSupported
70.3%
13
Gemini 3.7 FlashGoogle · ClosedSupported
69.7%
14
Grok 4.5xAI · ClosedSupported
69.5%
15
GPT-5.4OpenAI · ClosedSupported
69.2%
16
GPT-5.6 TerraOpenAI · ClosedSupported
69.2%
17
Qwen3.8 MaxAlibaba · Open weightSupported
68.8%
18
Gemini 3.6 FlashGoogle · ClosedSupported
67.9%
19
Gemini 3.5 FlashGoogle · ClosedSupported
67.8%
20
Claude Sonnet 5Anthropic · ClosedSupported
66.6%
21
Gemini 3.1 ProGoogle · ClosedSupported
66%
22
GLM-5.3-FlashZ.AI · Open weightSupported
66.0%
23
GPT-5.3 CodexOpenAI · ClosedEstimated
65.7%
24
Muse SparkMeta · ClosedSupported
65.4%
25
Gemini 3 ProGoogle · ClosedEstimated
64.7%

Top AI Models for KnowledgeSeptember 2026

As of September 2026, Claude Fable 5.1 leads the BenchAlign knowledge leaderboard with a score of 86.7, followed by Claude Fable 5 (83.3) and Claude Opus 5 (82.2). BenchLM is currently showing 64 Supported and 119 Estimated models in this category.

What changed

Claude Fable 5.1 ranks #1 at 86.7 with a Supported evidence label.

Claude Fable 5 ranks #2 at 83.3 with a Supported evidence label.

Claude Opus 5 ranks #3 at 82.2 with a Supported evidence label.

Top models by benchmark

Expert-level questions in biology, physics, and chemistry(7% of category score)

Score in Context

What these scores mean

BenchAlign places direct benchmarks and independent external signals on a common calibrated scale. The score is relative to the current evidence universe; it is not a raw percentage from any single test.

Known limitations

Estimated rows have less diverse direct evidence and wider uncertainty. They remain ranked so a newly released model is not treated as weak merely because fewer benchmark publishers have evaluated it.

How we weight

This lens combines category-relevant external evidence with admitted benchmark protocols. Evidence sources are calibrated for difficulty before aggregation, and no generated benchmark row contributes to the score.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Knowledge benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
MMLUDisplay onlyTests knowledge across 57 academic subjects
GPQA7%WeightedExpert-level questions in biology, physics, and chemistry
GPQA-DDisplay onlyProvider-table reference for GPQA Diamond scores reported in first-party comparison charts.
SuperGPQA7%WeightedEnhanced version covering 285 disciplines
MMLU-Pro20%WeightedHarder version of MMLU with 10 answer choices and more reasoning-focused questions
HLE35%WeightedExtremely difficult questions contributed by domain experts worldwide to test frontier AI
FrontierScienceDisplay onlyResearch-level science and scientific reasoning benchmark
HLE w/o tools10%WeightedTool-free variant of Humanity's Last Exam used to isolate raw frontier reasoning without external aids
SimpleQA5%WeightedFactual question answering benchmark
HealthBench HardDisplay onlyA harder health reasoning benchmark subset used in first-party frontier model comparisons.
HealthBench ProfessionalDisplay onlyAn open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.
MedXpertQA (Text)Display onlyMedical multiple-choice benchmark covering many specialties with text-only questions.
FrontierScience ResearchDisplay onlyA research-oriented FrontierScience variant focused on scientific investigation and solution quality.
MMLU-Pro (Arcee)Display onlyDisplay-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.
MMMLUDisplay onlyA multilingual MMLU-style benchmark reported in provider evaluation tables.

About Knowledge Benchmarks

Tests knowledge across 57 academic subjects

Common questions

What is the best LLM for knowledge tasks?

The top LLMs for knowledge tasks are ranked by benchmarks like MMLU and GPQA, which test factual accuracy and expert-level understanding across dozens of subjects.

What is MMLU and how does it measure LLM knowledge?

MMLU (Massive Multitask Language Understanding) tests LLMs across 57 subjects from STEM to humanities, measuring broad factual knowledge and reasoning at varying difficulty levels.

What benchmarks test knowledge in AI models?

Key knowledge benchmarks include MMLU, MMLU-Pro, GPQA, SuperGPQA, HLE, and FrontierScience, each evaluating different depths of factual and scientific understanding.

How do knowledge benchmarks differ from reasoning benchmarks?

Knowledge benchmarks focus on factual recall and domain expertise, while reasoning benchmarks test logical deduction and multi-step problem solving independent of specific facts.

Knowledge benchmark updates

Know which model knows the most — updated every week.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Related