Skip to main content
BenchLM
Data
Recommendation

Best Factuality AI Models in 2026

As of October 5, 2026, the top model in best factuality ai models on the BenchLM leaderboard is Claude Opus 5.5 with a score of 64.4.

Bottom line: Factuality benchmarks are intentionally narrow — SimpleQA and HLE-no-tools are the primary signals. The category is still maturing, so treat the ranking as a shortlist.

Ranking data as of

Full Rankings (44 models)

1
Claude Opus 5.5
Anthropic·Proprietary·1M

64.4

sourced avg

2
Claude Fable 5.1
Anthropic·Proprietary·1M

60.9

sourced avg

3
Claude Mythos 5
Anthropic·Proprietary·1M+

59

sourced avg

4
DeepSeek V4 Pro 0813
DeepSeek·Open Weight·1M

57.9

sourced avg

5
Claude Sonnet 5.5
Anthropic·Proprietary·1M

56.9

sourced avg

6
Claude Opus 5
Anthropic·Proprietary·

56.3

sourced avg

7
Muse Spark 1.1
Meta·Proprietary·1M

52.2

sourced avg

8
Sakana Fugu-Ultra
Sakana AI·Proprietary·1M

50

sourced avg

9
Claude Opus 4.8
Anthropic·Proprietary·1M

49.8

sourced avg

10
Pareto 26.9
Unbiased·Proprietary·N/A

49

sourced avg

11
Sakana Fugu
Sakana AI·Proprietary·1M

47.2

sourced avg

12
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

46.9

sourced avg

13
Gemini 3.1 Pro
Google·Proprietary·1M

45.4

sourced avg

14
Ornith-1.5-397B
Ornith AI·Open Weight·262K

44.6

sourced avg

15
Qwen3.8 Max
Alibaba·Open Weight·1M

43.6

sourced avg

16
Kimi K3
Moonshot AI·Pending·1.05M

43.5

sourced avg

17
Hy4 preview
Tencent·Open Weight·1M

43.4

sourced avg

18
Claude Sonnet 5
Anthropic·Proprietary·1M

43.2

sourced avg

19
GPT-5.5 Pro
OpenAI·Proprietary·1M

43.1

sourced avg

20
Muse Spark
Meta·Proprietary·262K

42.8

sourced avg

21
GPT-5.4 Pro
OpenAI·Proprietary·1.05M

42.7

sourced avg

22
GPT-5.5
OpenAI·Proprietary·1M

41.4

sourced avg

23
GLM-5.2
Z.AI·Open Weight·1M

40.5

sourced avg

24
Claude Opus 4.6
Anthropic·Proprietary·1M

40

sourced avg

25
GPT-5.4
OpenAI·Proprietary·1.05M

39.8

sourced avg

26
Beam
Reflection AI·Pending·

36.2

sourced avg

27
Qwen3.8-Flash-Next
Alibaba·Open Weight·262K

35.9

sourced avg

28
DeepSeek V4 Flash 0731
DeepSeek·Open Weight·1M

34.1

sourced avg

29
MiMo-V2.5-Pro
Xiaomi·Proprietary·1M

34

sourced avg

30
Grok 4.20
xAI·Proprietary·2M

31.6

sourced avg

31
Inkling-Small
Thinking Machines Lab·Open Weight·1M

31.6

sourced avg

32
MAI-Thinking-1
Microsoft·Proprietary·256K

31

sourced avg

33
Qwen3.8-27B
Alibaba·Open Weight·262K

30.8

sourced avg

34
Inkling
Thinking Machines Lab·Open Weight·1M

30

sourced avg

35
Solar Open 2
Upstage·Open Weight·1M

28.8

sourced avg

36
GPT-5.4 mini
OpenAI·Proprietary·400K

28.2

sourced avg

37
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

26.7

sourced avg

38
Ornith-1.5-35B-A3B
Ornith AI·Open Weight·262K

25.6

sourced avg

39
GPT-5.4 nano
OpenAI·Proprietary·400K

24.3

sourced avg

40
Ornith-1.5-9B
Ornith AI·Open Weight·262K

20.2

sourced avg

41
Gemma 4 31B
Google·Open Weight·256K

19.5

sourced avg

42

10.5

sourced avg

43
Gemma 4 26B A4B
Google·Open Weight·256K

8.7

sourced avg

44
Gemma 4 12B
Google·Open Weight·256K

5.2

sourced avg

How to choose

Key Takeaways

The top model on this sourced reporting-family slice is Claude Opus 5.5 by Anthropic with an average of 64.4.

The best open-weight model is DeepSeek V4 Pro 0813 at position #4.

44 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

This is a reporting family ranking, not a weighted category. It averages sourced factuality benchmarks to give a focused view of this capability.

Known limitations

Models must have sourced results on at least a quarter of the benchmarks in this family to be included. Coverage varies — a model with 2 benchmark scores is less reliable than one with 5.

About this ranking

Ranking data as of October 5, 2026

This reporting page is intentionally narrow. It focuses on currently tracked sourced factuality signals such as SimpleQA, HLE without tools, and multimodal factuality. It is a reporting page, not a mature weighted category.

This page ranks models using only sourced factuality benchmarks in the reporting family.

Claude Opus 5.5 leads this ranking with a score of 64.4, followed by Claude Fable 5.1 (60.9) and Claude Mythos 5 (59). There is meaningful separation between the top models, suggesting genuine performance differences.

The best open-weight option is DeepSeek V4 Pro 0813 (ranked #4 with a score of 57.9). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.

This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Explore More

Last updated: October 5, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.