Skip to main content
Radar

Provider changes are easy to miss. Radar watches releases, pricing, deprecations, and incidents at the source.Provider changes are easy to miss.

See Radar

Benchmark profile

Artificial Analysis Omniscience Hallucination Rate (AA-Omniscience Hallucination Rate)

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Data verified 25 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on AA-Omniscience Hallucination Rate — August 6, 2026

BenchLM mirrors the published score view for AA-Omniscience Hallucination Rate. Command A+ leads the public snapshot at 14.1% , followed by MiniMax M3 (16.1%) and Qwen3.7 Max (22.9%). BenchLM does not use these results to rank models overall.

155 modelsKnowledgeCurrentDisplay onlyUpdated August 6, 2026

Benchmark score table (155 models)

Score
1
Command A+Cohere · Open weight
14.1%
2
MiniMax M3MiniMax · Open weight
16.1%
3
Qwen3.7 MaxAlibaba · Closed
22.9%
4
MiMo-V2.5-ProXiaomi · Closed
24.5%
5
Grok 4.3xAI · Closed
25.0%
6
Qwen3.7 PlusAlibaba · Closed
25.5%
7
GLM-5.2Z.AI · Open weight
28.1%
8
Nemotron 3 UltraNVIDIA · Open weight
28.5%
9
GLM-5.1Z.AI · Open weight
29.4%
10
MiMo-V2-ProXiaomi · Closed
29.9%
11
Gemma 4 E4BGoogle · Open weight
31.3%
12
Qwen3.6 PlusAlibaba · Closed
32.0%
13
Gemma 4 E2BGoogle · Open weight
32.9%
14
Gemini 3.5 Flash-LiteGoogle · Closed
33.5%
15
GLM-5Z.AI · Open weight
34.0%
16
MiniMax M2.7MiniMax · Open weight
34.4%
17
Claude Opus 4.8Anthropic · Closed
35.9%
18
Claude Opus 4.7 (Adaptive)Anthropic · Closed
36.2%
19
Claude Sonnet 5Anthropic · Closed
37.3%
20
GPT-4oOpenAI · Closed
37.9%
21
Muse Spark 1.1Meta · Closed
38.1%
22
Kimi K2.6Moonshot AI · Open weight
39.3%
23
Claude 4 SonnetAnthropic · Closed
40.8%
24
Qwen 3.6 Max (preview)Alibaba · Closed
44.2%
25
MiMo-V2-OmniXiaomi · Closed
44.4%
26
LFM2.5-8B-A1BLiquidAI · Open weight
47.0%
27
Qwen3.6-27BAlibaba · Open weight
48.3%
28
Qwen3.6-35B-A3BAlibaba · Open weight
49.7%
29
Gemini 3.1 ProGoogle · Closed
49.9%
30
Claude Opus 5Anthropic · Closed
50.1%
31
Kimi K3Moonshot AI · Closed
50.9%
32
Llama 3.1 405BMeta · Open weight
51.0%
33
GPT-5.1OpenAI · Closed
51.3%
34
Claude Opus 4.7Anthropic · Closed
51.9%
35
Gemini 3.6 FlashGoogle · Closed
53.5%
36
Grok 4.5xAI · Closed
53.5%
37
GPT-5 miniOpenAI · Closed
54.1%
38
Claude Fable 5Anthropic · Closed
54.9%
39
GPT-5 nanoOpenAI · Closed
56.3%
40
Inkling-SmallThinking Machines Lab · Open weight
56.9%
41
Claude Opus 4.5 ThinkingAnthropic · Closed
59.8%
42
Gemini 3.5 FlashGoogle · Closed
60.7%
43
Mistral Medium 3Mistral · Closed
60.9%
44
Claude Opus 4.6 (Adaptive)Anthropic · Closed
61.3%
45
GLM-5-TurboZ.AI · Closed
62.2%
46
InklingThinking Machines Lab · Open weight
63.1%
47
Grok 4xAI · Closed
64.2%
48
Kimi K2.5Moonshot AI · Open weight
64.6%
49
Kimi K2.5 (Reasoning)Moonshot AI · Closed
64.6%
50
Claude Sonnet 4.6Anthropic · Closed
65.9%
51
66.0%
52
GLM-4.6Z.AI · Open weight
66.1%
53
Mistral Small 4Mistral · Open weight
66.8%
54
Mistral Small 4 (Reasoning)Mistral · Open weight
66.8%
55
Mistral Large 2Mistral · Closed
67.8%
56
GLM-5V-TurboZ.AI · Closed
67.9%
57
o1OpenAI · Closed
69.3%
58
LFM2-24B-A2BLiquidAI · Closed
70.0%
59
72.4%
60
GPT-5.2-CodexOpenAI · Closed
72.8%
61
Hy3 PreviewTencent · Open weight
73.0%
62
Hy3Tencent · Open weight
73.0%
63
Muse SparkMeta · Closed
73.2%
64
GPT-5.4 nanoOpenAI · Closed
73.6%
65
Kimi K2Moonshot AI · Closed
74.2%
66
GPT-5.1-Codex-MaxOpenAI · Closed
74.4%
67
GPT-5.1-CodexOpenAI · Closed
74.4%
68
MiMo-V2-FlashXiaomi · Open weight
75.1%
69
Claude Opus 4.5Anthropic · Closed
75.4%
70
Claude Opus 4.6Anthropic · Closed
76.0%
71
Ministral 3 3B (Reasoning)Mistral · Open weight
77.5%
72
Ministral 3 3BMistral · Open weight
77.5%
73
Granite-4.0-350MIBM · Open weight
77.8%
74
Nova ProAmazon · Closed
77.9%
75
Claude 3 HaikuAnthropic · Closed
78.2%
76
Llama 4 ScoutMeta · Open weight
78.3%
77
Grok Code Fast 1xAI · Closed
78.5%
78
GPT-4.1OpenAI · Closed
79.6%
79
Qwen3.5-27BAlibaba · Open weight
79.7%
80
GPT-5.2OpenAI · Closed
79.7%
81
GPT-5 (medium)OpenAI · Closed
80.1%
82
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
80.3%
83
Kimi K2.7 CodeMoonshot AI · Open weight
80.3%
84
GPT-4.1 nanoOpenAI · Closed
80.4%
85
Phi-4Microsoft · Open weight
80.5%
86
Gemma 4 12BGoogle · Open weight
80.8%
87
Gemma 4 26B A4BGoogle · Open weight
80.9%
88
Exaone 4.0 32BLG AI Research · Open weight
81.0%
89
Gemini 3.1 Flash-LiteGoogle · Closed
81.6%
90
Gemma 4 31BGoogle · Open weight
81.6%
91
Nemotron Ultra 253BNVIDIA · Open weight
81.7%
92
Grok 4.1 FastxAI · Closed
81.8%
93
Mistral Medium 3.5 128BMistral · Open weight
82.0%
94
GPT-4.1 miniOpenAI · Closed
82.0%
95
GPT-5 (high)OpenAI · Closed
82.1%
96
Nemotron 3 Nano 30BNVIDIA · Open weight
82.9%
97
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
83.1%
98
Granite-4.0-H-1BIBM · Open weight
83.4%
99
DeepSeek V3.1DeepSeek · Open weight
83.5%
100
Mistral Large 3Mistral · Closed
83.7%
101
Qwen3.5-35B-A3BAlibaba · Open weight
84.0%
102
DeepSeek-R1DeepSeek · Open weight
84.0%
103
DeepSeek V4 Flash (Max)DeepSeek · Closed
84.4%
104
Step 3.7 FlashStepFun · Open weight
84.4%
105
LFM2.5-1.2B-InstructLiquidAI · Closed
84.8%
106
GPT-5.6 TerraOpenAI · Closed
85.2%
107
GPT-5.5OpenAI · Closed
85.5%
108
Qwen3.5-122B-A10BAlibaba · Open weight
85.5%
109
Trinity-Large-PreviewArcee AI · Open weight
86.6%
110
Trinity-Large-ThinkingArcee AI · Open weight
86.6%
111
MiniMax M1 80kMiniMax · Closed
86.8%
112
GPT-5.3 CodexOpenAI · Closed
86.9%
113
GPT-5.3-Codex-SparkOpenAI · Closed
86.9%
114
Nemotron 3 Super 120B A12BNVIDIA · Open weight
87.0%
115
o3OpenAI · Closed
87.1%
116
Llama 4 MaverickMeta · Open weight
87.3%
117
Gemini 2.5 ProGoogle · Closed
87.4%
118
GPT-5.4OpenAI · Closed
88.6%
119
DeepSeek V4 Pro (High)DeepSeek · Open weight
88.6%
120
GPT-5.6 SolOpenAI · Closed
88.8%
121
Ministral 3 8B (Reasoning)Mistral · Open weight
89.0%
122
Ministral 3 8BMistral · Open weight
89.0%
123
Qwen3.5 397BAlibaba · Open weight
89.1%
124
Qwen3.5 397B (Reasoning)Alibaba · Open weight
89.1%
125
K-ExaoneLG AI Research · Closed
89.1%
126
MiniMax M2.5MiniMax · Closed
89.3%
127
DeepSeek V3DeepSeek · Open weight
89.4%
128
Qwen3 MaxAlibaba · Closed
89.4%
129
Gemma 3 27BGoogle · Open weight
89.5%
130
GLM-4.7-FlashZ.AI · Open weight
89.5%
131
GPT-5.4 miniOpenAI · Closed
89.8%
132
GPT-5.6 LunaOpenAI · Closed
90.1%
133
Gemini 3 FlashGoogle · Closed
90.2%
134
Ministral 3 14B (Reasoning)Mistral · Open weight
90.2%
135
Ministral 3 14BMistral · Open weight
90.2%
136
GLM-4.7Z.AI · Open weight
90.3%
137
Gemini 3 ProGoogle · Closed
90.9%
138
GPT-OSS 120BOpenAI · Open weight
91.2%
139
Exaone 4.0 1.2BLG AI Research · Open weight
91.5%
140
Solar Pro 2Upstage · Closed
91.5%
141
Mercury 2Inception · Closed
91.5%
142
Step 3.5 FlashStepFun · Open weight
91.6%
143
GLM-4.5-AirZ.AI · Closed
92.3%
144
Gemini 2.5 FlashGoogle · Closed
93.3%
145
DeepSeek V3.2DeepSeek · Open weight
93.5%
146
Sarvam 105BSarvam · Open weight
93.5%
147
Granite-4.0-1BIBM · Open weight
93.5%
148
Celeris-1Celeris · Closed
93.8%
149
DeepSeek V4 Pro (Max)DeepSeek · Open weight
94.0%
150
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
94.0%
151
GPT-OSS 20BOpenAI · Open weight
94.1%
152
Granite-4.0-H-350MIBM · Open weight
94.4%
153
Ling 2.6 FlashInclusionAI · Open weight
95.8%
154
LFM2.5-1.2B-ThinkingLiquidAI · Closed
96.9%
155
Sarvam 30BSarvam · Open weight
97.0%

The published AA-Omniscience Hallucination Rate snapshot places Command A+ first at 14.1%. The third row is 8.8 points higher. The broader top-10 range is 15.8 points, so the table still separates the published systems.

155 models have been evaluated on AA-Omniscience Hallucination Rate. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-Omniscience Hallucination Rate is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-Omniscience Hallucination Rate

Year

2026

Tasks

Knowledge questions

Format

Hallucination rate

Difficulty

Factuality

BenchLM marks this row lower-is-better because a lower hallucination rate is preferable, even though the OpenRouter card displays the raw percentage.

BenchLM freshness & provenance

Version

AA-Omniscience Hallucination Rate 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-Omniscience Hallucination Rate measure?

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Which model scores highest on AA-Omniscience Hallucination Rate?

Command A+ by Cohere currently leads with a score of 14.1% on AA-Omniscience Hallucination Rate.

How many models are evaluated on AA-Omniscience Hallucination Rate?

155 AI models have been evaluated on AA-Omniscience Hallucination Rate on BenchLM.

Last updated: August 6, 2026 · BenchLM version AA-Omniscience Hallucination Rate 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.