Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Artificial Analysis Omniscience Hallucination Rate (AA-Omniscience Hallucination Rate)

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on AA-Omniscience Hallucination Rate — September 10, 2026

We mirror the published score view for AA-Omniscience Hallucination Rate. Command A+ leads the public snapshot at 14.2%, followed by LFM2.5-2.6B (16.0%) and MiniMax M3 (18.4%). We do not use these results to rank models overall.

190 modelsKnowledgeCurrentDisplay onlyUpdated September 10, 2026

Benchmark score table (190 models)

Score
1
Command A+Cohere · Open weight
14.2%
2
LFM2.5-2.6BLiquidAI · Open weight
16.0%
3
MiniMax M3MiniMax · Open weight
18.4%
4
Quasar 438BMultiverse Computing · Closed
21.4%
5
MiniCPM5-2BOpenBMB · Open weight
21.9%
6
Ling 3.0 Flash VLInclusionAI · Open weight
22.0%
7
K-EXAONE 2.0LG AI Research · Open weight
22.6%
8
Solar Pro 4Upstage · Closed
24.4%
9
MiMo-V2.5-ProXiaomi · Closed
24.7%
10
Grok 4.3xAI · Closed
25.0%
11
Qwen3.7 MaxAlibaba · Closed
25.6%
12
Granite 4.2 30BIBM · Open weight
25.6%
13
GLM-5.2Z.AI · Open weight
26.3%
14
Granite 4.2 3BIBM · Open weight
26.3%
15
Qwen3.7 PlusAlibaba · Closed
27.7%
16
GLM-5.3Z.AI · Open weight
29.6%
17
Nemotron 3 UltraNVIDIA · Open weight
29.7%
18
GLM-5.1Z.AI · Open weight
29.9%
19
MiMo-V2-ProXiaomi · Closed
30.0%
20
Qwen3.8-27BAlibaba · Open weight
30.3%
21
Ling 3.0 TinyInclusionAI · Open weight
30.5%
22
Gemma 4 E4BGoogle · Open weight
30.9%
23
Granite 4.2 8BIBM · Open weight
32.0%
24
Gemma 4 E2BGoogle · Open weight
32.4%
25
Muse Spark 1.3Meta · Closed
32.9%
26
A.X K2SK Telecom · Open weight
33.0%
27
Muse Spark 1.2Meta · Closed
33.3%
28
Grok 4.6xAI · Closed
34.3%
29
Gemini 3.5 Flash-LiteGoogle · Closed
34.4%
30
Qwen3.6 PlusAlibaba · Closed
34.6%
31
GLM-5Z.AI · Open weight
35.3%
32
MiniMax M2.7MiniMax · Open weight
35.6%
33
37.6%
34
GPT-4oOpenAI · Closed
37.9%
35
Claude Opus 4.8Anthropic · Closed
39.3%
36
Claude Sonnet 5Anthropic · Closed
39.4%
37
Kimi K2.6Moonshot AI · Open weight
40.5%
38
Claude 4 SonnetAnthropic · Closed
41.0%
39
Qwen3.8 Max PreviewAlibaba · Closed
41.7%
40
Claude Opus 4.7 (Adaptive)Anthropic · Closed
42.3%
41
Ling 3.0 FlashInclusionAI · Open weight
44.1%
42
Ling 3.0 Flash FP8InclusionAI · Open weight
44.1%
43
Qwen3.8-Flash-NextAlibaba · Open weight
45.3%
44
Qwen 3.6 Max (preview)Alibaba · Closed
46.2%
45
LFM2.5-8B-A1BLiquidAI · Open weight
46.9%
46
MiMo-V2-OmniXiaomi · Closed
48.9%
47
Qwen3.6-27BAlibaba · Open weight
49.3%
48
Muse Spark 1.1Meta · Closed
50.0%
49
Qwen3.6-35B-A3BAlibaba · Open weight
50.5%
50
Gemini 3.1 ProGoogle · Closed
50.9%
51
GPT-6 AstraOpenAI · Closed
51.3%
52
GPT-5.1OpenAI · Closed
51.9%
53
Llama 3.1 405BMeta · Open weight
52.4%
54
Kimi K3Moonshot AI · Closed
53.2%
55
Grok 4.5xAI · Closed
54.1%
56
Claude Opus 4.7Anthropic · Closed
54.1%
57
Gemini 3.8 FlashGoogle · Closed
55.2%
58
Gemini 3.6 FlashGoogle · Closed
55.6%
59
GPT-5 miniOpenAI · Closed
56.4%
60
GPT-5 nanoOpenAI · Closed
59.0%
61
Gemini 3.5 FlashGoogle · Closed
60.7%
62
Claude Opus 5Anthropic · Closed
60.8%
63
Mistral Medium 3Mistral · Closed
60.9%
64
Claude Opus 4.5 ThinkingAnthropic · Closed
61.0%
65
GLM-5-TurboZ.AI · Closed
62.6%
66
Claude Opus 4.6 (Adaptive)Anthropic · Closed
62.8%
67
Inkling-SmallThinking Machines Lab · Open weight
63.0%
68
Claude Fable 5Anthropic · Closed
63.6%
69
Gemini 3.7 FlashGoogle · Closed
64.5%
70
Grok 4xAI · Closed
64.5%
71
Kimi K2.5Moonshot AI · Open weight
65.7%
72
Kimi K2.5 (Reasoning)Moonshot AI · Closed
65.7%
73
Mistral Small 4Mistral · Open weight
66.5%
74
Mistral Small 4 (Reasoning)Mistral · Open weight
66.5%
75
Mercury 2.5Inception · Closed
67.0%
76
GLM-4.6Z.AI · Open weight
67.6%
77
InklingThinking Machines Lab · Open weight
67.7%
78
Mistral Large 2Mistral · Closed
67.7%
79
68.3%
80
Claude Sonnet 4.6Anthropic · Closed
68.5%
81
GLM-5V-TurboZ.AI · Closed
68.8%
82
LFM2-24B-A2BLiquidAI · Closed
69.0%
83
o1OpenAI · Closed
69.6%
84
Claude Fable 5.1Anthropic · Closed
72.6%
85
Hy3 PreviewTencent · Open weight
73.0%
86
GPT-5.2-CodexOpenAI · Closed
73.4%
87
73.4%
88
Hy3Tencent · Open weight
74.1%
89
GPT-5.4 nanoOpenAI · Closed
74.2%
90
MiMo-V2-FlashXiaomi · Open weight
76.0%
91
Granite-4.0-350MIBM · Open weight
76.1%
92
Claude Opus 4.5Anthropic · Closed
76.2%
93
Kimi K2Moonshot AI · Closed
76.6%
94
GPT-5.1-CodexOpenAI · Closed
77.2%
95
GPT-5.1-Codex-MaxOpenAI · Closed
77.2%
96
Nova ProAmazon · Closed
77.7%
97
Apodex 1.1Apodex · Closed
78.4%
98
Apodex 1.1 MiniApodex · Open weight
78.4%
99
Grok Code Fast 1xAI · Closed
79.3%
100
Llama 4 ScoutMeta · Open weight
79.4%
101
Claude Opus 4.6Anthropic · Closed
80.1%
102
Ministral 3 3B (Reasoning)Mistral · Open weight
80.2%
103
Ministral 3 3BMistral · Open weight
80.2%
104
Claude 3 HaikuAnthropic · Closed
80.5%
105
Gemma 4 12BGoogle · Open weight
81.0%
106
GPT-5.2OpenAI · Closed
81.2%
107
Phi-4Microsoft · Open weight
81.2%
108
Nemotron Ultra 253BNVIDIA · Open weight
81.2%
109
Qwen3.5-27BAlibaba · Open weight
81.5%
110
Mistral Medium 3.5 128BMistral · Open weight
81.6%
111
Granite-4.0-H-1BIBM · Open weight
81.7%
112
Muse Glimmer 30BMeta · Open weight
81.9%
113
Exaone 4.0 32BLG AI Research · Open weight
82.1%
114
GPT-5 (high)OpenAI · Closed
82.2%
115
Grok 4.1 FastxAI · Closed
82.3%
116
Kimi K2.7 CodeMoonshot AI · Open weight
82.4%
117
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
82.5%
118
GPT-4.1 nanoOpenAI · Closed
82.6%
119
Qwen3.5 397BAlibaba · Open weight
82.7%
120
Qwen3.5 397B (Reasoning)Alibaba · Open weight
82.7%
121
GPT-5 (medium)OpenAI · Closed
83.2%
122
North Mini CodeCohere · Open weight
83.2%
123
Nemotron 3 Nano 30BNVIDIA · Open weight
83.3%
124
DeepSeek-R1DeepSeek · Open weight
83.4%
125
Muse SparkMeta · Closed
84.2%
126
Gemma 4 31BGoogle · Open weight
85.0%
127
Step 3.7 FlashStepFun · Open weight
85.0%
128
LFM2.5-1.2B-InstructLiquidAI · Closed
85.1%
129
Qwen3.5-35B-A3BAlibaba · Open weight
85.4%
130
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
85.7%
131
DeepSeek V3.1DeepSeek · Open weight
85.7%
132
Trinity-Large-ThinkingArcee AI · Open weight
85.9%
133
Trinity-Large-PreviewArcee AI · Open weight
85.9%
134
Mistral Large 3Mistral · Closed
86.0%
135
Gemma 4 26B A4BGoogle · Open weight
86.4%
136
Nemotron 3 Super 100BNVIDIA · Open weight
87.0%
137
Nemotron 3 Super 120B A12BNVIDIA · Open weight
87.0%
138
Qwen3.5-122B-A10BAlibaba · Open weight
87.1%
139
GPT-5.6 TerraOpenAI · Closed
87.9%
140
o3OpenAI · Closed
88.1%
141
Granite-4.0-H-350MIBM · Open weight
88.1%
142
MiniMax M2.5MiniMax · Closed
88.1%
143
Solar Pro 3Upstage · Closed
88.2%
144
DeepSeek V4 Pro (High)DeepSeek · Open weight
88.7%
145
K-ExaoneLG AI Research · Closed
88.8%
146
Llama 4 MaverickMeta · Open weight
88.9%
147
GPT-5.5OpenAI · Closed
89.0%
148
Qwen3-Omni-30B-A3B-ThinkingAlibaba · Open weight
89.0%
149
GPT-5.3 CodexOpenAI · Closed
89.2%
150
GPT-5.3-Codex-SparkOpenAI · Closed
89.2%
151
Qwen3 MaxAlibaba · Closed
89.9%
152
DeepSeek V3DeepSeek · Open weight
90.0%
153
GPT-5.4 miniOpenAI · Closed
90.2%
154
Ultravox v0.6 Llama 3.3 70BFixie AI · Open weight
90.2%
155
GPT-OSS 120BOpenAI · Open weight
90.8%
156
Gemini 2.5 ProGoogle · Closed
90.9%
157
Mercury 2Inception · Closed
91.2%
158
Step 3.5 FlashStepFun · Open weight
91.3%
159
Gemini 3 ProGoogle · Closed
91.5%
160
MiniMax M1 80kMiniMax · Closed
91.5%
161
GPT-5.4OpenAI · Closed
91.7%
162
DeepSeek V4 Flash 0731DeepSeek · Closed
91.7%
163
Exaone 4.0 1.2BLG AI Research · Open weight
91.7%
164
Gemma 3 27BGoogle · Open weight
92.1%
165
GPT-5.6 SolOpenAI · Closed
92.2%
166
Gemini 3 FlashGoogle · Closed
92.4%
167
Ministral 3 14B (Reasoning)Mistral · Open weight
92.5%
168
Ministral 3 14BMistral · Open weight
92.5%
169
GPT-5.6 LunaOpenAI · Closed
92.6%
170
GPT-4.1 miniOpenAI · Closed
92.7%
171
Celeris-1Celeris · Closed
92.8%
172
GLM-4.5-AirZ.AI · Closed
92.9%
173
GLM-4.7Z.AI · Open weight
93.0%
174
Gemini 2.5 FlashGoogle · Closed
93.0%
175
Solar Pro 2Upstage · Closed
93.0%
176
DeepSeek V3.2DeepSeek · Open weight
93.3%
177
GPT-4.1OpenAI · Closed
93.3%
178
Sarvam 105BSarvam · Open weight
93.4%
179
Granite-4.0-1BIBM · Open weight
93.5%
180
GLM-4.7-FlashZ.AI · Open weight
93.9%
181
DeepSeek V4 Pro 0813DeepSeek · Closed
94.1%
182
GPT-OSS 20BOpenAI · Open weight
94.1%
183
Ministral 3 8B (Reasoning)Mistral · Open weight
94.2%
184
Ministral 3 8BMistral · Open weight
94.2%
185
LFM2.5-1.2B-ThinkingLiquidAI · Closed
94.4%
186
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
95.7%
187
Sarvam 30BSarvam · Open weight
96.3%
188
DeepSeek V4.1 FlashDeepSeek · Open weight
96.5%
189
Ling 2.6 FlashInclusionAI · Open weight
96.7%
190
Qwen3-Omni-30B-A3B-InstructAlibaba · Open weight
97.6%

The published AA-Omniscience Hallucination Rate snapshot places Command A+ first at 14.2%. The third row is 4.2 points higher. The broader top-10 range is 10.8 points, so the table still separates the published systems.

190 models have been evaluated on AA-Omniscience Hallucination Rate. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-Omniscience Hallucination Rate is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-Omniscience Hallucination Rate

Year

2026

Tasks

Knowledge questions

Format

Hallucination rate

Difficulty

Factuality

BenchLM marks this row lower-is-better because a lower hallucination rate is preferable, even though the OpenRouter card displays the raw percentage.

BenchLM freshness & provenance

Version

AA-Omniscience Hallucination Rate 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-Omniscience Hallucination Rate measure?

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Which model scores highest on AA-Omniscience Hallucination Rate?

Command A+ by Cohere currently leads with a score of 14.2% on AA-Omniscience Hallucination Rate.

How many models are evaluated on AA-Omniscience Hallucination Rate?

190 AI models have been evaluated on AA-Omniscience Hallucination Rate on BenchLM.

Last updated: September 10, 2026 · BenchLM version AA-Omniscience Hallucination Rate 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.