Skip to main content

Benchmark profile

Artificial Analysis Omniscience Hallucination Rate (AA-Omniscience Hallucination Rate)

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Data verified

Benchmark score on AA-Omniscience Hallucination Rate — July 29, 2026

BenchLM mirrors the published score view for AA-Omniscience Hallucination Rate. Command A+ leads the public snapshot at 14.1% , followed by MiniMax M3 (16.1%) and Qwen3.7 Max (22.9%). BenchLM does not use these results to rank models overall.

154 modelsKnowledgeCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (154 models)

Score
1
Command A+Cohere · Open weight
14.1%
2
MiniMax M3MiniMax · Open weight
16.1%
3
Qwen3.7 MaxAlibaba · Closed
22.9%
4
MiMo-V2.5-ProXiaomi · Closed
24.5%
5
Grok 4.3xAI · Closed
25.0%
6
Qwen3.7 PlusAlibaba · Closed
25.5%
7
GLM-5.2Z.AI · Open weight
28.1%
8
Nemotron 3 UltraNVIDIA · Open weight
28.5%
9
GLM-5.1Z.AI · Open weight
29.4%
10
MiMo-V2-ProXiaomi · Closed
29.9%
11
Gemma 4 E4BGoogle · Open weight
31.3%
12
Qwen3.6 PlusAlibaba · Closed
32.0%
13
Gemma 4 E2BGoogle · Open weight
32.9%
14
Gemini 3.5 Flash-LiteGoogle · Closed
33.5%
15
GLM-5Z.AI · Open weight
34.0%
16
MiniMax M2.7MiniMax · Open weight
34.4%
17
Claude Opus 4.8Anthropic · Closed
35.9%
18
Claude Opus 4.7 (Adaptive)Anthropic · Closed
36.2%
19
Claude Sonnet 5Anthropic · Closed
37.3%
20
GPT-4oOpenAI · Closed
37.9%
21
Muse Spark 1.1Meta · Closed
38.1%
22
Kimi K2.6Moonshot AI · Open weight
39.3%
23
Claude 4 SonnetAnthropic · Closed
40.8%
24
Qwen 3.6 Max (preview)Alibaba · Closed
44.2%
25
MiMo-V2-OmniXiaomi · Closed
44.4%
26
LFM2.5-8B-A1BLiquidAI · Open weight
47.0%
27
Qwen3.6-27BAlibaba · Open weight
48.3%
28
Qwen3.6-35B-A3BAlibaba · Open weight
49.7%
29
Gemini 3.1 ProGoogle · Closed
49.9%
30
Claude Opus 5Anthropic · Closed
50.1%
31
Kimi K3Moonshot AI · Closed
50.9%
32
Llama 3.1 405BMeta · Open weight
51.0%
33
GPT-5.1OpenAI · Closed
51.3%
34
Claude Opus 4.7Anthropic · Closed
51.9%
35
Gemini 3.6 FlashGoogle · Closed
53.5%
36
Grok 4.5xAI · Closed
53.5%
37
GPT-5 miniOpenAI · Closed
54.1%
38
Claude Fable 5Anthropic · Closed
54.9%
39
GPT-5 nanoOpenAI · Closed
56.3%
40
Claude Opus 4.5 ThinkingAnthropic · Closed
59.8%
41
Gemini 3.5 FlashGoogle · Closed
60.7%
42
Mistral Medium 3Mistral · Closed
60.9%
43
Claude Opus 4.6 (Adaptive)Anthropic · Closed
61.3%
44
GLM-5-TurboZ.AI · Closed
62.2%
45
InklingThinking Machines Lab · Open weight
63.1%
46
Grok 4xAI · Closed
64.2%
47
Kimi K2.5Moonshot AI · Open weight
64.6%
48
Kimi K2.5 (Reasoning)Moonshot AI · Closed
64.6%
49
Claude Sonnet 4.6Anthropic · Closed
65.9%
50
66.0%
51
GLM-4.6Z.AI · Open weight
66.1%
52
Mistral Small 4Mistral · Open weight
66.8%
53
Mistral Small 4 (Reasoning)Mistral · Open weight
66.8%
54
Mistral Large 2Mistral · Closed
67.8%
55
GLM-5V-TurboZ.AI · Closed
67.9%
56
o1OpenAI · Closed
69.3%
57
LFM2-24B-A2BLiquidAI · Closed
70.0%
58
72.4%
59
GPT-5.2-CodexOpenAI · Closed
72.8%
60
Hy3 PreviewTencent · Open weight
73.0%
61
Hy3Tencent · Open weight
73.0%
62
Muse SparkMeta · Closed
73.2%
63
GPT-5.4 nanoOpenAI · Closed
73.6%
64
Kimi K2Moonshot AI · Closed
74.2%
65
GPT-5.1-Codex-MaxOpenAI · Closed
74.4%
66
GPT-5.1-CodexOpenAI · Closed
74.4%
67
MiMo-V2-FlashXiaomi · Open weight
75.1%
68
Claude Opus 4.5Anthropic · Closed
75.4%
69
Claude Opus 4.6Anthropic · Closed
76.0%
70
Ministral 3 3B (Reasoning)Mistral · Open weight
77.5%
71
Ministral 3 3BMistral · Open weight
77.5%
72
Granite-4.0-350MIBM · Open weight
77.8%
73
Nova ProAmazon · Closed
77.9%
74
Claude 3 HaikuAnthropic · Closed
78.2%
75
Llama 4 ScoutMeta · Open weight
78.3%
76
Grok Code Fast 1xAI · Closed
78.5%
77
GPT-4.1OpenAI · Closed
79.6%
78
Qwen3.5-27BAlibaba · Open weight
79.7%
79
GPT-5.2OpenAI · Closed
79.7%
80
GPT-5 (medium)OpenAI · Closed
80.1%
81
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
80.3%
82
Kimi K2.7 CodeMoonshot AI · Open weight
80.3%
83
GPT-4.1 nanoOpenAI · Closed
80.4%
84
Phi-4Microsoft · Open weight
80.5%
85
Gemma 4 12BGoogle · Open weight
80.8%
86
Gemma 4 26B A4BGoogle · Open weight
80.9%
87
Exaone 4.0 32BLG AI Research · Open weight
81.0%
88
Gemini 3.1 Flash-LiteGoogle · Closed
81.6%
89
Gemma 4 31BGoogle · Open weight
81.6%
90
Nemotron Ultra 253BNVIDIA · Open weight
81.7%
91
Grok 4.1 FastxAI · Closed
81.8%
92
Mistral Medium 3.5 128BMistral · Open weight
82.0%
93
GPT-4.1 miniOpenAI · Closed
82.0%
94
GPT-5 (high)OpenAI · Closed
82.1%
95
Nemotron 3 Nano 30BNVIDIA · Open weight
82.9%
96
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
83.1%
97
Granite-4.0-H-1BIBM · Open weight
83.4%
98
DeepSeek V3.1DeepSeek · Open weight
83.5%
99
Mistral Large 3Mistral · Closed
83.7%
100
Qwen3.5-35B-A3BAlibaba · Open weight
84.0%
101
DeepSeek-R1DeepSeek · Open weight
84.0%
102
Step 3.7 FlashStepFun · Open weight
84.4%
103
LFM2.5-1.2B-InstructLiquidAI · Closed
84.8%
104
GPT-5.6 TerraOpenAI · Closed
85.2%
105
GPT-5.5OpenAI · Closed
85.5%
106
Qwen3.5-122B-A10BAlibaba · Open weight
85.5%
107
Trinity-Large-PreviewArcee AI · Open weight
86.6%
108
Trinity-Large-ThinkingArcee AI · Open weight
86.6%
109
MiniMax M1 80kMiniMax · Closed
86.8%
110
GPT-5.3 CodexOpenAI · Closed
86.9%
111
GPT-5.3-Codex-SparkOpenAI · Closed
86.9%
112
Nemotron 3 Super 120B A12BNVIDIA · Open weight
87.0%
113
o3OpenAI · Closed
87.1%
114
Llama 4 MaverickMeta · Open weight
87.3%
115
Gemini 2.5 ProGoogle · Closed
87.4%
116
GPT-5.4OpenAI · Closed
88.6%
117
DeepSeek V4 Pro (High)DeepSeek · Open weight
88.6%
118
GPT-5.6 SolOpenAI · Closed
88.8%
119
Ministral 3 8B (Reasoning)Mistral · Open weight
89.0%
120
Ministral 3 8BMistral · Open weight
89.0%
121
Qwen3.5 397BAlibaba · Open weight
89.1%
122
Qwen3.5 397B (Reasoning)Alibaba · Open weight
89.1%
123
K-ExaoneLG AI Research · Closed
89.1%
124
MiniMax M2.5MiniMax · Closed
89.3%
125
DeepSeek V3DeepSeek · Open weight
89.4%
126
Qwen3 MaxAlibaba · Closed
89.4%
127
Gemma 3 27BGoogle · Open weight
89.5%
128
GLM-4.7-FlashZ.AI · Open weight
89.5%
129
DeepSeek V4 Flash (High)DeepSeek · Open weight
89.7%
130
GPT-5.4 miniOpenAI · Closed
89.8%
131
GPT-5.6 LunaOpenAI · Closed
90.1%
132
Gemini 3 FlashGoogle · Closed
90.2%
133
Ministral 3 14B (Reasoning)Mistral · Open weight
90.2%
134
Ministral 3 14BMistral · Open weight
90.2%
135
GLM-4.7Z.AI · Open weight
90.3%
136
Gemini 3 ProGoogle · Closed
90.9%
137
GPT-OSS 120BOpenAI · Open weight
91.2%
138
Exaone 4.0 1.2BLG AI Research · Open weight
91.5%
139
Solar Pro 2Upstage · Closed
91.5%
140
Mercury 2Inception · Closed
91.5%
141
Step 3.5 FlashStepFun · Open weight
91.6%
142
GLM-4.5-AirZ.AI · Closed
92.3%
143
Gemini 2.5 FlashGoogle · Closed
93.3%
144
DeepSeek V3.2DeepSeek · Open weight
93.5%
145
Sarvam 105BSarvam · Open weight
93.5%
146
Granite-4.0-1BIBM · Open weight
93.5%
147
DeepSeek V4 Pro (Max)DeepSeek · Open weight
94.0%
148
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
94.0%
149
GPT-OSS 20BOpenAI · Open weight
94.1%
150
Granite-4.0-H-350MIBM · Open weight
94.4%
151
DeepSeek V4 Flash (Max)DeepSeek · Open weight
95.8%
152
Ling 2.6 FlashInclusionAI · Open weight
95.8%
153
LFM2.5-1.2B-ThinkingLiquidAI · Closed
96.9%
154
Sarvam 30BSarvam · Open weight
97.0%

The published AA-Omniscience Hallucination Rate snapshot places Command A+ first at 14.1%. The third row is 8.8 points higher. The broader top-10 range is 15.8 points, so the table still separates the published systems.

154 models have been evaluated on AA-Omniscience Hallucination Rate. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-Omniscience Hallucination Rate is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-Omniscience Hallucination Rate

Year

2026

Tasks

Knowledge questions

Format

Hallucination rate

Difficulty

Factuality

BenchLM marks this row lower-is-better because a lower hallucination rate is preferable, even though the OpenRouter card displays the raw percentage.

BenchLM freshness & provenance

Version

AA-Omniscience Hallucination Rate 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-Omniscience Hallucination Rate measure?

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Which model scores highest on AA-Omniscience Hallucination Rate?

Command A+ by Cohere currently leads with a score of 14.1% on AA-Omniscience Hallucination Rate.

How many models are evaluated on AA-Omniscience Hallucination Rate?

154 AI models have been evaluated on AA-Omniscience Hallucination Rate on BenchLM.

Last updated: July 29, 2026 · BenchLM version AA-Omniscience Hallucination Rate 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.