Skip to main content

Benchmark profile

Artificial Analysis Omniscience Accuracy (AA-Omniscience Accuracy)

A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.

Data verified

Benchmark score on AA-Omniscience Accuracy — July 29, 2026

BenchLM mirrors the published score view for AA-Omniscience Accuracy. Claude Fable 5 leads the public snapshot at 61.4% , followed by GPT-5.6 Sol (58.5%) and GPT-5.5 (56.9%). BenchLM does not use these results to rank models overall.

154 modelsKnowledgeCurrentDisplay onlyUpdated July 29, 2026

Benchmark score table (154 models)

Score
1
Claude Fable 5Anthropic · Closed
61.4%
2
GPT-5.6 SolOpenAI · Closed
58.5%
3
GPT-5.5OpenAI · Closed
56.9%
4
Gemini 3 ProGoogle · Closed
55.9%
5
Gemini 3.1 ProGoogle · Closed
55.3%
6
Claude Opus 5Anthropic · Closed
54.2%
7
Grok 4.5xAI · Closed
52.1%
8
Gemini 3.5 FlashGoogle · Closed
51.9%
9
GPT-5.3 CodexOpenAI · Closed
51.8%
10
GPT-5.3-Codex-SparkOpenAI · Closed
51.8%
11
Gemini 3.6 FlashGoogle · Closed
50.2%
12
GPT-5.4OpenAI · Closed
50.0%
13
Claude Opus 4.8Anthropic · Closed
46.6%
14
Claude Opus 4.6 (Adaptive)Anthropic · Closed
46.4%
15
Kimi K3Moonshot AI · Closed
46.0%
16
GPT-5.6 TerraOpenAI · Closed
45.9%
17
Claude Opus 4.7 (Adaptive)Anthropic · Closed
45.8%
18
Claude Opus 4.5 ThinkingAnthropic · Closed
45.7%
19
Gemini 3 FlashGoogle · Closed
45.5%
20
Claude Opus 4.6Anthropic · Closed
45.2%
21
Muse SparkMeta · Closed
44.6%
22
GPT-5.2OpenAI · Closed
43.8%
23
Claude Opus 4.7Anthropic · Closed
43.5%
24
DeepSeek V4 Pro (Max)DeepSeek · Open weight
43.3%
25
DeepSeek V4 Pro (High)DeepSeek · Open weight
41.8%
26
GPT-5.6 LunaOpenAI · Closed
41.5%
27
Grok 4xAI · Closed
41.4%
28
Claude Opus 4.5Anthropic · Closed
40.7%
29
GPT-5 (high)OpenAI · Closed
40.7%
30
GPT-5.2-CodexOpenAI · Closed
40.7%
31
Muse Spark 1.1Meta · Closed
40.6%
32
InklingThinking Machines Lab · Open weight
40.0%
33
GPT-5.1-Codex-MaxOpenAI · Closed
39.2%
34
GPT-5.1-CodexOpenAI · Closed
39.2%
35
Gemini 2.5 ProGoogle · Closed
39.0%
36
GPT-5 (medium)OpenAI · Closed
38.9%
37
Kimi K2.7 CodeMoonshot AI · Open weight
38.6%
38
o3OpenAI · Closed
38.4%
39
Claude Sonnet 5Anthropic · Closed
38.3%
40
Claude Sonnet 4.6Anthropic · Closed
38.0%
41
Qwen 3.6 Max (preview)Alibaba · Closed
37.7%
42
GPT-5.1OpenAI · Closed
37.6%
43
GPT-5.4 miniOpenAI · Closed
37.5%
44
DeepSeek V4 Flash (Max)DeepSeek · Open weight
37.2%
45
Gemini 3.1 Flash-LiteGoogle · Closed
36.4%
46
DeepSeek V4 Flash (High)DeepSeek · Open weight
35.5%
47
o1OpenAI · Closed
34.7%
48
Grok 4.3xAI · Closed
34.6%
49
Kimi K2.5Moonshot AI · Open weight
34.3%
50
Kimi K2.5 (Reasoning)Moonshot AI · Closed
34.3%
51
Kimi K2.6Moonshot AI · Open weight
32.8%
52
Hy3 PreviewTencent · Open weight
31.5%
53
Hy3Tencent · Open weight
31.5%
54
Qwen3.5 397BAlibaba · Open weight
31.4%
55
Qwen3.5 397B (Reasoning)Alibaba · Open weight
31.4%
56
DeepSeek-R1DeepSeek · Open weight
31.0%
57
Gemini 3.5 Flash-LiteGoogle · Closed
30.3%
58
Qwen3.7 MaxAlibaba · Closed
30.1%
59
GLM-4.7Z.AI · Open weight
29.3%
60
GLM-5V-TurboZ.AI · Closed
29.1%
61
GLM-5-TurboZ.AI · Closed
29.0%
62
DeepSeek V3.1 (Reasoning)DeepSeek · Open weight
28.8%
63
GLM-5Z.AI · Open weight
26.9%
64
Kimi K2Moonshot AI · Closed
26.8%
65
MiMo-V2-ProXiaomi · Closed
26.8%
66
Gemini 2.5 FlashGoogle · Closed
26.5%
67
Qwen3.6 PlusAlibaba · Closed
26.2%
68
MiniMax M2.5MiniMax · Closed
26.2%
69
MiniMax M2.7MiniMax · Open weight
26.1%
70
GPT-5.4 nanoOpenAI · Closed
25.4%
71
Step 3.7 FlashStepFun · Open weight
25.4%
72
DeepSeek V3DeepSeek · Open weight
25.4%
73
25.3%
74
GLM-5.2Z.AI · Open weight
25.1%
75
Mistral Medium 3.5 128BMistral · Open weight
25.1%
76
Step 3.5 FlashStepFun · Open weight
25.0%
77
Qwen3.5-122B-A10BAlibaba · Open weight
24.7%
78
Qwen3 MaxAlibaba · Closed
24.4%
79
Llama 4 MaverickMeta · Open weight
24.3%
80
GLM-5.1Z.AI · Open weight
24.2%
81
DeepSeek V3.2DeepSeek · Open weight
24.2%
82
GPT-4.1OpenAI · Closed
24.2%
83
Mistral Large 3Mistral · Closed
24.1%
84
GPT-5 miniOpenAI · Closed
24.0%
85
Nemotron 3 Super 120B A12BNVIDIA · Open weight
24.0%
86
Grok Code Fast 1xAI · Closed
23.8%
87
DeepSeek V3.1DeepSeek · Open weight
23.1%
88
Trinity-Large-PreviewArcee AI · Open weight
22.8%
89
Trinity-Large-ThinkingArcee AI · Open weight
22.8%
90
MiMo-V2.5-ProXiaomi · Closed
22.6%
91
22.6%
92
Claude 4 SonnetAnthropic · Closed
22.4%
93
Llama 3.1 405BMeta · Open weight
22.3%
94
Qwen3.7 PlusAlibaba · Closed
22.2%
95
Mistral Small 4Mistral · Open weight
22.1%
96
Mistral Small 4 (Reasoning)Mistral · Open weight
22.1%
97
Nemotron 3 UltraNVIDIA · Open weight
21.6%
98
GPT-OSS 120BOpenAI · Open weight
21.5%
99
MiniMax M1 80kMiniMax · Closed
21.1%
100
Qwen3.5-27BAlibaba · Open weight
21.0%
101
GLM-4.6Z.AI · Open weight
20.8%
102
Qwen3.5-35B-A3BAlibaba · Open weight
20.5%
103
Mercury 2Inception · Closed
20.5%
104
Mistral Large 2Mistral · Closed
20.1%
105
Gemma 4 31BGoogle · Open weight
19.9%
106
Nemotron Ultra 253BNVIDIA · Open weight
19.9%
107
GPT-4oOpenAI · Closed
19.7%
108
Qwen3.6-27BAlibaba · Open weight
19.2%
109
Qwen3.6-35B-A3BAlibaba · Open weight
18.9%
110
MiMo-V2-OmniXiaomi · Closed
18.7%
111
Mistral Medium 3Mistral · Closed
18.3%
112
GPT-5 nanoOpenAI · Closed
18.3%
113
Gemma 4 26B A4BGoogle · Open weight
18.2%
114
Sarvam 105BSarvam · Open weight
17.6%
115
GPT-4.1 miniOpenAI · Closed
17.5%
116
Claude 3 HaikuAnthropic · Closed
17.2%
117
Nemotron 3 Nano 30BNVIDIA · Open weight
17.1%
118
Grok 4.1 FastxAI · Closed
17.0%
119
Nova ProAmazon · Closed
17.0%
120
K-ExaoneLG AI Research · Closed
16.5%
121
Gemma 4 12BGoogle · Open weight
16.0%
122
GLM-4.7-FlashZ.AI · Open weight
15.9%
123
Solar Pro 2Upstage · Closed
15.6%
124
GLM-4.5-AirZ.AI · Closed
15.5%
125
GPT-OSS 20BOpenAI · Open weight
15.5%
126
Ling 2.6 FlashInclusionAI · Open weight
15.4%
127
MiMo-V2-FlashXiaomi · Open weight
15.2%
128
MiniMax M3MiniMax · Open weight
15.0%
129
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
14.8%
130
Llama 4 ScoutMeta · Open weight
14.6%
131
GPT-4.1 nanoOpenAI · Closed
13.3%
132
Phi-4Microsoft · Open weight
13.2%
133
Sarvam 30BSarvam · Open weight
12.7%
134
Gemma 3 27BGoogle · Open weight
12.5%
135
Ministral 3 14B (Reasoning)Mistral · Open weight
12.3%
136
Ministral 3 14BMistral · Open weight
12.3%
137
Ministral 3 8B (Reasoning)Mistral · Open weight
11.2%
138
Ministral 3 8BMistral · Open weight
11.2%
139
Exaone 4.0 32BLG AI Research · Open weight
10.4%
140
LFM2.5-8B-A1BLiquidAI · Open weight
9.4%
141
Command A+Cohere · Open weight
8.9%
142
Gemma 4 E4BGoogle · Open weight
8.6%
143
Ministral 3 3B (Reasoning)Mistral · Open weight
7.6%
144
Ministral 3 3BMistral · Open weight
7.6%
145
Gemma 4 E2BGoogle · Open weight
6.7%
146
LFM2.5-1.2B-ThinkingLiquidAI · Closed
6.6%
147
LFM2-24B-A2BLiquidAI · Closed
6.4%
148
Granite-4.0-1BIBM · Open weight
6.1%
149
LFM2.5-1.2B-InstructLiquidAI · Closed
6.0%
150
Granite-4.0-H-1BIBM · Open weight
5.3%
151
LFM2.5-VL-1.6B-ExtractLiquidAI · Open weight
5.2%
152
Exaone 4.0 1.2BLG AI Research · Open weight
4.7%
153
Granite-4.0-H-350MIBM · Open weight
3.7%
154
Granite-4.0-350MIBM · Open weight
3.2%

The published AA-Omniscience Accuracy snapshot places Claude Fable 5 first at 61.4%. The third row is 4.5 points behind. The broader top-10 range is 9.6 points, so many of the published results sit in a relatively narrow band.

154 models have been evaluated on AA-Omniscience Accuracy. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. AA-Omniscience Accuracy is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AA-Omniscience Accuracy

Year

2026

Tasks

Knowledge questions

Format

Accuracy

Difficulty

Broad knowledge

BenchLM stores AA-Omniscience Accuracy as a display-only row when a model page publishes the exact Artificial Analysis benchmark card value.

BenchLM freshness & provenance

Version

AA-Omniscience Accuracy 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AA-Omniscience Accuracy measure?

A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.

Which model scores highest on AA-Omniscience Accuracy?

Claude Fable 5 by Anthropic currently leads with a score of 61.4% on AA-Omniscience Accuracy.

How many models are evaluated on AA-Omniscience Accuracy?

154 AI models have been evaluated on AA-Omniscience Accuracy on BenchLM.

Last updated: July 29, 2026 · BenchLM version AA-Omniscience Accuracy 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.