Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

BullshitBench v2

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

How BenchLM shows BullshitBench v2

BenchLM mirrors the published BullshitBench v2 leaderboard using the official snapshot generated on August 24, 2026 at 9:50 PM UTC. The public view reports per-model clear-pushback rates across 100 nonsense prompts, scored by a 3-judge panel.

BullshitBench is a useful reasoning sanity check, but BenchLM currently keeps it display only rather than weighted. The public leaderboard is highly variant-specific and exposes reasoning-effort settings directly, so BenchLM treats it as a mirrored external benchmark instead of a canonical ranking input.

Snapshot

208 model variants118 base models100 nonsense prompts3 judgesDisplay only

Clear pushback rate on BullshitBench v2 — August 24, 2026 at 9:50 PM UTC

We mirror the published clear pushback rate view for BullshitBench v2. Claude Opus 4.8 leads the public snapshot at 95%, followed by Claude Opus 4.8 (94%) and Qwen3.8 Max (94%). We do not use these results to rank models overall.

208 modelsReasoningCurrentDisplay onlyUpdated August 24, 2026 at 9:50 PM UTC

Clear pushback rate table (208 models)

Score
1
Claude Opus 4.8Anthropic · Closed
95%
2
Claude Opus 4.8Anthropic · Closed
94%
3
Qwen3.8 MaxAlibaba · Open weight
94%
4
Claude Sonnet 4.6Anthropic · Closed
91%
5
Claude Opus 4.5Anthropic · Closed
90%
6
Claude Sonnet 4.6Anthropic · Closed
89%
7
Claude Opus 4.6Anthropic · Closed
87%
8
Claude Opus 4.6Anthropic · Closed
83%
9
Claude Opus 4.7Anthropic · Closed
83%
10
Claude Sonnet 5Anthropic · Closed
80%
11
Claude Sonnet 4.5Anthropic · Closed
79%
12
Claude Opus 4.5Anthropic · Closed
79%
13
Claude Sonnet 5Anthropic · Closed
78%
14
Qwen3.5 397B (Reasoning)Alibaba · Open weight
78%
15
Qwen3.8-27BAlibaba · Open weight
78%
16
Claude Haiku 4.5Anthropic · Closed
77%
17
Claude Opus 4.7Anthropic · Closed
74%
18
Claude Sonnet 4.5Anthropic · Closed
74%
19
Claude Opus 5Anthropic · Closed
73%
20
Kimi K3Moonshot AI · Closed
73%
21
Qwen3.6 PlusAlibaba · Closed
72%
22
Kimi K3Moonshot AI · Closed
71%
23
Claude Haiku 4.5Anthropic · Closed
71%
24
GLM-5.3Z.AI · Open weight
71%
25
Qwen3.7 MaxAlibaba · Closed
71%
26
Claude Opus 5Anthropic · Closed
70%
27
Qwen3.5 397BAlibaba · Open weight
69%
28
67%
29
GLM-5.3Z.AI · Open weight
65%
30
Gemini 3.5 Flash-LiteGoogle · Closed
65%
31
Kimi K2.6Moonshot AI · Open weight
65%
32
Grok 4.6xAI · Closed
65%
34
64%
35
Qwen3.6 PlusAlibaba · Closed
63%
36
MiniMax M3MiniMax · Open weight
63%
37
MiniMax M3MiniMax · Open weight
62%
38
MiMo-V2.5-ProXiaomi · Closed
62%
39
Qwen3.6 PlusAlibaba · Closed
59%
40
Gemini 3.5 Flash-LiteGoogle · Closed
59%
41
Qwen3.7 MaxAlibaba · Closed
56%
42
Grok 4.20xAI · Closed
56%
43
Grok 4.6xAI · Closed
55%
44
Claude Fable 5Anthropic · Closed
54%
45
Grok 4.5xAI · Closed
54%
46
Grok 4.5xAI · Closed
54%
47
Grok 4.20xAI · Closed
54%
48
Nemotron 3 Super 120B A12BNVIDIA · Open weight
54%
50
GPT-5.6 TerraOpenAI · Closed
53%
51
Kimi K2.5Moonshot AI · Open weight
52%
52
Grok 4.3xAI · Closed
50%
53
Kimi K2.6Moonshot AI · Open weight
50%
54
Muse Spark 1.2Meta · Closed
50%
57
Muse Spark 1.2Meta · Closed
49%
59
GPT-5.4OpenAI · Closed
48%
60
Gemini 3 ProGoogle · Closed
48%
61
GPT-5.6 SolOpenAI · Closed
47%
62
GPT-5.5OpenAI · Closed
47%
63
Nemotron 3 Super 120B A12BNVIDIA · Open weight
47%
64
GPT-5.6 SolOpenAI · Closed
46%
65
GPT-5.6 TerraOpenAI · Closed
46%
66
Qwen3.6 PlusAlibaba · Closed
46%
67
Grok 4.3xAI · Closed
46%
68
GPT-5.5OpenAI · Closed
45%
69
GPT-5.5OpenAI · Closed
45%
70
GPT-5.2-CodexOpenAI · Closed
45%
71
Claude 3.5 SonnetAnthropic · Closed
45%
72
GPT-5.1OpenAI · Closed
45%
73
Claude Fable 5Anthropic · Closed
44%
74
Claude 4.1 OpusAnthropic · Closed
43%
76
Nemotron 3 Super 120B A12BNVIDIA · Open weight
43%
78
GPT-5.4OpenAI · Closed
42%
79
Claude 4.1 OpusAnthropic · Closed
42%
80
GPT-5.6 LunaOpenAI · Closed
40%
81
GPT-5.3 InstantOpenAI · Closed
40%
84
GPT-5.2-CodexOpenAI · Closed
39%
85
Gemini 3.6 FlashGoogle · Closed
39%
86
GPT-5.2OpenAI · Closed
38%
87
DeepSeek V4 Flash 0731DeepSeek · Closed
38%
88
MiMo-V2.5-ProXiaomi · Closed
38%
89
Gemini 3.1 ProGoogle · Closed
37%
90
GPT-5.2-CodexOpenAI · Closed
37%
92
GPT-5.5 ProOpenAI · Closed
36%
93
GPT-5.6 LunaOpenAI · Closed
36%
94
Gemini 3 Pro Deep ThinkGoogle · Closed
36%
95
DeepSeek V4 Pro 0813DeepSeek · Closed
35%
96
Gemini 3.7 FlashGoogle · Closed
35%
97
MiMo-V2.5Xiaomi · Closed
35%
99
GPT-5.5 ProOpenAI · Closed
34%
100
GPT-5.5OpenAI · Closed
34%
102
GPT-5.4 miniOpenAI · Closed
32%
103
GPT-5.4 miniOpenAI · Closed
32%
104
DeepSeek V4 Pro 0813DeepSeek · Closed
32%
105
GPT-5.1-Codex-MaxOpenAI · Closed
32%
106
GPT-5.4 miniOpenAI · Closed
31%
107
Kimi K2.5 (Reasoning)Moonshot AI · Closed
31%
108
Gemini 3.1 ProGoogle · Closed
31%
109
Gemini 3.7 FlashGoogle · Closed
31%
110
GLM-5-TurboZ.AI · Closed
31%
111
Nemotron 3 Super 120B A12BNVIDIA · Open weight
31%
112
GLM-5.2Z.AI · Open weight
31%
113
Claude 4 SonnetAnthropic · Closed
30%
114
Claude 4 SonnetAnthropic · Closed
29%
115
GPT-5.2OpenAI · Closed
28%
116
Qwen3.8-27BAlibaba · Open weight
28%
117
Llama 4 MaverickMeta · Open weight
28%
118
Gemini 3.6 FlashGoogle · Closed
28%
119
GLM-5 (Reasoning)Z.AI · Open weight
28%
121
GPT-5.2 InstantOpenAI · Closed
27%
122
Qwen3.8 MaxAlibaba · Open weight
26%
123
o3OpenAI · Closed
26%
125
GPT-5.1OpenAI · Closed
25%
126
Gemma 4 31BGoogle · Open weight
25%
127
GPT-5.3 CodexOpenAI · Closed
24%
128
MiMo-V2.5Xiaomi · Closed
24%
129
GLM-5-TurboZ.AI · Closed
23%
130
GLM-5.1Z.AI · Open weight
22%
131
Step 3.5 FlashStepFun · Open weight
22%
132
Laguna S 2.1Poolside · Open weight
22%
133
21%
134
Gemma 4 26B A4BGoogle · Open weight
21%
135
GPT-5.3 CodexOpenAI · Closed
20%
136
DeepSeek V4 Flash 0731DeepSeek · Closed
20%
138
Gemini 2.5 ProGoogle · Closed
20%
139
Laguna S 2.1Poolside · Open weight
20%
140
GLM-5Z.AI · Open weight
20%
141
Gemma 4 31BGoogle · Open weight
20%
142
Gemini 3.5 FlashGoogle · Closed
20%
143
GPT-5.3 CodexOpenAI · Closed
19%
144
Grok 4.1 FastxAI · Closed
19%
145
Llama 4 ScoutMeta · Open weight
19%
146
Gemini 2.5 FlashGoogle · Closed
19%
147
Gemini 3.5 FlashGoogle · Closed
19%
148
18%
149
DeepSeek V4 Flash 0731DeepSeek · Closed
18%
150
GLM-5.1Z.AI · Open weight
18%
151
GLM-5.2Z.AI · Open weight
17%
152
Trinity-Large-ThinkingArcee AI · Open weight
17%
153
MiMo-V2-FlashXiaomi · Open weight
16%
154
Hy3Tencent · Open weight
16%
156
DeepSeek V4 Pro 0813DeepSeek · Closed
14%
158
GPT-5.4 nanoOpenAI · Closed
14%
159
DeepSeek V4 Pro 0813DeepSeek · Closed
14%
160
GPT-4.1OpenAI · Closed
14%
161
DeepSeek V4 Flash 0731DeepSeek · Closed
14%
162
GPT-5.4 nanoOpenAI · Closed
13%
163
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
13%
164
Step 3.5 FlashStepFun · Open weight
13%
165
Trinity-Large-ThinkingArcee AI · Open weight
13%
166
MiMo-V2-FlashXiaomi · Open weight
13%
167
GPT-4oOpenAI · Closed
12%
168
Gemma 4 26B A4BGoogle · Open weight
11%
169
Gemini 3.1 Flash-LiteGoogle · Closed
11%
170
Seed 1.6ByteDance · Closed
11%
171
GPT-OSS 120BOpenAI · Open weight
11%
173
GPT-5.4 nanoOpenAI · Closed
10%
174
Gemini 3 FlashGoogle · Closed
10%
175
DeepSeek V3.2DeepSeek · Open weight
10%
176
Claude 3 HaikuAnthropic · Closed
10%
177
Gemini 3 FlashGoogle · Closed
10%
179
Kimi K2Moonshot AI · Closed
10%
180
Grok 4.1 FastxAI · Closed
10%
181
MiniMax M2.5MiniMax · Closed
9%
182
Hy3Tencent · Open weight
8%
183
MiniMax M2.5MiniMax · Closed
8%
184
GLM-4.5Z.AI · Closed
8%
185
MiniMax M2.7MiniMax · Open weight
8%
186
DeepSeek-R1DeepSeek · Open weight
8%
187
o4-mini (high)OpenAI · Closed
8%
188
Seed 1.6ByteDance · Closed
7%
189
MiniMax M2.7MiniMax · Open weight
7%
190
DeepSeek-R1DeepSeek · Open weight
7%
194
GLM-4.5Z.AI · Closed
6%
195
GPT-OSS 120BOpenAI · Open weight
5%
198
5%
199
o4-mini (high)OpenAI · Closed
4%
208
GPT-4o miniOpenAI · Closed
2%

The published BullshitBench v2 snapshot places Claude Opus 4.8 first at 95%. The third row is 1.0 points behind. The broader top-10 range is 15.0 points, so the table still separates the published systems.

208 models have been evaluated on BullshitBench v2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. BullshitBench v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About BullshitBench v2

Year

2025

Tasks

Nonsensical and flawed prompts across multiple domains

Format

Prompt challenge and refusal evaluation

Difficulty

Robustness and critical reasoning

BullshitBench evaluates a crucial real-world capability: knowing when NOT to answer. Models that score highly recognize flawed premises, impossible physics scenarios, and logical contradictions rather than hallucinating plausible-sounding responses. V2 includes harder and more diverse challenge categories.

BenchLM freshness & provenance

Version

BullshitBench v2 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does BullshitBench v2 measure?

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

Which model leads the published BullshitBench v2 snapshot?

Claude Opus 4.8 currently leads the published BullshitBench v2 snapshot with 95% clear pushback rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on BullshitBench v2?

The August 24, 2026 at 9:50 PM UTC contains 208 AI models.

Last updated: August 24, 2026 at 9:50 PM UTC · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.