Skip to main content

Benchmark profile

BullshitBench v2

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

How BenchLM shows BullshitBench v2

BenchLM mirrors the published BullshitBench v2 leaderboard using the official snapshot generated on July 24, 2026 at 11:57 PM UTC. The public view reports per-model clear-pushback rates across 100 nonsense prompts, scored by a 3-judge panel.

BullshitBench is a useful reasoning sanity check, but BenchLM currently keeps it display only rather than weighted. The public leaderboard is highly variant-specific and exposes reasoning-effort settings directly, so BenchLM treats it as a mirrored external benchmark instead of a canonical ranking input.

190 model variants109 base models100 nonsense prompts3 judgesDisplay only

Clear pushback rate on BullshitBench v2 — July 24, 2026 at 11:57 PM UTC

BenchLM mirrors the published clear pushback rate view for BullshitBench v2. Claude Opus 4.8 (none) leads the public snapshot at 95% , followed by Claude Opus 4.8 (xhigh) (94%) and Claude Sonnet 4.6 (high) (91%). BenchLM does not use these results to rank models overall.

190 modelsReasoningCurrentDisplay onlyUpdated July 24, 2026 at 11:57 PM UTC

Clear pushback rate table (190 models)

Score
1
Claude Opus 4.8 (none)Anthropic · Closed
95%
2
Claude Opus 4.8 (xhigh)Anthropic · Closed
94%
3
Claude Sonnet 4.6 (high)Anthropic · Closed
91%
4
Claude Opus 4.5 (high)Anthropic · Closed
90%
5
Claude Sonnet 4.6 (none)Anthropic · Closed
89%
6
Claude Opus 4.6 (high)Anthropic · Closed
87%
7
Claude Opus 4.6 (none)Anthropic · Closed
83%
8
Claude Opus 4.7 (none)Anthropic · Closed
83%
9
Claude Sonnet 5 (low)Anthropic · Closed
80%
10
Claude Sonnet 4.5 (high)Anthropic · Closed
79%
11
Claude Opus 4.5 (none)Anthropic · Closed
79%
12
Claude Sonnet 5 (max)Anthropic · Closed
78%
13
Qwen3.5 397B (Reasoning) (high)Alibaba · Open weight
78%
14
Claude Haiku 4.5 (high)Anthropic · Closed
77%
15
Claude Opus 4.7 (max)Anthropic · Closed
74%
16
Claude Sonnet 4.5 (none)Anthropic · Closed
74%
17
Claude Opus 5 (low)Anthropic · Closed
73%
18
Kimi K3 (xhigh)Moonshot AI · Closed
73%
19
Qwen3.6 Plus (none)Alibaba · Closed
72%
20
Kimi K3 (minimal)Moonshot AI · Closed
71%
21
Claude Haiku 4.5 (none)Anthropic · Closed
71%
22
Qwen3.7 Max (none)Alibaba · Closed
71%
23
Claude Opus 5 (xhigh)Anthropic · Closed
70%
24
Qwen3.5 397B (none)Alibaba · Open weight
69%
26
65%
27
Kimi K2.6 (none)Moonshot AI · Open weight
65%
29
Qwen3.6 Plus (high)Alibaba · Closed
63%
30
MiniMax M3 (xhigh)MiniMax · Open weight
63%
31
MiniMax M3 (none)MiniMax · Open weight
62%
32
MiMo-V2.5-Pro (xhigh)Xiaomi · Closed
62%
33
Qwen3.6 Plus (xhigh)Alibaba · Closed
59%
34
59%
35
Qwen3.7 Max (xhigh)Alibaba · Closed
56%
37
Claude Fable 5 (xhigh)Anthropic · Closed
54%
38
Grok 4.5 (low)xAI · Closed
54%
39
Grok 4.5 (high)xAI · Closed
54%
41
54%
42
GPT-5.6 Terra (max)OpenAI · Closed
53%
43
Kimi K2.5 (none)Moonshot AI · Open weight
52%
44
Grok 4.3 (minimal)xAI · Closed
50%
45
Kimi K2.6 (xhigh)Moonshot AI · Open weight
50%
49
GPT-5.4 (none)OpenAI · Closed
48%
50
Gemini 3 Pro (low)Google · Closed
48%
51
GPT-5.6 Sol (low)OpenAI · Closed
47%
52
GPT-5.5 (xhigh)OpenAI · Closed
47%
53
47%
54
GPT-5.6 Sol (max)OpenAI · Closed
46%
55
GPT-5.6 Terra (low)OpenAI · Closed
46%
56
Qwen3.6 Plus (none)Alibaba · Closed
46%
57
Grok 4.3 (xhigh)xAI · Closed
46%
58
GPT-5.5 (none)OpenAI · Closed
45%
59
GPT-5.5 (low)OpenAI · Closed
45%
60
GPT-5.2-Codex (low)OpenAI · Closed
45%
61
Claude 3.5 SonnetAnthropic · Closed
45%
62
GPT-5.1OpenAI · Closed
45%
63
Claude Fable 5 (low)Anthropic · Closed
44%
64
Claude 4.1 Opus (none)Anthropic · Closed
43%
66
43%
68
GPT-5.4 (xhigh)OpenAI · Closed
42%
69
Claude 4.1 Opus (high)Anthropic · Closed
42%
70
GPT-5.6 Luna (max)OpenAI · Closed
40%
71
GPT-5.3 InstantOpenAI · Closed
40%
74
GPT-5.2-Codex (xhigh)OpenAI · Closed
39%
75
39%
76
GPT-5.2 (none)OpenAI · Closed
38%
77
MiMo-V2.5-Pro (none)Xiaomi · Closed
38%
78
Gemini 3.1 Pro (low)Google · Closed
37%
79
GPT-5.2-Codex (high)OpenAI · Closed
37%
81
GPT-5.5 Pro (xhigh)OpenAI · Closed
36%
82
GPT-5.6 Luna (low)OpenAI · Closed
36%
83
36%
84
MiMo-V2.5 (xhigh)Xiaomi · Closed
35%
86
GPT-5.5 Pro (medium)OpenAI · Closed
34%
87
GPT-5.5OpenAI · Closed
34%
89
GPT-5.4 mini (high)OpenAI · Closed
32%
90
GPT-5.4 mini (none)OpenAI · Closed
32%
91
GPT-5.1-Codex-MaxOpenAI · Closed
32%
92
GPT-5.4 mini (xhigh)OpenAI · Closed
31%
93
Kimi K2.5 (Reasoning) (high)Moonshot AI · Closed
31%
94
Gemini 3.1 Pro (high)Google · Closed
31%
95
GLM-5-Turbo (high)Z.AI · Closed
31%
96
31%
97
GLM-5.2 (xhigh)Z.AI · Open weight
31%
98
Claude 4 Sonnet (high)Anthropic · Closed
30%
99
Claude 4 Sonnet (none)Anthropic · Closed
29%
100
GPT-5.2 (high)OpenAI · Closed
28%
101
Llama 4 MaverickMeta · Open weight
28%
102
Gemini 3.6 Flash (xhigh)Google · Closed
28%
103
GLM-5 (Reasoning) (high)Z.AI · Open weight
28%
105
GPT-5.2 InstantOpenAI · Closed
27%
106
o3OpenAI · Closed
26%
108
GPT-5.1OpenAI · Closed
25%
109
Gemma 4 31B (high)Google · Open weight
25%
110
GPT-5.3 Codex (low)OpenAI · Closed
24%
111
MiMo-V2.5 (none)Xiaomi · Closed
24%
112
GLM-5-Turbo (none)Z.AI · Closed
23%
113
GLM-5.1 (xhigh)Z.AI · Open weight
22%
114
Step 3.5 Flash (xhigh)StepFun · Open weight
22%
115
Laguna S 2.1 (none)Poolside · Open weight
22%
116
21%
117
Gemma 4 26B A4B (xhigh)Google · Open weight
21%
118
GPT-5.3 Codex (high)OpenAI · Closed
20%
120
Gemini 2.5 ProGoogle · Closed
20%
121
Laguna S 2.1 (xhigh)Poolside · Open weight
20%
122
GLM-5 (none)Z.AI · Open weight
20%
123
Gemma 4 31B (none)Google · Open weight
20%
124
Gemini 3.5 Flash (xhigh)Google · Closed
20%
125
GPT-5.3 Codex (xhigh)OpenAI · Closed
19%
126
19%
127
Llama 4 ScoutMeta · Open weight
19%
128
Gemini 2.5 FlashGoogle · Closed
19%
129
19%
130
18%
131
DeepSeek V4 Flash (none)DeepSeek · Open weight
18%
132
GLM-5.1 (none)Z.AI · Open weight
18%
133
GLM-5.2 (none)Z.AI · Open weight
17%
134
Trinity-Large-Thinking (minimal)Arcee AI · Open weight
17%
135
MiMo-V2-Flash (none)Xiaomi · Open weight
16%
136
Hy3 (none)Tencent · Open weight
16%
138
DeepSeek V4 Pro (xhigh)DeepSeek · Open weight
14%
140
GPT-5.4 nano (high)OpenAI · Closed
14%
141
DeepSeek V4 Pro (none)DeepSeek · Open weight
14%
142
GPT-4.1OpenAI · Closed
14%
143
DeepSeek V4 Flash (xhigh)DeepSeek · Open weight
14%
144
GPT-5.4 nano (none)OpenAI · Closed
13%
145
DeepSeek V3.2 (Thinking) (high)DeepSeek · Open weight
13%
146
Step 3.5 Flash (minimal)StepFun · Open weight
13%
147
Trinity-Large-Thinking (xhigh)Arcee AI · Open weight
13%
148
MiMo-V2-Flash (high)Xiaomi · Open weight
13%
150
Gemma 4 26B A4B (none)Google · Open weight
11%
151
Gemini 3.1 Flash-LiteGoogle · Closed
11%
152
Seed 1.6 (none)ByteDance · Closed
11%
153
GPT-OSS 120B (low)OpenAI · Open weight
11%
155
GPT-5.4 nano (xhigh)OpenAI · Closed
10%
156
Gemini 3 Flash (high)Google · Closed
10%
157
DeepSeek V3.2 (none)DeepSeek · Open weight
10%
158
Claude 3 HaikuAnthropic · Closed
10%
159
Gemini 3 Flash (none)Google · Closed
10%
161
Kimi K2Moonshot AI · Closed
10%
162
10%
163
MiniMax M2.5 (low)MiniMax · Closed
9%
164
Hy3 (xhigh)Tencent · Open weight
8%
165
MiniMax M2.5 (high)MiniMax · Closed
8%
166
GLM-4.5 (xhigh)Z.AI · Closed
8%
167
MiniMax M2.7 (high)MiniMax · Open weight
8%
168
DeepSeek-R1 (xhigh)DeepSeek · Open weight
8%
169
o4-mini (high) (low)OpenAI · Closed
8%
170
Seed 1.6 (high)ByteDance · Closed
7%
171
MiniMax M2.7 (low)MiniMax · Open weight
7%
172
DeepSeek-R1 (none)DeepSeek · Open weight
7%
176
GLM-4.5 (none)Z.AI · Closed
6%
177
GPT-OSS 120B (high)OpenAI · Open weight
5%
180
5%
181
o4-mini (high) (high)OpenAI · Closed
4%

The published BullshitBench v2 snapshot places Claude Opus 4.8 (none) first at 95%. The third row is 4.0 points behind. The broader top-10 range is 16.0 points, so the table still separates the published systems.

190 models have been evaluated on BullshitBench v2. The benchmark falls in the Reasoning category. This category carries a 17% weight in BenchLM.ai's overall scoring system. BullshitBench v2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About BullshitBench v2

Year

2025

Tasks

Nonsensical and flawed prompts across multiple domains

Format

Prompt challenge and refusal evaluation

Difficulty

Robustness and critical reasoning

BullshitBench evaluates a crucial real-world capability: knowing when NOT to answer. Models that score highly recognize flawed premises, impossible physics scenarios, and logical contradictions rather than hallucinating plausible-sounding responses. V2 includes harder and more diverse challenge categories.

BenchLM freshness & provenance

Version

BullshitBench v2 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does BullshitBench v2 measure?

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

Which model leads the published BullshitBench v2 snapshot?

Claude Opus 4.8 (none) currently leads the published BullshitBench v2 snapshot with 95% clear pushback rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on BullshitBench v2?

190 AI models are included in BenchLM's mirrored BullshitBench v2 snapshot, based on the public leaderboard captured on July 24, 2026 at 11:57 PM UTC.

Last updated: July 24, 2026 at 11:57 PM UTC · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.