Benchmark profile
xAI general refusal compliance (General refusal compliance)
Compliance with harmful requests on xAI's general refusal evaluation.
Data verified 26 confirmed releases in the last 30 daysStart free briefBenchmark score on General refusal compliance — August 12, 2026
BenchLM mirrors the published score view for General refusal compliance. Grok 4.6 leads the public snapshot at 0.93% , followed by Grok 4.5 (1.10%). We do not use these results to rank models overall.
Grok 4.6
xAI
grok-4-6
Grok 4.5
xAI
grok-4-5
About General refusal compliance
Year
2026
Tasks
General harmful requests
Format
Compliance rate
Difficulty
General safety
The exact task set is not public. Lower harmful compliance is better; the lane is display-only.
BenchLM freshness & provenance
Version
General refusal compliance 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does General refusal compliance measure?
Compliance with harmful requests on xAI's general refusal evaluation.
Which model scores highest on General refusal compliance?
Grok 4.6 by xAI currently leads with a score of 0.93% on General refusal compliance.
How many models are evaluated on General refusal compliance?
2 AI models have been evaluated on General refusal compliance on BenchLM.
Compare Top Models on General refusal compliance
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.