Benchmark profile
xAI factuality hallucination rate (Factuality hallucination rate)
The share of factuality-evaluation answers classified as hallucinations in the Grok 4.6 model card.
Data verified 26 confirmed releases in the last 30 daysStart free briefBenchmark score on Factuality hallucination rate — August 12, 2026
BenchLM mirrors the published score view for Factuality hallucination rate. Grok 4.5 leads the public snapshot at 0.98% , followed by GPT-5.5 (1.10%) and Grok 4.6 (1.70%). We do not use these results to rank models overall.
Grok 4.5
xAI
grok-4-5
GPT-5.5
OpenAI
gpt-5-5
Grok 4.6
xAI
grok-4-6
Benchmark score table (4 models)
ScoreThe published Factuality hallucination rate snapshot places Grok 4.5 first at 0.98%. The third row is 0.72 points higher. The broader top-10 range is 2.42 points, so many of the published results sit in a relatively narrow band.
4 models have been evaluated on Factuality hallucination rate. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Factuality hallucination rate is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Factuality hallucination rate
Year
2026
Tasks
Internal factuality questions
Format
Hallucination rate
Difficulty
Factual reliability
No public task set or matching result page was found. BenchLM preserves the exact card values as display-only internal-evaluation evidence.
BenchLM freshness & provenance
Version
Factuality hallucination rate 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Factuality hallucination rate measure?
The share of factuality-evaluation answers classified as hallucinations in the Grok 4.6 model card.
Which model scores highest on Factuality hallucination rate?
Grok 4.5 by xAI currently leads with a score of 0.98% on Factuality hallucination rate.
How many models are evaluated on Factuality hallucination rate?
4 AI models have been evaluated on Factuality hallucination rate on BenchLM.
Compare Top Models on Factuality hallucination rate
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.