Benchmark profile
xAI sycophancy rate (Sycophancy)
The rate of sycophantic responses in xAI's internal evaluation.
Data verified 26 confirmed releases in the last 30 daysStart free briefBenchmark score on Sycophancy — August 12, 2026
BenchLM mirrors the published score view for Sycophancy. Grok 4.5 leads the public snapshot at 0.01% , followed by Grok 4.6 (0.04%). We do not use these results to rank models overall.
Grok 4.5
xAI
grok-4-5
Grok 4.6
xAI
grok-4-6
About Sycophancy
Year
2026
Tasks
Prompts testing agreement with a user's stated view
Format
Sycophancy rate
Difficulty
Behavioral robustness
Public sycophancy benchmarks exist, but no matching artifact for xAI's exact implementation was found. Lower is better; the lane is display-only.
BenchLM freshness & provenance
Version
Sycophancy 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Sycophancy measure?
The rate of sycophantic responses in xAI's internal evaluation.
Which model scores highest on Sycophancy?
Grok 4.5 by xAI currently leads with a score of 0.01% on Sycophancy.
How many models are evaluated on Sycophancy?
2 AI models have been evaluated on Sycophancy on BenchLM.
Compare Top Models on Sycophancy
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.