Benchmark profile
EQ-Bench 4
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
Data verifiedHow BenchLM shows EQ-Bench 4
BenchLM mirrors the official EQ-Bench 4 snapshot generated on July 20, 2026. Each model plays 120 personas through 16-turn conversations, and three judge models compare the transcripts across six emotional and social ability dimensions.
EQ-Bench 4 reports pairwise Elo rather than a bounded accuracy percentage. The benchmark uses synthetic personas and LLM judges without human validation, so BenchLM preserves the public table and uncertainty fields as display-only behavioral evidence.
EQ-Bench 4 Elo on EQ-Bench 4 — July 20, 2026
BenchLM mirrors the published eq-bench 4 elo view for EQ-Bench 4. Claude Fable 5 leads the public snapshot at 1349.5 , followed by Kimi K3 (1349.2) and GPT-5.5 (1325.8). BenchLM does not use these results to rank models overall.
Claude Fable 5
Anthropic
claude-fable-5
Kimi K3
Moonshot AI
moonshotai/kimi-k3
GPT-5.5
OpenAI
openai/gpt-5.5
EQ-Bench 4 Elo table (25 models)
ScoreThe published EQ-Bench 4 snapshot places Claude Fable 5 first at 1349.5. The third row is 23.7 score units behind. The broader top-10 range is 115.0 score units, so the table still separates the published systems.
25 models have been evaluated on EQ-Bench 4. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. EQ-Bench 4 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About EQ-Bench 4
Year
2026
Tasks
120 multi-turn persona scenarios
Format
Pairwise Elo
Difficulty
Emotional and social intelligence
EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.
BenchLM freshness & provenance
Version
EQ-Bench 4 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does EQ-Bench 4 measure?
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
Which model leads the published EQ-Bench 4 snapshot?
Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with 1349.5 eq-bench 4 elo. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on EQ-Bench 4?
25 AI models are included in BenchLM's mirrored EQ-Bench 4 snapshot, based on the public leaderboard captured on July 20, 2026.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.