Benchmark profile
EQ-Bench 4
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
Data verifiedHow BenchLM shows EQ-Bench 4
BenchLM mirrors the official EQ-Bench 4 snapshot generated on July 20, 2026. Each model plays 120 personas through 16-turn conversations, and three judge models compare the transcripts across six emotional and social ability dimensions.
EQ-Bench 4 reports pairwise Elo rather than a bounded accuracy percentage. The benchmark uses synthetic personas and LLM judges without human validation, so BenchLM preserves the public table and uncertainty fields as display-only behavioral evidence.
EQ-Bench 4 Elo on EQ-Bench 4 — July 20, 2026
BenchLM mirrors the published eq-bench 4 elo view for EQ-Bench 4. Claude Fable 5 leads the public snapshot at 1349.5 , followed by Kimi K3 (1349.2) and GPT-5.5 (1325.8). BenchLM does not use these results to rank models overall.
Claude Fable 5
Anthropic
claude-fable-5
Kimi K3
Moonshot AI
moonshotai/kimi-k3
GPT-5.5
OpenAI
openai/gpt-5.5
EQ-Bench 4 Elo table (25 models)
ScoreThe published EQ-Bench 4 snapshot places Claude Fable 5 first at 1349.5. The third row is 23.7 score units behind. The broader top-10 range is 115.0 score units, so the table still separates the published systems.
25 models have been evaluated on EQ-Bench 4. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. EQ-Bench 4 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About EQ-Bench 4
Year
2026
Tasks
120 multi-turn persona scenarios
Format
Pairwise Elo
Difficulty
Emotional and social intelligence
EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.
BenchLM freshness & provenance
Version
EQ-Bench 4 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does EQ-Bench 4 measure?
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
Which model leads the published EQ-Bench 4 snapshot?
Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with 1349.5 eq-bench 4 elo. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on EQ-Bench 4?
25 AI models are included in BenchLM's mirrored EQ-Bench 4 snapshot, based on the public leaderboard captured on July 20, 2026.