EQ-Bench 4
We show this table for reference; we do not rank on it.
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
EQ-Bench 4 Elo on EQ-Bench 4 — July 20, 2026
We mirror the published eq-bench 4 elo view for EQ-Bench 4. Claude Fable 5 leads the public snapshot at 1349.5, followed by Kimi K3 (1349.2) and GPT-5.5 (1325.8). We do not use these results to rank models overall.
Claude Fable 5
Anthropic
Kimi K3
Moonshot AI
GPT-5.5
OpenAI
25 modelsAgenticCurrentDisplay onlyUpdated July 20, 2026
EQ-Bench 4 Elo table (25 models)
ScoreHow EQ-Bench 4 is shown here
BenchLM mirrors the official EQ-Bench 4 snapshot generated on July 20, 2026. Each model plays 120 personas through 16-turn conversations, and three judge models compare the transcripts across six emotional and social ability dimensions.
EQ-Bench 4 reports pairwise Elo rather than a bounded accuracy percentage. The benchmark uses synthetic personas and LLM judges without human validation, so BenchLM preserves the public table and uncertainty fields as display-only behavioral evidence.
Snapshot
The published EQ-Bench 4 snapshot places Claude Fable 5 first at 1349.5. The third row is 23.7 score units behind. The broader top-10 range is 115.0 score units, so the table still separates the published systems.
25 models have been evaluated on EQ-Bench 4. The benchmark falls in the Agentic category. EQ-Bench 4 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About EQ-Bench 4
Year
2026
Tasks
120 multi-turn persona scenarios
Format
Pairwise Elo
Difficulty
Emotional and social intelligence
EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.
Freshness and provenance
Version
EQ-Bench 4 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does EQ-Bench 4 measure?
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
Which model leads the published EQ-Bench 4 snapshot?
Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with 1349.5 eq-bench 4 elo. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on EQ-Bench 4?
The July 20, 2026 snapshot contains 25 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.