Skip to main content

Benchmark profile

EQ-Bench 4

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

Data verified

How BenchLM shows EQ-Bench 4

BenchLM mirrors the official EQ-Bench 4 snapshot generated on July 20, 2026. Each model plays 120 personas through 16-turn conversations, and three judge models compare the transcripts across six emotional and social ability dimensions.

EQ-Bench 4 reports pairwise Elo rather than a bounded accuracy percentage. The benchmark uses synthetic personas and LLM judges without human validation, so BenchLM preserves the public table and uncertainty fields as display-only behavioral evidence.

25 models120 personas16 turns3 judgesDisplay only

EQ-Bench 4 Elo on EQ-Bench 4 — July 20, 2026

BenchLM mirrors the published eq-bench 4 elo view for EQ-Bench 4. Claude Fable 5 leads the public snapshot at 1349.5 , followed by Kimi K3 (1349.2) and GPT-5.5 (1325.8). BenchLM does not use these results to rank models overall.

25 modelsAgenticCurrentDisplay onlyUpdated July 20, 2026

EQ-Bench 4 Elo table (25 models)

Score
1
1349.5
2
Kimi K3Moonshot AI
1349.2
3
GPT-5.5OpenAI
1325.8
4
1321.1
5
1291.5
6
GPT-5.4OpenAI
1283.4
7
1262.6
8
1245.9
9
1244.3
10
InklingThinking Machines Lab
1234.5
11
1233.4
12
1229.3
13
1217.9
14
Kimi K2.6Moonshot AI
1212.6
15
1176.9
16
1165.8
17
MiniMax M3MiniMax
1160.9
19
1130.0
20
1120.1
21
1097.5
22
1084.1
23
1073.9
24
1035.2
25
1002.5

The published EQ-Bench 4 snapshot places Claude Fable 5 first at 1349.5. The third row is 23.7 score units behind. The broader top-10 range is 115.0 score units, so the table still separates the published systems.

25 models have been evaluated on EQ-Bench 4. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. EQ-Bench 4 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About EQ-Bench 4

Year

2026

Tasks

120 multi-turn persona scenarios

Format

Pairwise Elo

Difficulty

Emotional and social intelligence

EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.

BenchLM freshness & provenance

Version

EQ-Bench 4 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does EQ-Bench 4 measure?

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

Which model leads the published EQ-Bench 4 snapshot?

Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with 1349.5 eq-bench 4 elo. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on EQ-Bench 4?

25 AI models are included in BenchLM's mirrored EQ-Bench 4 snapshot, based on the public leaderboard captured on July 20, 2026.

Last updated: July 20, 2026 · mirrored from the public benchmark leaderboard