Skip to main content
BenchLM

EQ-Bench 4

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

EQ-Bench 4 Elo on EQ-Bench 4 — July 20, 2026

We mirror the published eq-bench 4 elo view for EQ-Bench 4. Claude Fable 5 leads the public snapshot at 1349.5, followed by Kimi K3 (1349.2) and GPT-5.5 (1325.8). We do not use these results to rank models overall.

25 modelsAgenticCurrentDisplay onlyUpdated July 20, 2026

EQ-Bench 4 Elo table (25 models)

Score
1
Claude Fable 5Anthropic · Closed
1349.5
2
Kimi K3Moonshot AI · Closed
1349.2
3
GPT-5.5OpenAI · Closed
1325.8
4
Claude Opus 4.7Anthropic · Closed
1321.1
5
Claude Opus 4.8Anthropic · Closed
1291.5
6
GPT-5.4OpenAI · Closed
1283.4
7
GPT-5.6 SolOpenAI · Closed
1262.6
8
Claude Sonnet 5Anthropic · Closed
1245.9
9
GPT-5.6 TerraOpenAI · Closed
1244.3
10
InklingThinking Machines Lab · Open weight
1234.5
11
Claude Opus 4.6Anthropic · Closed
1233.4
12
GLM-5.2Z.AI · Open weight
1229.3
13
Claude Sonnet 4.6Anthropic · Closed
1217.9
14
Kimi K2.6Moonshot AI · Open weight
1212.6
15
DeepSeek V4 Pro 0813DeepSeek · Open weight
1176.9
16
GPT-5.6 LunaOpenAI · Closed
1165.8
17
MiniMax M3MiniMax · Open weight
1160.9
19
1130.0
20
Qwen3.7 MaxAlibaba · Closed
1120.1
21
Gemini 3.5 FlashGoogle · Closed
1097.5
22
Grok 4.3xAI · Closed
1084.1
23
Claude Haiku 4.5Anthropic · Closed
1073.9
24
Qwen3.6-27BAlibaba · Open weight
1035.2
25
1002.5

How EQ-Bench 4 is shown here

BenchLM mirrors the official EQ-Bench 4 snapshot generated on July 20, 2026. Each model plays 120 personas through 16-turn conversations, and three judge models compare the transcripts across six emotional and social ability dimensions.

EQ-Bench 4 reports pairwise Elo rather than a bounded accuracy percentage. The benchmark uses synthetic personas and LLM judges without human validation, so BenchLM preserves the public table and uncertainty fields as display-only behavioral evidence.

Snapshot

25 models120 personas16 turns3 judgesDisplay only

The published EQ-Bench 4 snapshot places Claude Fable 5 first at 1349.5. The third row is 23.7 score units behind. The broader top-10 range is 115.0 score units, so the table still separates the published systems.

25 models have been evaluated on EQ-Bench 4. The benchmark falls in the Agentic category. EQ-Bench 4 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About EQ-Bench 4

Year

2026

Tasks

120 multi-turn persona scenarios

Format

Pairwise Elo

Difficulty

Emotional and social intelligence

EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.

Freshness and provenance

Version

EQ-Bench 4 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does EQ-Bench 4 measure?

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

Which model leads the published EQ-Bench 4 snapshot?

Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with 1349.5 eq-bench 4 elo. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on EQ-Bench 4?

The July 20, 2026 snapshot contains 25 AI models.

Last updated: July 20, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.