# EQ-Bench 4

> A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

Canonical page: https://benchlm.ai/benchmarks/eqbench4

- Category: [Agentic](/agentic)
- Last updated: July 20, 2026

## About EQ-Bench 4

- Year: 2026
- Tasks: 120 multi-turn persona scenarios
- Format: Pairwise Elo
- Difficulty: Emotional and social intelligence
- Paper: [EQ-Bench 4](https://eqbench.com/)

EQ-Bench 4 runs 120 personas for 16 turns and uses three judge models to compare transcripts across six ability dimensions. BenchLM mirrors its pairwise Elo table as display-only behavioral evidence because the personas and judgments are model-generated and have not been validated against human ratings.

EQ-Bench 4 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | Anthropic | 1349.5 |
| 2 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 1349.2 |
| 3 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 1325.8 |
| 4 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 1321.1 |
| 5 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 1291.5 |
| 6 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 1283.4 |
| 7 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 1262.6 |
| 8 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 1245.9 |
| 9 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 1244.3 |
| 10 | [Inkling](/models/inkling) | Thinking Machines Lab | 1234.5 |
| 11 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 1233.4 |
| 12 | [GLM-5.2](/models/glm-5-2) | Z.AI | 1229.3 |
| 13 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 1217.9 |
| 14 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 1212.6 |
| 15 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 1176.9 |
| 16 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 1165.8 |
| 17 | [MiniMax M3](/models/minimax-m3) | MiniMax | 1160.9 |
| 18 | [Gemini 3.1 Pro Preview](https://eqbench.com/) | Google | 1152.1 |
| 19 | [Gemma 4 31B IT](https://eqbench.com/) | Google | 1130.0 |
| 20 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 1120.1 |
| 21 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 1097.5 |
| 22 | [Grok 4.3](/models/grok-4-3) | xAI | 1084.1 |
| 23 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 1073.9 |
| 24 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 1035.2 |
| 25 | [Mistral Medium 3.5](https://eqbench.com/) | Mistral AI | 1002.5 |

## FAQ

### What does EQ-Bench 4 measure?

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

### Which model leads the published EQ-Bench 4 snapshot?

Claude Fable 5 currently leads the published EQ-Bench 4 snapshot with a score of 1349.5.

### How many models are evaluated on EQ-Bench 4?

The July 20, 2026 contains 25 AI models.

### Does EQ-Bench 4 affect BenchLM's overall score?

Not directly. EQ-Bench 4 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
