# CAIS AI Dashboard Text Capabilities Index (CAIS Text Leaderboard)

> A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

Canonical page: https://benchlm.ai/benchmarks/caistextleaderboard

- Category: [external](/external)
- Last updated: June 2026 dashboard snapshot

## About CAIS Text Leaderboard

- Year: 2025
- Tasks: HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests
- Format: Average component score
- Difficulty: Composite frontier text capability
- Paper: [CAIS AI Dashboard](https://dashboard.safe.ai/)

BenchLM mirrors the text-capability portion of the CAIS AI Dashboard as a display-only composite. The displayed score is the average of the public HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests component scores.

CAIS Text Leaderboard is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 54.1% |
| 2 | [Opus 4.8](https://dashboard.safe.ai/) | Anthropic | 53.8% |
| 3 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 52.9% |
| 4 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 49.3% |
| 5 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 48.8% |
| 6 | [Opus 4.7](https://dashboard.safe.ai/) | Anthropic | 46.9% |
| 7 | [Opus 4.6](https://dashboard.safe.ai/) | Anthropic | 44.0% |
| 8 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 38.4% |
| 9 | [Opus 4.5](https://dashboard.safe.ai/) | Anthropic | 36.6% |
| 10 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 35.6% |
| 11 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 33.8% |
| 12 | [Sonnet 4.6](https://dashboard.safe.ai/) | Anthropic | 32.6% |
| 13 | [Grok 4.2](https://dashboard.safe.ai/) | xAI | 32.5% |
| 14 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 32.1% |
| 15 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 31.4% |
| 16 | [GLM-5.1](/models/glm-5-1) | Z.AI | 29.8% |
| 17 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 29.0% |
| 18 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 26.1% |
| 19 | [Sonnet 4.5](https://dashboard.safe.ai/) | Anthropic | 25.4% |
| 20 | [Grok 4.3](/models/grok-4-3) | xAI | 24.7% |
| 21 | [GPT-5.4-mini](https://dashboard.safe.ai/) | OpenAI | 24.2% |
| 22 | [GPT-5 (high)](/models/gpt-5-high) | OpenAI | 20.9% |
| 23 | [Grok 4](/models/grok-4) | xAI | 20.8% |
| 24 | [o3](/models/o3) | OpenAI | 20.5% |
| 25 | [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) | DeepSeek | 20.3% |

## FAQ

### What does CAIS Text Leaderboard measure?

A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

### Which model leads the published CAIS Text Leaderboard snapshot?

GPT-5.5 currently leads the published CAIS Text Leaderboard snapshot with a score of 54.1%.

### How many models are evaluated on CAIS Text Leaderboard?

The June 2026 dashboard snapshot contains 25 AI models.

### Does CAIS Text Leaderboard affect BenchLM's overall score?

Not directly. CAIS Text Leaderboard is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
