Skip to main content
BenchLM

CAIS AI Dashboard Text Capabilities Index (CAIS Text Leaderboard)

We show this table for reference; we do not rank on it.

A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

Text average on CAIS Text Leaderboard — June 2026 dashboard snapshot

We mirror the published text average view for CAIS Text Leaderboard. GPT-5.5 leads the public snapshot at 54.1%, followed by Opus 4.8 (53.8%) and Gemini 3.1 Pro (52.9%). We do not use these results to rank models overall.

25 modelsCross-domainCurrentDisplay onlyUpdated June 2026 dashboard snapshot

Text average table (25 models)

Score
1
GPT-5.5OpenAI · Closed
54.1%
2
Opus 4.8Anthropic
53.8%
3
Gemini 3.1 ProGoogle · Closed
52.9%
4
GPT-5.4OpenAI · Closed
49.3%
5
Gemini 3.5 FlashGoogle · Closed
48.8%
6
Opus 4.7Anthropic
46.9%
7
Opus 4.6Anthropic
44.0%
8
Gemini 3 ProGoogle · Closed
38.4%
9
Opus 4.5Anthropic
36.6%
10
Gemini 3 FlashGoogle · Closed
35.6%
11
GPT-5.2OpenAI · Closed
33.8%
12
Sonnet 4.6Anthropic
32.6%
13
32.5%
14
DeepSeek V4 Pro 0813DeepSeek · Open weight
32.1%
15
Kimi K2.6Moonshot AI · Open weight
31.4%
16
GLM-5.1Z.AI · Open weight
29.8%
17
GPT-5.1OpenAI · Closed
29.0%
18
Kimi K2.5Moonshot AI · Open weight
26.1%
19
Sonnet 4.5Anthropic
25.4%
20
Grok 4.3xAI · Closed
24.7%
21
24.2%
22
GPT-5 (high)OpenAI · Closed
20.9%
23
Grok 4xAI · Closed
20.8%
24
o3OpenAI · Closed
20.5%
25
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
20.3%

How the CAIS Text Leaderboard is shown here

BenchLM mirrors the CAIS AI Dashboard text-capability view as a simple average over hle, arc_agi_2, swebench_pro, textquests. The source dashboard publishes the component benchmark scores and model metadata used here.

The CAIS Text Leaderboard is display only on BenchLM. It is a composite dashboard view rather than a single benchmark-native task set, so BenchLM keeps it out of weighted rankings.

Snapshot

25 mirrored rows4 text componentsCAIS AI DashboardComposite scoreDisplay only

The published CAIS Text Leaderboard snapshot places GPT-5.5 first at 54.1%. The third row is 1.2 points behind. The broader top-10 range is 18.5 points, so the table still separates the published systems.

25 models have been evaluated on CAIS Text Leaderboard. The benchmark falls in the Cross-domain category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. CAIS Text Leaderboard is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CAIS Text Leaderboard

Year

2025

Tasks

HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests

Format

Average component score

Difficulty

Composite frontier text capability

BenchLM mirrors the text-capability portion of the CAIS AI Dashboard as a display-only composite. The displayed score is the average of the public HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests component scores.

Freshness and provenance

Version

CAIS Text Leaderboard 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CAIS Text Leaderboard measure?

A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

Which model leads the published CAIS Text Leaderboard snapshot?

GPT-5.5 currently leads the published CAIS Text Leaderboard snapshot with 54.1% text average. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on CAIS Text Leaderboard?

The June 2026 dashboard snapshot snapshot contains 25 AI models.

Last updated: June 2026 dashboard snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.