Skip to main content
BenchLM

Vals Terminal-Bench 4.0 (Terminal-Bench 4.0 (Vals))

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Vals AI’s independent run of the Terminal-Bench 4.0 frontier terminal suite, with per-domain splits.

Vals Terminal-Bench 4.0 mirror score on Terminal-Bench 4.0 (Vals) — September 27, 2026

We mirror the published vals terminal-bench 4.0 mirror score view for Terminal-Bench 4.0 (Vals). Claude Opus 5.5 leads the public snapshot at 61.62%, followed by GPT-6 Astra (57.07%) and Claude Sonnet 5.5 (53.03%). We do not use these results to rank models overall.

37 modelsAgenticCurrentDisplay onlyUpdated September 27, 2026

Vals Terminal-Bench 4.0 mirror score table (37 models)

Score
1
Claude Opus 5.5Anthropic · Closed
61.62%
2
GPT-6 AstraOpenAI · Closedmax reasoning
57.07%
3
Claude Sonnet 5.5Anthropic · Closed
53.03%
4
Claude Fable 5.1Anthropic · Closed
49.49%
5
Claude Opus 5Anthropic · Closed
45.45%
6
GPT-6 SolOpenAI · Closedmax reasoning
34.34%
7
Grok 4.7xAI · Closedxhigh reasoning
28.28%
8
GPT-5.6 SolOpenAI · Closedmax reasoning
27.78%
9
Muse Spark 1.3 MaxMetamax reasoning
27.78%
10
GPT-5.6 TerraOpenAI · Closedmax reasoning
26.26%
11
GLM-5.3Z.AI · Open weightmax reasoning
25.25%
12
MiMo-V2.6-ProXiaomi · Open weight
24.75%
13
Qwen3.8 MaxAlibaba · Open weight
24.75%
14
Claude Fable 5Anthropic · Closed
22.73%
15
MiMo-V2.6-FlashXiaomi · Open weight
21.21%
16
GLM-5.3-FlashZ.AI · Open weightmax reasoning
19.70%
17
Grok 4.6xAI · Closedhigh reasoning
17.17%
18
Claude Opus 4.8Anthropic · Closed
16.16%
19
Muse Spark 1.3Meta · Closedxhigh reasoning
15.15%
20
Gemini 3.8 FlashGoogle · Closedhigh reasoning
13.13%
21
Kimi K3Moonshot AI · Closedmax reasoning
12.63%
22
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
11.62%
23
GPT-6 LunaOpenAI · Closedmax reasoning
9.60%
24
DeepSeek V4 Flash 0731DeepSeek · Open weighthigh reasoning
9.09%
25
Claude Sonnet 5Anthropic · Closed
8.08%
26
Grok 4.5xAI · Closedhigh reasoning
6.57%
27
Gemini 3.7 FlashGoogle · Closedhigh reasoning
6.06%
28
Muse Spark 1.2Meta · Closedxhigh reasoning
5.56%
29
Hy4 previewTencent · Open weight
5.05%
30
GPT-5.6 LunaOpenAI · Closedmax reasoning
4.54%
31
Gemini 3.6 FlashGoogle · Closedhigh reasoning
4.54%
32
Gemini 3.5 FlashGoogle · Closedhigh reasoning
4.04%
33
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
4.04%
34
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
1.51%
35
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
1.01%
36
Mercury 2.5Inception · Closedhigh reasoning
0.00%
37
InklingThinking Machines Lab · Open weight0.99 reasoning
0.00%

How Vals Terminal-Bench 4.0 mirror is shown here

BenchLM mirrors the public Vals AI Vals Terminal-Bench 4.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-4 and updated by Vals on September 27, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 4.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

37 Vals rows8 task viewspublic datasetTasks: Overall, Software, Science, ML, OperationsDisplay only

The published Terminal-Bench 4.0 (Vals) snapshot places Claude Opus 5.5 first at 61.62%. The third row is 8.59 points behind. The broader top-10 range is 35.35 points, so the table still separates the published systems.

37 models have been evaluated on Terminal-Bench 4.0 (Vals). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Terminal-Bench 4.0 (Vals) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench 4.0 (Vals)

Year

2026

Tasks

Frontier-difficulty terminal tasks across seven domains

Format

Accuracy score

Difficulty

Frontier terminal-agent execution

Vals publishes overall accuracy plus software, science, ML, operations, hardware, security, and media splits with standard error, latency, and cost per test. BenchLM mirrors the board on a dedicated key so a Vals run never overwrites the official Terminal-Bench 4.0 rows, and keeps it display only and outside weighted rankings.

Freshness and provenance

Version

Terminal-Bench 4.0 (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Terminal-Bench 4.0 (Vals) measure?

Vals AI’s independent run of the Terminal-Bench 4.0 frontier terminal suite, with per-domain splits.

Which model leads the published Terminal-Bench 4.0 (Vals) snapshot?

Claude Opus 5.5 currently leads the published Terminal-Bench 4.0 (Vals) snapshot with 61.62% vals terminal-bench 4.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench 4.0 (Vals)?

The September 27, 2026 snapshot contains 37 AI models.

Last updated: September 27, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.