Skip to main content
BenchLM

Vals-hosted Terminal-Bench 2.0 mirror (Vals Terminal-Bench 2.0 mirror)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

Vals Terminal-Bench 2.0 mirror score on Vals Terminal-Bench 2.0 mirror — June 4, 2026

We mirror the published vals terminal-bench 2.0 mirror score view for Vals Terminal-Bench 2.0 mirror. GPT-5.5 leads the public snapshot at 73.20%, followed by Claude Opus 4.8 (70.04%) and Claude Opus 4.7 (68.54%). We do not use these results to rank models overall.

67 modelsAgenticCurrentDisplay onlyUpdated June 4, 2026

Vals Terminal-Bench 2.0 mirror score table (67 models)

Score
1
GPT-5.5OpenAI · Closedhigh reasoning
73.20%
2
Claude Opus 4.8Anthropic · Closed
70.04%
3
Claude Opus 4.7Anthropic · Closed
68.54%
4
Gemini 3.5 FlashGoogle · Closedhigh reasoning
67.42%
5
Gemini 3.1 Pro PreviewGooglehigh reasoning
67.42%
6
GPT-5.3 CodexOpenAI · Closedhigh reasoning
64.05%
7
Muse SparkMeta · Closed
59.55%
8
Claude Sonnet 4.6Anthropic · Closed
59.55%
9
Qwen3.7 MaxAlibaba · Closed
59.18%
10
Claude Opus 4.5Anthropic · Closed
58.43%
11
Claude Opus 4.6 (Adaptive)Anthropic · Closed
58.43%
12
GPT-5.4OpenAI · Closedhigh reasoning
58.43%
13
Kimi K2.6Moonshot AI · Open weight
57.30%
14
DeepSeek V4 Pro 0813DeepSeek · Open weight
56.18%
15
Gemini 3 Pro PreviewGooglehigh reasoning
55.06%
16
Claude Opus 4.5 ThinkingAnthropic · Closed
53.93%
17
GLM-5.1Z.AI · Open weight
53.93%
18
Gemini 3 Flash PreviewGooglehigh reasoning
51.69%
19
GPT-5.2OpenAI · Closedhigh reasoning
51.69%
20
Qwen 3.6 Max (preview)Alibaba · Closed
51.69%
21
49.44%
22
MiniMax M2.7MiniMax · Open weight
47.19%
23
MiniMax M3MiniMax · Open weight
46.07%
24
Qwen3.6 PlusAlibaba · Closed
44.94%
25
GPT-5.1OpenAI · Closedhigh reasoning
44.94%
26
Qwen3.6-27BAlibaba · Open weight
44.94%
27
GPT-5.4 miniOpenAI · Closedhigh reasoning
44.94%
28
Grok 4.3xAI · Closed
43.45%
29
Claude Sonnet 4.5 ThinkingAnthropic · Closed
41.57%
30
MiniMax M2.5MiniMax · Closed
41.57%
31
41.57%
33
40.45%
34
GPT-5.4 nanoOpenAI · Closedhigh reasoning
39.89%
35
Gemma 4 31b ItGooglehigh reasoning
39.33%
36
Claude Haiku 4.5 ThinkingAnthropic · Closed
38.20%
37
GLM-4.7Z.AI · Open weight
38.20%
38
37.08%
39
GPT-5OpenAIhigh reasoning
37.08%
40
Kimi K2 ThinkingMoonshot AI
37.08%
41
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
35.95%
42
DeepSeek V3p2Fireworks AInone reasoning
34.83%
43
Laguna M.1Poolside · Closed
31.46%
44
Gemini 2.5 ProGoogle · Closed
30.34%
45
Mistral Medium 3.5Mistral AIhigh reasoning
30.34%
46
29.21%
47
Laguna XS.2Poolside · Open weight
28.09%
48
GLM-4.6Z.AI · Open weight
28.09%
49
28.09%
50
GPT-5 miniOpenAI · Closedhigh reasoning
26.97%
51
25.84%
52
24.72%
53
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
24.72%
54
Qwen3.5 FlashAlibaba · Closed
24.72%
55
Qwen3 MaxAlibaba · Closed
24.72%
56
DeepSeek V3p1Fireworks AI
22.47%
58
Qwen3 MaxAlibaba · Closed
20.23%
59
GPT-OSS 120BOpenAI · Open weight
19.10%
61
Trinity-Large-ThinkingArcee AI · Open weight
17.98%
62
Mistral Small 2603Mistral AIhigh reasoning
16.85%
63
GPT-4.1OpenAI · Closedhigh reasoning
14.61%
64
13.48%
65
8.99%
66
2.25%

How Vals Terminal-Bench 2.0 mirror is shown here

BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2 and updated by Vals on June 4, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 2.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

67 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

The published Vals Terminal-Bench 2.0 mirror snapshot places GPT-5.5 first at 73.20%. The third row is 4.66 points behind. The broader top-10 range is 14.77 points, so the table still separates the published systems.

67 models have been evaluated on Vals Terminal-Bench 2.0 mirror. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Terminal-Bench 2.0 mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals Terminal-Bench 2.0 mirror

Year

2026

Tasks

Terminal task difficulty splits

Format

Accuracy score

Difficulty

Terminal-based agent execution

BenchLM mirrors this Vals-hosted Terminal-Bench view as display-only secondary context.

Freshness and provenance

Version

Vals Terminal-Bench 2.0 mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Vals Terminal-Bench 2.0 mirror measure?

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

Which model leads the published Vals Terminal-Bench 2.0 mirror snapshot?

GPT-5.5 currently leads the published Vals Terminal-Bench 2.0 mirror snapshot with 73.20% vals terminal-bench 2.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals Terminal-Bench 2.0 mirror?

The June 4, 2026 snapshot contains 67 AI models.

Last updated: June 4, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.