Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

Vals Terminal-Bench 2.1 (Terminal-Bench 2.1)

State-of-the-art set of difficult terminal-based tasks

Data verified 24 confirmed releases in the last 30 daysStart the free Radar Brief

How BenchLM shows Vals Terminal-Bench 2.1 mirror

BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.1 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2-1 and updated by Vals on August 19, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 2.1 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

57 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals Terminal-Bench 2.1 mirror score on Terminal-Bench 2.1 — August 19, 2026

We mirror the published vals terminal-bench 2.1 mirror score view for Terminal-Bench 2.1. GPT-5.6 Sol leads the public snapshot at 85.77%, followed by Claude Opus 5 (84.64%) and Kimi K3 (80.90%). We do not use these results to rank models overall.

57 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated August 19, 2026

Vals Terminal-Bench 2.1 mirror score table (57 models)

Score
1
GPT-5.6 SolOpenAI · Closedmax reasoning
85.77%
2
Claude Opus 5Anthropic · Closed
84.64%
3
Kimi K3Moonshot AI · Closed
80.90%
4
Claude Fable 5Anthropic · Closed
80.52%
5
GPT-5.6 LunaOpenAI · Closedmax reasoning
79.03%
6
Grok 4.6xAI · Closedhigh reasoning
78.28%
7
Gemini 3.7 FlashGoogle · Closedhigh reasoning
77.53%
8
GPT-5.6 TerraOpenAI · Closedmax reasoning
77.53%
9
GPT-5.5OpenAI · Closedhigh reasoning
76.40%
10
Claude Sonnet 5Anthropic · Closed
74.53%
11
Gemini 3.5 FlashGoogle · Closedhigh reasoning
74.16%
12
Gemini 3.6 FlashGoogle · Closedhigh reasoning
73.78%
13
Claude Opus 4.8Anthropic · Closed
71.91%
14
GLM-5.3Z.AI · Closedmax reasoning
71.54%
15
Gemini 3.1 Pro PreviewGooglehigh reasoning
70.79%
16
Claude Opus 4.8 Claude CodeAnthropicClaude CodeClaude Code
69.66%
17
Muse Spark 1.2Meta · Closedxhigh reasoning
69.66%
18
Muse Spark 1.1Meta · Closedxhigh reasoning
69.29%
19
Claude Opus 4.7Anthropic · Closed
68.54%
20
Grok 4.5xAI · Closedhigh reasoning
67.79%
21
GLM-5.2Z.AI · Open weightmax reasoning
67.79%
22
Qwen3.8 MaxAlibaba · Open weight
67.42%
23
DeepSeek V4 Flash 0731DeepSeek · Closed
67.04%
24
Kimi K2.7 CodeMoonshot AI · Open weight
67.04%
25
Qwen3.7 MaxAlibaba · Closed
61.05%
26
MiMo-V2.5Xiaomi · Closed
60.67%
27
Composer 2.5Cursor · ClosedCursor CLICursor CLI
58.43%
28
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
58.43%
29
GPT-5.5 CodexOpenAICodexCodex
57.30%
30
GPT-5.5 FactoryOpenAIFactoryFactory
57.30%
31
Claude Sonnet 4.6Anthropic · Closed
57.30%
32
MiMo-V2.5-ProXiaomi · Closed
57.30%
33
GLM-5.1Z.AI · Open weight
56.93%
34
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
55.06%
35
GPT-5.4 miniOpenAI · Closedhigh reasoning
54.68%
36
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
54.68%
37
Gemini 3 Flash PreviewGooglehigh reasoning
53.93%
38
Kimi K2.6Moonshot AI · Open weight
53.56%
39
MiniMax M3MiniMax · Open weight
53.56%
40
Qwen3.6 PlusAlibaba · Closed
53.18%
41
Qwen3.7 PlusAlibaba · Closed
52.81%
43
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
50.19%
45
DeepSeek V4 Pro 0813DeepSeek · Closed
50.19%
46
MiniMax M2.7MiniMax · Open weight
48.69%
47
InklingThinking Machines Lab · Open weight0.99 reasoning
47.57%
49
Claude Haiku 4.5 ThinkingAnthropic · Closed
43.82%
50
Grok 4.3xAI · Closedhigh reasoning
41.95%
51
41.95%
52
GPT-5.4 nanoOpenAI · Closedhigh reasoning
41.57%
53
Mistral Medium 3.5Mistral AIhigh reasoning
38.95%
54
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
34.08%
55
Laguna M.1Poolside · Closed
34.08%
56
Laguna XS.2Poolside · Open weight
25.84%
57
10.86%

The published Terminal-Bench 2.1 snapshot places GPT-5.6 Sol first at 85.77%. The third row is 4.87 points behind. The broader top-10 range is 11.24 points, so the table still separates the published systems.

57 models have been evaluated on Terminal-Bench 2.1. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Terminal-Bench 2.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench 2.1

Year

2026

Tasks

Terminal-based task execution

Format

Accuracy score

Difficulty

Frontier terminal-agent execution

BenchLM mirrors the public Vals AI Terminal-Bench 2.1 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

BenchLM freshness & provenance

Version

Terminal-Bench 2.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Terminal-Bench 2.1 measure?

State-of-the-art set of difficult terminal-based tasks

Which model leads the published Terminal-Bench 2.1 snapshot?

GPT-5.6 Sol currently leads the published Terminal-Bench 2.1 snapshot with 85.77% vals terminal-bench 2.1 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench 2.1?

The August 19, 2026 contains 57 AI models.

Last updated: August 19, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.