Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Terminal-Bench 2.1, Vals AI run (Terminal-Bench 2.1 (Vals))

Vals AI’s independent run of the Terminal-Bench 2.1 terminal-task suite with published easy, medium, and hard splits.

Current release

Terminal-Bench 4.0 is the current release. It changes task resources and the task set, so its freshly run scores are not directly comparable with this Terminal-Bench 2.x table.

View Terminal-Bench 4.0
Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

How BenchLM shows Vals Terminal-Bench 2.1 mirror

BenchLM mirrors the public Vals AI Vals Terminal-Bench 2.1 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-2-1 and updated by Vals on September 10, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 2.1 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

64 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals Terminal-Bench 2.1 mirror score on Terminal-Bench 2.1 (Vals) — September 10, 2026

We mirror the published vals terminal-bench 2.1 mirror score view for Terminal-Bench 2.1 (Vals). GPT-6 Astra leads the public snapshot at 87.3%, followed by GPT-5.6 Sol (85.8%) and Claude Fable 5.1 (85.0%). We do not use these results to rank models overall.

64 modelsAgenticCurrentDisplay onlyUpdated September 10, 2026

Vals Terminal-Bench 2.1 mirror score table (64 models)

Score
1
GPT-6 AstraOpenAI · Closedmax reasoning
87.3%
2
GPT-5.6 SolOpenAI · Closedmax reasoning
85.8%
3
Claude Fable 5.1Anthropic · Closed
85.0%
4
Claude Opus 5Anthropic · Closed
84.6%
5
Gemini 3.8 FlashGoogle · Closedhigh reasoning
81.3%
6
Kimi K3Moonshot AI · Closed
80.9%
7
Claude Fable 5Anthropic · Closed
80.5%
8
GPT-5.6 LunaOpenAI · Closedmax reasoning
79.0%
9
Muse Spark 1.3 MaxMetamax reasoning
79.0%
10
Grok 4.6xAI · Closedhigh reasoning
78.3%
11
Gemini 3.7 FlashGoogle · Closedhigh reasoning
77.5%
12
GPT-5.6 TerraOpenAI · Closedmax reasoning
77.5%
13
GPT-5.5OpenAI · Closedhigh reasoning
76.4%
14
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
74.5%
15
Claude Sonnet 5Anthropic · Closed
74.5%
16
Gemini 3.5 FlashGoogle · Closedhigh reasoning
74.2%
17
Gemini 3.6 FlashGoogle · Closedhigh reasoning
73.8%
18
Muse Spark 1.3Meta · Closedxhigh reasoning
72.3%
19
Claude Opus 4.8Anthropic · Closed
71.9%
20
GLM-5.3Z.AI · Open weightmax reasoning
71.5%
21
Gemini 3.1 Pro PreviewGooglehigh reasoning
70.8%
22
Claude Opus 4.8 Claude CodeAnthropicClaude CodeClaude Code
69.7%
23
Muse Spark 1.2Meta · Closedxhigh reasoning
69.7%
24
Muse Spark 1.1Meta · Closedxhigh reasoning
69.3%
25
Claude Opus 4.7Anthropic · Closed
68.5%
26
Grok 4.5xAI · Closedhigh reasoning
67.8%
27
GLM-5.2Z.AI · Open weightmax reasoning
67.8%
28
Qwen3.8 MaxAlibaba · Open weight
67.4%
29
DeepSeek V4 Flash 0731DeepSeek · Closed
67.0%
30
Kimi K2.7 CodeMoonshot AI · Open weight
67.0%
31
GLM-5.3-FlashZ.AI · Open weightmax reasoning
62.9%
32
Qwen3.7 MaxAlibaba · Closed
61.0%
33
MiMo-V2.5Xiaomi · Closed
60.7%
34
Composer 2.5Cursor · ClosedCursor CLICursor CLI
58.4%
35
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
58.4%
36
GPT-5.5 CodexOpenAICodexCodex
57.3%
37
GPT-5.5 FactoryOpenAIFactoryFactory
57.3%
38
Claude Sonnet 4.6Anthropic · Closed
57.3%
39
MiMo-V2.5-ProXiaomi · Closed
57.3%
40
GLM-5.1Z.AI · Open weight
56.9%
41
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
55.1%
42
GPT-5.4 miniOpenAI · Closedhigh reasoning
54.7%
43
DeepSeek V4 Pro 0813DeepSeek · Closedmax reasoning
54.7%
44
Gemini 3 Flash PreviewGooglehigh reasoning
53.9%
45
Kimi K2.6Moonshot AI · Open weight
53.6%
46
MiniMax M3MiniMax · Open weight
53.6%
47
Qwen3.6 PlusAlibaba · Closed
53.2%
48
Qwen3.7 PlusAlibaba · Closed
52.8%
50
Gemini 3.5 Flash-LiteGoogle · Closedhigh reasoning
50.2%
52
DeepSeek V4 Pro 0813DeepSeek · Closed
50.2%
53
MiniMax M2.7MiniMax · Open weight
48.7%
54
InklingThinking Machines Lab · Open weight0.99 reasoning
47.6%
56
Claude Haiku 4.5 ThinkingAnthropic · Closed
43.8%
57
Grok 4.3xAI · Closedhigh reasoning
41.9%
58
41.9%
59
GPT-5.4 nanoOpenAI · Closedhigh reasoning
41.6%
60
Mistral Medium 3.5Mistral AIhigh reasoning
39.0%
61
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
34.1%
62
Laguna M.1Poolside · Closed
34.1%
63
Laguna XS.2Poolside · Open weight
25.8%

The published Terminal-Bench 2.1 (Vals) snapshot places GPT-6 Astra first at 87.3%. The third row is 2.2 points behind. The broader top-10 range is 9.0 points, so many of the published results sit in a relatively narrow band.

64 models have been evaluated on Terminal-Bench 2.1 (Vals). The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Terminal-Bench 2.1 (Vals) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench 2.1 (Vals)

Year

2026

Tasks

Difficult terminal tasks

Format

Task success rate

Difficulty

Frontier agentic

BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).

BenchLM freshness & provenance

Version

Terminal-Bench 2.1 (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Terminal-Bench 2.1 (Vals) measure?

Vals AI’s independent run of the Terminal-Bench 2.1 terminal-task suite with published easy, medium, and hard splits.

Which model leads the published Terminal-Bench 2.1 (Vals) snapshot?

GPT-6 Astra currently leads the published Terminal-Bench 2.1 (Vals) snapshot with 87.3% vals terminal-bench 2.1 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench 2.1 (Vals)?

The September 10, 2026 snapshot contains 64 AI models.

Last updated: September 10, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.