Skip to main content
BenchLM

Vals-hosted Terminal-Bench 1.0 mirror (Vals Terminal-Bench 1.0 mirror)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.

Vals Terminal-Bench 1.0 mirror score on Vals Terminal-Bench 1.0 mirror — January 12, 2026

We mirror the published vals terminal-bench 1.0 mirror score view for Vals Terminal-Bench 1.0 mirror. GPT-5.2 leads the public snapshot at 63.75%, followed by Claude Sonnet 4.5 Thinking (61.25%) and Gemini 3 Flash Preview (60.00%). We do not use these results to rank models overall.

47 modelsAgenticCurrentDisplay onlyUpdated January 12, 2026

Vals Terminal-Bench 1.0 mirror score table (47 models)

Score
1
GPT-5.2OpenAI · Closedxhigh reasoning
63.75%
2
Claude Sonnet 4.5 ThinkingAnthropic · Closed
61.25%
3
Gemini 3 Flash PreviewGooglehigh reasoning
60.00%
4
GPT-5 CodexOpenAIhigh reasoning
58.75%
5
Claude Opus 4.5 ThinkingAnthropic · Closed
57.50%
6
GPT-5.1-CodexOpenAI · Closedhigh reasoning
57.50%
7
Claude Opus 4.5Anthropic · Closed
56.25%
8
GPT-5.1-Codex-MaxOpenAI · Closedhigh reasoning
53.75%
9
Gemini 3 Pro PreviewGooglehigh reasoning
51.25%
10
Claude Haiku 4.5 ThinkingAnthropic · Closed
50.00%
11
DeepSeek V3p2Fireworks AInone reasoning
50.00%
12
GLM-4.7Z.AI · Open weight
50.00%
13
GPT-5OpenAIhigh reasoning
48.75%
14
GPT-5.1OpenAI · Closedhigh reasoning
47.50%
16
Devstral 2512Mistral AI
43.75%
17
GLM-4.6Z.AI · Open weight
42.50%
18
GLM-4.5Z.AI · Closed
41.25%
19
Gemini 2.5 ProGoogle · Closed
41.25%
20
41.25%
21
DeepSeek V3p1Fireworks AI
41.25%
22
Kimi K2 ThinkingMoonshot AI
40.00%
23
40.00%
24
40.00%
25
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
40.00%
26
38.75%
27
37.50%
28
36.25%
29
Qwen3 MaxAlibaba · Closed
36.25%
30
GPT-4.1OpenAI · Closedhigh reasoning
33.75%
32
GPT-5 miniOpenAI · Closedhigh reasoning
30.00%
33
28.75%
34
28.75%
35
27.50%
37
GPT-OSS 120BOpenAI · Open weight
22.50%
39
21.25%
41
Grok Code Fast 1xAI · Closed
20.00%
43
17.50%
44
15.00%
45
DeepSeek-R1DeepSeek · Open weight
13.75%
46
6.25%
47
6.25%

How Vals Terminal-Bench 1.0 mirror is shown here

BenchLM mirrors the public Vals AI Vals Terminal-Bench 1.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench and updated by Vals on January 12, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 1.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

47 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

The published Vals Terminal-Bench 1.0 mirror snapshot places GPT-5.2 first at 63.75%. The third row is 3.75 points behind. The broader top-10 range is 13.75 points, so the table still separates the published systems.

47 models have been evaluated on Vals Terminal-Bench 1.0 mirror. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Terminal-Bench 1.0 mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals Terminal-Bench 1.0 mirror

Year

2026

Tasks

Terminal tasks split by easy, medium, and hard difficulty

Format

Accuracy score

Difficulty

Terminal-agent execution

We mirror the 47-row Vals table as historical, display-only evidence. It stays separate from current Terminal-Bench 2.x and 3.0 results because the benchmark version and evaluation setup differ.

Freshness and provenance

Version

Vals Terminal-Bench 1.0 mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Vals Terminal-Bench 1.0 mirror measure?

A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.

Which model leads the published Vals Terminal-Bench 1.0 mirror snapshot?

GPT-5.2 currently leads the published Vals Terminal-Bench 1.0 mirror snapshot with 63.75% vals terminal-bench 1.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals Terminal-Bench 1.0 mirror?

The January 12, 2026 snapshot contains 47 AI models.

Last updated: January 12, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.