Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Vals-hosted Terminal-Bench 1.0 mirror (Vals Terminal-Bench 1.0 mirror)

Vals AI hosted Terminal-Bench 1.0 view with easy, medium, and hard task splits.

Data verified 28 confirmed releases in the last 30 daysStart free brief

How BenchLM shows Vals Terminal-Bench 1.0 mirror

BenchLM mirrors the public Vals AI Vals Terminal-Bench 1.0 mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench and updated by Vals on January 12, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench 1.0 mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

47 Vals rows4 task viewspublic datasetTasks: Overall, Easy, Medium, HardDisplay only

Vals Terminal-Bench 1.0 mirror score on Vals Terminal-Bench 1.0 mirror — January 12, 2026

BenchLM mirrors the published vals terminal-bench 1.0 mirror score view for Vals Terminal-Bench 1.0 mirror. GPT-5.2 leads the public snapshot at 63.75% , followed by Claude Sonnet 4.5 20250929 Thinking (61.25%) and Gemini 3 Flash Preview (60.00%). We do not use these results to rank models overall.

47 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated January 12, 2026

Vals Terminal-Bench 1.0 mirror score table (47 models)

Score
1
GPT-5.2OpenAI
63.75%
4
58.75%
6
57.50%
7
56.25%
8
53.75%
11
DeepSeek V3p2Fireworks
50.00%
12
50.00%
13
GPT-5OpenAI
48.75%
14
GPT-5.1OpenAI
47.50%
16
43.75%
17
42.50%
18
41.25%
19
41.25%
20
41.25%
21
DeepSeek V3p1Fireworks
41.25%
22
Kimi K2 ThinkingMoonshot AI
40.00%
23
40.00%
25
40.00%
26
38.75%
27
37.50%
28
36.25%
29
Qwen3 MaxAlibaba
36.25%
30
GPT-4.1OpenAI
33.75%
32
30.00%
34
28.75%
37
GPT Oss 120bFireworks
22.50%
39
21.25%
41
20.00%
44
15.00%
45
DeepSeek R1Fireworks
13.75%
46
6.25%
47
6.25%

The published Vals Terminal-Bench 1.0 mirror snapshot places GPT-5.2 first at 63.75%. The third row is 3.75 points behind. The broader top-10 range is 13.75 points, so the table still separates the published systems.

47 models have been evaluated on Vals Terminal-Bench 1.0 mirror. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vals Terminal-Bench 1.0 mirror is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vals Terminal-Bench 1.0 mirror

Year

2026

Tasks

Terminal task difficulty splits

Format

Accuracy score

Difficulty

Terminal-based agent execution

BenchLM mirrors this Vals-hosted Terminal-Bench 1.0 view as display-only historical context. It remains separate from benchmark-native Terminal-Bench records and weighted rankings.

BenchLM freshness & provenance

Version

Vals Terminal-Bench 1.0 mirror 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Vals Terminal-Bench 1.0 mirror measure?

Vals AI hosted Terminal-Bench 1.0 view with easy, medium, and hard task splits.

Which model leads the published Vals Terminal-Bench 1.0 mirror snapshot?

GPT-5.2 currently leads the published Vals Terminal-Bench 1.0 mirror snapshot with 63.75% vals terminal-bench 1.0 mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vals Terminal-Bench 1.0 mirror?

47 AI models are included in BenchLM's mirrored Vals Terminal-Bench 1.0 mirror snapshot, based on the public leaderboard captured on January 12, 2026.

Last updated: January 12, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.