Skip to main content
BenchLM

Vals Terminal-Bench Science 0.1 (Terminal-Bench Science (Vals))

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Vals AI’s fixed-harness, single-trial run of 70 expert-authored scientific research workflows.

70-task pass@1 on Terminal-Bench Science (Vals) — September 22, 2026

We mirror the published 70-task pass@1 view for Terminal-Bench Science (Vals). GPT-6 Astra leads the public snapshot at 65.71%, followed by Claude Opus 5.5 (48.57%) and Claude Fable 5.1 (34.29%). We do not use these results to rank models overall.

29 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026

70-task pass@1 table (29 models)

Score
1
GPT-6 AstraOpenAI · Closedmax reasoning
65.71%
2
Claude Opus 5.5Anthropic · Closed
48.57%
3
Claude Fable 5.1Anthropic · Closed
34.29%
4
Claude Opus 5Anthropic · Closed
22.86%
5
Muse Spark 1.3 MaxMetamax reasoning
14.29%
6
Claude Fable 5Anthropic · Closed
12.86%
7
GPT-5.6 SolOpenAI · Closedmax reasoning
12.86%
8
Grok 4.7xAI · Closedxhigh reasoning
11.43%
9
Claude Opus 4.8Anthropic · Closed
10.00%
10
GPT-5.6 TerraOpenAI · Closedmax reasoning
10.00%
11
Gemini 3.8 FlashGoogle · Closedhigh reasoning
8.57%
12
Muse Spark 1.3Meta · Closedxhigh reasoning
8.57%
13
Gemini 3.7 FlashGoogle · Closedhigh reasoning
5.71%
14
GPT-5.6 LunaOpenAI · Closedmax reasoning
5.71%
15
Gemini 3.5 FlashGoogle · Closedhigh reasoning
4.29%
16
Grok 4.6xAI · Closedhigh reasoning
4.29%
17
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
4.29%
18
GLM-5.3Z.AI · Open weightmax reasoning
4.29%
19
Claude Sonnet 5Anthropic · Closed
2.86%
20
MiMo-V2.6-ProXiaomi · Open weight
2.86%
21
Kimi K3Moonshot AI · Closedmax reasoning
2.86%
22
Gemini 3.6 FlashGoogle · Closedhigh reasoning
1.43%
23
Qwen3.8 MaxAlibaba · Open weight
1.43%
24
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
1.43%
25
Hy4 previewTencent · Open weight
1.43%
26
Mercury 2.5Inception · Closedhigh reasoning
0.00%
27
DeepSeek V4 Flash 0731DeepSeek · Open weighthigh reasoning
0.00%
28
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
0.00%
29
GLM-5.3-FlashZ.AI · Open weightmax reasoning
0.00%

How Vals Terminal-Bench Science mirror is shown here

BenchLM mirrors the public Vals AI Vals Terminal-Bench Science mirror leaderboard captured from https://www.vals.ai/benchmarks/terminal-bench-science and updated by Vals on September 22, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Vals Terminal-Bench Science mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

29 Vals rows6 task viewspublic datasetTasks: Overall, Life Sciences, Physical Sciences, Earth Sciences, Mathematical SciencesDisplay only

The published Terminal-Bench Science (Vals) snapshot places GPT-6 Astra first at 65.71%. The third row is 31.43 points behind. The broader top-10 range is 55.71 points, so the table still separates the published systems.

29 models have been evaluated on Terminal-Bench Science (Vals). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Terminal-Bench Science (Vals) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench Science (Vals)

Year

2026

Tasks

70 scientific research workflows across five domains

Format

Strict pass@1 accuracy

Difficulty

Expert scientific research workflows

Vals runs all 70 version-0.1 tasks through Terminus 2 and reports strict pass@1 by scientific domain. The fixed harness and single trial differ from the official native-agent, three-trial leaderboard; the Vals rows remain display-only and separate.

Freshness and provenance

Version

Terminal-Bench Science (Vals) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Terminal-Bench Science (Vals) measure?

Vals AI’s fixed-harness, single-trial run of 70 expert-authored scientific research workflows.

Which model leads the published Terminal-Bench Science (Vals) snapshot?

GPT-6 Astra currently leads the published Terminal-Bench Science (Vals) snapshot with 65.71% 70-task pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench Science (Vals)?

The September 22, 2026 snapshot contains 29 AI models.

Last updated: September 22, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.