# Vals Terminal-Bench 4.0 (Terminal-Bench 4.0 (Vals))

> Vals AI’s independent run of the Terminal-Bench 4.0 frontier terminal suite, with per-domain splits.

Canonical page: https://benchlm.ai/benchmarks/vals-terminal-bench-4

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Terminal-Bench 4.0 (Vals)

- Year: 2026
- Tasks: Frontier-difficulty terminal tasks across seven domains
- Format: Accuracy score
- Difficulty: Frontier terminal-agent execution
- Paper: [Vals Terminal-Bench 4.0](https://www.vals.ai/benchmarks/terminal-bench-4)

Vals publishes overall accuracy plus software, science, ML, operations, hardware, security, and media splits with standard error, latency, and cost per test. BenchLM mirrors the board on a dedicated key so a Vals run never overwrites the official Terminal-Bench 4.0 rows, and keeps it display only and outside weighted rankings.

Terminal-Bench 4.0 (Vals) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (37 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | — | Anthropic | 61.62% |
| 2 | [GPT-6 Astra](/models/gpt-6-astra) | max reasoning | OpenAI | 57.07% |
| 3 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | — | Anthropic | 53.03% |
| 4 | [Claude Fable 5.1](/models/claude-fable-5-1) | — | Anthropic | 49.49% |
| 5 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 45.45% |
| 6 | [GPT-6 Sol](/models/gpt-6-sol) | max reasoning | OpenAI | 34.34% |
| 7 | [Grok 4.7](/models/grok-4-7) | xhigh reasoning | xAI | 28.28% |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 27.78% |
| 9 | [Muse Spark 1.3 Max](https://www.vals.ai/models/meta_muse_spark_1_3_max) | max reasoning | Meta | 27.78% |
| 10 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | max reasoning | OpenAI | 26.26% |
| 11 | [GLM-5.3](/models/glm-5-3) | max reasoning | Z.AI | 25.25% |
| 12 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | — | Xiaomi | 24.75% |
| 13 | [Qwen3.8 Max](/models/qwen3-8-max) | — | Alibaba | 24.75% |
| 14 | [Claude Fable 5](/models/claude-fable) | — | Anthropic | 22.73% |
| 15 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | — | Xiaomi | 21.21% |
| 16 | [GLM-5.3-Flash](/models/glm-5-3-flash) | max reasoning | Z.AI | 19.70% |
| 17 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 17.17% |
| 18 | [Claude Opus 4.8](/models/claude-opus-4-8) | — | Anthropic | 16.16% |
| 19 | [Muse Spark 1.3](/models/muse-spark-1-3) | xhigh reasoning | Meta | 15.15% |
| 20 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | high reasoning | Google | 13.13% |
| 21 | [Kimi K3](/models/kimi-k3) | max reasoning | Moonshot AI | 12.63% |
| 22 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | high reasoning | DeepSeek | 11.62% |
| 23 | [GPT-6 Luna](/models/gpt-6-luna) | max reasoning | OpenAI | 9.60% |
| 24 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | high reasoning | DeepSeek | 9.09% |
| 25 | [Claude Sonnet 5](/models/claude-sonnet-5) | — | Anthropic | 8.08% |
| 26 | [Grok 4.5](/models/grok-4-5) | high reasoning | xAI | 6.57% |
| 27 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | high reasoning | Google | 6.06% |
| 28 | [Muse Spark 1.2](/models/muse-spark-1-2) | xhigh reasoning | Meta | 5.56% |
| 29 | [Hy4 preview](/models/hy4-preview) | — | Tencent | 5.05% |
| 30 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 4.54% |
| 31 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | high reasoning | Google | 4.54% |
| 32 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | high reasoning | Google | 4.04% |
| 33 | [Qwen3.8-27B](/models/qwen3-8-27b) | xhigh reasoning | Alibaba | 4.04% |
| 34 | [Inkling-Small](/models/inkling-small) | 0.99 reasoning | Thinking Machines Lab | 1.51% |
| 35 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 1.01% |
| 36 | [Mercury 2.5](/models/mercury-2-5) | high reasoning | Inception | 0.00% |
| 37 | [Inkling](/models/inkling) | 0.99 reasoning | Thinking Machines Lab | 0.00% |

## FAQ

### What does Terminal-Bench 4.0 (Vals) measure?

Vals AI’s independent run of the Terminal-Bench 4.0 frontier terminal suite, with per-domain splits.

### Which model leads the published Terminal-Bench 4.0 (Vals) snapshot?

Claude Opus 5.5 currently leads the published Terminal-Bench 4.0 (Vals) snapshot with a score of 61.62%.

### How many models are evaluated on Terminal-Bench 4.0 (Vals)?

The September 27, 2026 contains 37 AI models.

### Does Terminal-Bench 4.0 (Vals) affect BenchLM's overall score?

Not directly. Terminal-Bench 4.0 (Vals) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
