# Vals-hosted Terminal-Bench 1.0 mirror (Vals Terminal-Bench 1.0 mirror)

> A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.

Canonical page: https://benchlm.ai/benchmarks/valsterminalbench1

- Category: [Agentic](/agentic)
- Last updated: January 12, 2026

## About Vals Terminal-Bench 1.0 mirror

- Year: 2026
- Tasks: Terminal tasks split by easy, medium, and hard difficulty
- Format: Accuracy score
- Difficulty: Terminal-agent execution
- Paper: [Terminal-Bench 1.0](https://www.vals.ai/benchmarks/terminal-bench)

We mirror the 47-row Vals table as historical, display-only evidence. It stays separate from current Terminal-Bench 2.x and 3.0 results because the benchmark version and evaluation setup differ.

Vals Terminal-Bench 1.0 mirror is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (47 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-5.2](/models/gpt-5-2) | xhigh reasoning | OpenAI | 63.75% |
| 2 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | — | Anthropic | 61.25% |
| 3 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | high reasoning | Google | 60.00% |
| 4 | [GPT-5 Codex](https://www.vals.ai/models/openai_gpt-5-codex) | high reasoning | OpenAI | 58.75% |
| 5 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | — | Anthropic | 57.50% |
| 6 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | high reasoning | OpenAI | 57.50% |
| 7 | [Claude Opus 4.5](/models/claude-opus-4-5) | — | Anthropic | 56.25% |
| 8 | [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) | high reasoning | OpenAI | 53.75% |
| 9 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | high reasoning | Google | 51.25% |
| 10 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | — | Anthropic | 50.00% |
| 11 | [DeepSeek V3p2](https://www.vals.ai/models/fireworks_deepseek-v3p2) | none reasoning | Fireworks AI | 50.00% |
| 12 | [GLM-4.7](/models/glm-4-7) | — | Z.AI | 50.00% |
| 13 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | high reasoning | OpenAI | 48.75% |
| 14 | [GPT-5.1](/models/gpt-5-1) | high reasoning | OpenAI | 47.50% |
| 15 | [Claude Sonnet 4 20250514 Thinking](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514-thinking) | — | Anthropic | 45.00% |
| 16 | [Devstral 2512](https://www.vals.ai/models/mistralai_devstral-2512) | — | Mistral AI | 43.75% |
| 17 | [GLM-4.6](/models/glm-4-6) | — | Z.AI | 42.50% |
| 18 | [GLM-4.5](/models/glm-4-5) | — | Z.AI | 41.25% |
| 19 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | — | Google | 41.25% |
| 20 | [MiniMax M2.1](https://www.vals.ai/models/minimax_MiniMax-M2.1) | — | MiniMax | 41.25% |
| 21 | [DeepSeek V3p1](https://www.vals.ai/models/fireworks_deepseek-v3p1) | — | Fireworks AI | 41.25% |
| 22 | [Kimi K2 Thinking](https://www.vals.ai/models/kimi_kimi-k2-thinking) | — | Moonshot AI | 40.00% |
| 23 | [Moonshotai Kimi K2 Instruct](https://www.vals.ai/models/together_moonshotai/Kimi-K2-Instruct) | — | Together AI | 40.00% |
| 24 | [Labs Devstral Small 2512](https://www.vals.ai/models/mistralai_labs-devstral-small-2512) | — | Mistral AI | 40.00% |
| 25 | [DeepSeek V3p2 Thinking](https://www.vals.ai/models/fireworks_deepseek-v3p2-thinking) | high reasoning | Fireworks AI | 40.00% |
| 26 | [Grok 4 0709](https://www.vals.ai/models/grok_grok-4-0709) | — | xAI | 38.75% |
| 27 | [Kimi K2 Instruct 0905](https://www.vals.ai/models/fireworks_kimi-k2-instruct-0905) | — | Fireworks AI | 37.50% |
| 28 | [Qwen3 Max Preview](https://www.vals.ai/models/alibaba_qwen3-max-preview) | — | Alibaba | 36.25% |
| 29 | [Qwen3 Max](/models/qwen3-max) | — | Alibaba | 36.25% |
| 30 | [GPT-4.1](/models/gpt-4-1) | high reasoning | OpenAI | 33.75% |
| 31 | [Gemini 2.5 Flash Preview 09 2025](https://www.vals.ai/models/google_gemini-2.5-flash-preview-09-2025) | — | Google | 31.25% |
| 32 | [GPT-5 mini](/models/gpt-5-mini) | high reasoning | OpenAI | 30.00% |
| 33 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | — | xAI | 28.75% |
| 34 | [Magistral Medium 2509](https://www.vals.ai/models/mistralai_magistral-medium-2509) | — | Mistral AI | 28.75% |
| 35 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | — | xAI | 27.50% |
| 36 | [Gemini 2.5 Flash Preview 09 2025 Thinking](https://www.vals.ai/models/google_gemini-2.5-flash-preview-09-2025-thinking) | — | Google | 26.25% |
| 37 | [GPT-OSS 120B](/models/gpt-oss-120b) | — | OpenAI | 22.50% |
| 38 | [Grok 4.1 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-1-fast-non-reasoning) | — | xAI | 21.25% |
| 39 | [Mistral Large 2512](https://www.vals.ai/models/mistralai_mistral-large-2512) | — | Mistral AI | 21.25% |
| 40 | [Gemini 2.5 Flash Thinking](https://www.vals.ai/models/google_gemini-2.5-flash-thinking) | — | Google | 21.25% |
| 41 | [Grok Code Fast 1](/models/grok-code-fast-1) | — | xAI | 20.00% |
| 42 | [Grok 4 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-fast-non-reasoning) | — | xAI | 18.75% |
| 43 | [Llama4 Maverick Instruct Basic](https://www.vals.ai/models/fireworks_llama4-maverick-instruct-basic) | — | Fireworks AI | 17.50% |
| 44 | [Magistral Small 2509](https://www.vals.ai/models/mistralai_magistral-small-2509) | — | Mistral AI | 15.00% |
| 45 | [DeepSeek-R1](/models/deepseek-r1) | — | DeepSeek | 13.75% |
| 46 | [Command A 03 2025](https://www.vals.ai/models/cohere_command-a-03-2025) | — | Cohere | 6.25% |
| 47 | [Jamba Large 1.7](https://www.vals.ai/models/ai21labs_jamba-large-1.7) | — | AI21 Labs | 6.25% |

## FAQ

### What does Vals Terminal-Bench 1.0 mirror measure?

A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.

### Which model leads the published Vals Terminal-Bench 1.0 mirror snapshot?

GPT-5.2 currently leads the published Vals Terminal-Bench 1.0 mirror snapshot with a score of 63.75%.

### How many models are evaluated on Vals Terminal-Bench 1.0 mirror?

The January 12, 2026 contains 47 AI models.

### Does Vals Terminal-Bench 1.0 mirror affect BenchLM's overall score?

Not directly. Vals Terminal-Bench 1.0 mirror is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
