# Terminal-Bench 3.0

> A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

Canonical page: https://benchlm.ai/benchmarks/terminal-bench-3

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026 snapshot

Latest rankings: [the latest benchmark rankings](/benchmarks/terminal-bench-4)

## About Terminal-Bench 3.0

- Year: 2026
- Tasks: 74 professional computer-work tasks across 7 domains
- Format: Task completion rate
- Difficulty: Frontier autonomous knowledge work
- Paper: [Terminal-Bench 3.0](https://www.frontierbench.ai/)

Terminal-Bench 3.0, formerly FrontierBench, launched as version 0.1 with 74 tasks across seven domains. Each public row combines a model with an agent harness, so we keep the raw results display-only. Its normalized score counts as one external agentic benchmark family and cannot make a category eligible without independent corroboration.

The raw Terminal-Bench 3.0 system leaderboard is excluded from direct benchmark weighting. Its normalized result contributes one external agentic consensus family, subject to BenchLM's source-diversity and corroboration rules.

## Leaderboard (12 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | max reasoning | Anthropic | 42.7% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 34.6% |
| 3 | [Claude Fable 5](/models/claude-fable) | max reasoning | Anthropic | 34.0% |
| 4 | [GLM-5.3](/models/glm-5-3) | max reasoning | Z.AI | 32.4% |
| 5 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 26.5% |
| 6 | [Claude Opus 4.8](/models/claude-opus-4-8) | max reasoning | Anthropic | 21.1% |
| 7 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | max reasoning | OpenAI | 20.8% |
| 8 | [SWE-1.7 Lightning](https://www.frontierbench.ai/) | — | Cognition | 18.6% |
| 9 | [Grok 4.5](/models/grok-4-5) | xhigh reasoning | xAI | 15.7% |
| 10 | [Claude Sonnet 5](/models/claude-sonnet-5) | max reasoning | Anthropic | 14.6% |
| 11 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 14.3% |
| 12 | [GLM-5.2](/models/glm-5-2) | max reasoning | Z.AI | 4.6% |

## FAQ

### What does Terminal-Bench 3.0 measure?

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

### Which model leads the published Terminal-Bench 3.0 snapshot?

Claude Opus 5 currently leads the published Terminal-Bench 3.0 snapshot with a score of 42.7%.

### How many models are evaluated on Terminal-Bench 3.0?

The September 18, 2026 snapshot contains 12 AI models.

### Does Terminal-Bench 3.0 affect BenchLM's overall score?

Indirectly. The raw Terminal-Bench 3.0 system leaderboard is not a directly weighted benchmark, but its normalized result contributes one external agentic consensus family. It still requires corroboration from another benchmark family and evaluator before it can affect a model's category score.
