Skip to main content
BenchLM

Terminal-Bench 3.0

We show this table for reference; we do not rank on it.

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

Current release

Terminal-Bench 4.0 is the current release. It recalibrates task resources, fixes 19 tasks, and removes eight, so its scores are not directly comparable with Terminal-Bench 3.0.

View Terminal-Bench 4.0

Tasks completed on Terminal-Bench 3.0 — September 23, 2026 snapshot

We mirror the published tasks completed view for Terminal-Bench 3.0. Claude Opus 5 leads the public snapshot at 42.7%, followed by GPT-5.6 Sol (34.6%) and Claude Fable 5 (34.0%). We do not weight the raw table directly. A normalized version can contribute through our external agentic consensus gate.

12 modelsAgentic3% of Agentic reference weightCurrentUpdated September 23, 2026 snapshot

Tasks completed table (12 models)

Score
1
Claude Opus 5Anthropic · Closedmax reasoning
42.7%
2
GPT-5.6 SolOpenAI · Closedmax reasoning
34.6%
3
Claude Fable 5Anthropic · Closedmax reasoning
34.0%
4
GLM-5.3Z.AI · Open weightmax reasoning
32.4%
5
Grok 4.6xAI · Closedhigh reasoning
26.5%
6
Claude Opus 4.8Anthropic · Closedmax reasoning
21.1%
7
GPT-5.6 TerraOpenAI · Closedmax reasoning
20.8%
8
18.6%
9
Grok 4.5xAI · Closedxhigh reasoning
15.7%
10
Claude Sonnet 5Anthropic · Closedmax reasoning
14.6%
11
GPT-5.6 LunaOpenAI · Closedmax reasoning
14.3%
12
GLM-5.2Z.AI · Open weightmax reasoning
4.6%

How we use Terminal-Bench 3.0

We mirror the Terminal-Bench 3.0 version 3.0 table: 74 professional computer-work tasks across 7 domains. The Terminal-Bench and Harbor team first released this benchmark under the FrontierBench name.

Each row keeps the harness and settings the source published. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.

The published Terminal-Bench 3.0 snapshot places Claude Opus 5 first at 42.7%. The third row is 8.7 points behind. The broader top-10 range is 28.1 points, so the table still separates the published systems.

12 models have been evaluated on Terminal-Bench 3.0. The benchmark falls in the Agentic category. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking. BenchAlign v5.7 gives Terminal-Bench 3.0 3% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About Terminal-Bench 3.0

Year

2026

Tasks

74 professional computer-work tasks across 7 domains

Format

Task completion rate

Difficulty

Frontier autonomous knowledge work

Terminal-Bench 3.0, formerly FrontierBench, launched as version 0.1 with 74 tasks across seven domains. Each public row combines a model with an agent harness.

Freshness and provenance

Version

Terminal-Bench 3.0

Refresh cadence

Continuous

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Terminal-Bench 3.0 measure?

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

Which model leads the published Terminal-Bench 3.0 snapshot?

Claude Opus 5 currently leads the published Terminal-Bench 3.0 snapshot with 42.7% tasks completed. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.

How many models are evaluated on Terminal-Bench 3.0?

The September 23, 2026 snapshot snapshot contains 12 AI models.

Last updated: September 23, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.