Skip to main content
BenchLM

Terminal-Bench 2.1

Data verified 35 confirmed releases in the last 30 daysFollow model changes

Terminal-Bench 2.1 results, mostly provider-reported, stored separately from the Terminal-Bench 2.0 lane.

Current release

Terminal-Bench 4.0 is the current release. It changes task resources and the task set, so its freshly run scores are not directly comparable with this Terminal-Bench 2.x table.

View Terminal-Bench 4.0

Top models on Terminal-Bench 2.1 — September 25, 2026

As of September 25, 2026, SWE-2 leads the Terminal-Bench 2.1 leaderboard with 92.8% , followed by GPT-5.6 Sol (91.9%) and DeepSeek V4.1 Flash (90.6%).

62 modelsAgentic8% of Agentic reference weightCurrentUpdated September 25, 2026

Leaderboard (62 models)

Score
1
SWE-2Cognition · Closed
92.8%
2
GPT-5.6 SolOpenAI · Closed
91.9%
3
DeepSeek V4.1 FlashDeepSeek · Open weight
90.6%
4
MiMo-V2.6-ProXiaomi · Open weight
89.9%
5
Gemini 3.8 FlashGoogle · Closed
89.4%
6
Muse Spark 1.3Meta · Closed
88.8%
7
Kimi K3Moonshot AI · Closed
88.3%
8
GLM-5.3Z.AI · Open weight
88.2%
9
Claude Mythos 5Anthropic · Closed
88.0%
10
DeepSeek V4 Pro 0813DeepSeek · Open weight
87.9%
11
MiMo-V2.6-FlashXiaomi · Open weight
87.6%
12
GPT-5.6 TerraOpenAI · Closed
87.4%
13
Qwen3.8 MaxAlibaba · Open weight
86.6%
14
Ornith-1.5-397BOrnith AI · Open weight
86.1%
15
Gemini 3.7 FlashGoogle · Closed
85.8%
16
Hy4 previewTencent · Open weight
85.4%
17
Step 5 PreviewStepFun · Closed
85.0%
18
GPT-5.6 LunaOpenAI · Closed
84.7%
19
Claude Fable 5Anthropic · Closed
84.3%
20
GLM-5.3-FlashZ.AI · Open weight
84.3%
21
Grok 4.5xAI · Closed
83.3%
22
Muse Spark 1.2Meta · Closed
82.9%
23
DeepSeek V4 Flash 0731DeepSeek · Open weight
82.7%
24
Sakana Fugu-UltraSakana AI · Closed
82.1%
25
SWE-1.7Cognition · Closed
81.5%
26
GLM-5.2Z.AI · Open weight
81.0%
27
Claude Sonnet 5Anthropic · Closed
80.4%
28
Sakana FuguSakana AI · Closed
80.2%
29
Muse Spark 1.1Meta · Closed
80.0%
30
Atria Dawn PreviewShanghai Artificial Intelligence Laboratory · Open weight
78.3%
31
Ornith-1.0-397BDeepReinforce AI · Open weight
77.5%
32
Gemini 3.5 FlashGoogle · Closed
76.2%
33
dots3-note PreviewDots Studio · Open weight
75.1%
34
Claude Opus 4.8Anthropic · Closed
74.6%
35
Qwen3.8-27BAlibaba · Open weight
73.0%
36
Seed 2.1 ProByteDance · Closed
71.0%
37
Apodex 1.1Apodex · Closed
70.8%
38
Laguna S 2.1Poolside · Open weight
70.2%
39
Quasar 438BMultiverse Computing · Closed
69.3%
40
Ornith-1.5-35B-A3BOrnith AI · Open weight
67.8%
41
Seed 2.1 TurboByteDance · Closed
67.6%
42
MiniMax M3MiniMax · Open weight
66.0%
43
Pokee-Isaac 28BPokee AI · Closed
65.1%
44
Inkling-SmallThinking Machines Lab · Open weight
64.7%
45
Ornith-1.0-35BDeepReinforce AI · Open weight
64.2%
46
InklingThinking Machines Lab · Open weight
63.8%
47
MAI-Code-1.1-FlashMicrosoft · Closed
62.9%
48
Step 3.7 FlashStepFun · Open weight
59.5%
49
Ling 3.0 FlashInclusionAI · Open weight
57.0%
50
Solar Pro 4Upstage · Closed
57.0%
51
Nemotron 3 UltraNVIDIA · Open weight
56.4%
52
Gemini 3.5 Flash-LiteGoogle · Closed
54.0%
53
Ternary Bonsai 2 27BPrism ML · Open weight
52.8%
54
Ornith-1.5-9BOrnith AI · Open weight
46.2%
55
K-EXAONE 2.0LG AI Research · Open weight
43.8%
56
Ornith-1.0-9BDeepReinforce AI · Open weight
43.1%
57
A.X K2SK Telecom · Open weight
36.0%
58
Granite 4.2 30BIBM · Open weight
29.2%
59
Mercury 2Inception · Closed
27.0%
60
23.5%
61
Granite 4.2 8BIBM · Open weight
20.6%
62
MiniCPM5-2BOpenBMB · Open weight
8.6%

According to BenchLM.ai, SWE-2 leads the Terminal-Bench 2.1 benchmark with a score of 92.8%, followed by GPT-5.6 Sol (91.9%) and DeepSeek V4.1 Flash (90.6%). The top models are clustered within 2.2 points, suggesting this benchmark is nearing saturation for frontier models.

62 models have been evaluated on Terminal-Bench 2.1. The benchmark falls in the Agentic category. BenchAlign v5.7 gives Terminal-Bench 2.1 8% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.

About Terminal-Bench 2.1

Year

2026

Tasks

Terminal-based software-agent tasks

Format

Interactive task success rate

Difficulty

Professional software engineering

BenchLM stores Terminal-Bench 2.1 results on their own key, apart from Terminal-Bench 2.0, because the two versions are not directly comparable. Harness and effort settings vary by row.

Freshness and provenance

Version

Terminal-Bench 2.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Terminal-Bench 2.1 measure?

Terminal-Bench 2.1 results, mostly provider-reported, stored separately from the Terminal-Bench 2.0 lane.

Which model scores highest on Terminal-Bench 2.1?

SWE-2 by Cognition currently leads with a score of 92.8% on Terminal-Bench 2.1.

How many models are evaluated on Terminal-Bench 2.1?

62 AI models have been evaluated on Terminal-Bench 2.1 on BenchLM.

Last updated: September 25, 2026 · BenchLM version Terminal-Bench 2.1 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.