Terminal-Bench 2.1
Terminal-Bench 2.1 results, mostly provider-reported, stored separately from the Terminal-Bench 2.0 lane.
Current release
Terminal-Bench 4.0 is the current release. It changes task resources and the task set, so its freshly run scores are not directly comparable with this Terminal-Bench 2.x table.
View Terminal-Bench 4.0Top models on Terminal-Bench 2.1 — September 25, 2026
As of September 25, 2026, SWE-2 leads the Terminal-Bench 2.1 leaderboard with 92.8% , followed by GPT-5.6 Sol (91.9%) and DeepSeek V4.1 Flash (90.6%).
SWE-2
Cognition
GPT-5.6 Sol
OpenAI
DeepSeek V4.1 Flash
DeepSeek
62 modelsAgentic8% of Agentic reference weightCurrentUpdated September 25, 2026
Leaderboard (62 models)
ScoreAccording to BenchLM.ai, SWE-2 leads the Terminal-Bench 2.1 benchmark with a score of 92.8%, followed by GPT-5.6 Sol (91.9%) and DeepSeek V4.1 Flash (90.6%). The top models are clustered within 2.2 points, suggesting this benchmark is nearing saturation for frontier models.
62 models have been evaluated on Terminal-Bench 2.1. The benchmark falls in the Agentic category. BenchAlign v5.7 gives Terminal-Bench 2.1 8% of the Agentic reference weight, so it moves the Agentic leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.
About Terminal-Bench 2.1
Year
2026
Tasks
Terminal-based software-agent tasks
Format
Interactive task success rate
Difficulty
Professional software engineering
BenchLM stores Terminal-Bench 2.1 results on their own key, apart from Terminal-Bench 2.0, because the two versions are not directly comparable. Harness and effort settings vary by row.
Freshness and provenance
Version
Terminal-Bench 2.1 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Terminal-Bench 2.1 measure?
Terminal-Bench 2.1 results, mostly provider-reported, stored separately from the Terminal-Bench 2.0 lane.
Which model scores highest on Terminal-Bench 2.1?
SWE-2 by Cognition currently leads with a score of 92.8% on Terminal-Bench 2.1.
How many models are evaluated on Terminal-Bench 2.1?
62 AI models have been evaluated on Terminal-Bench 2.1 on BenchLM.
Compare top models on Terminal-Bench 2.1
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.