# Terminal-Bench 4.0 — Mercor evaluation (Terminal-Bench 4.0 (Mercor))

> Completing command-line tasks on the Terminal-Bench 4.0 set. This table shows Mercor-run configurations for reference and is excluded from model rankings.

Canonical page: https://benchlm.ai/benchmarks/mercorterminalbench4

- Category: [Agentic](/agentic)
- Last updated: October 1, 2026 capture

Evaluation results published by [Mercor](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/). Published evaluation aggregates only; task contents and dataset license grants are not included.

We captured Mercor's Terminal-Bench 4.0 results on October 1, 2026. The table preserves 21 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.

This dated source table is display only and does not enter overall or category rankings. Source configurations retain their external identity, including Pro descriptors and different effort settings.

The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.

The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.

The general methodology lists programmatic unit tests (all must pass) and a judge of None. Its run-count label is k = 3. The linked result page takes precedence where its setup differs, as noted above.

- [Mercor Terminal-Bench 4.0 results](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/)
- [Mercor APEX benchmark index](https://www.mercor.com/apex/)
- [Mercor methodology](https://www.mercor.com/apex/methodology/)
- [Source paper](https://arxiv.org/abs/2601.11868)
- [Upstream evaluation code](https://github.com/harbor-framework/terminal-bench)
- [Upstream dataset](https://hub.harborframework.com/datasets/terminal-bench/terminal-bench/4?tab=tasks)

## About Terminal-Bench 4.0 (Mercor)

- Tasks: 66 source-reported tasks
- Format: Pass@1
- Difficulty: Source-specific evaluation
- Paper: [Terminal-Bench 4.0 evaluation results](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/)

The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.

Terminal-Bench 4.0 (Mercor) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (21 configurations)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Opus 5.5](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.9 · 192 reported samples | Anthropic | 58.30% |
| 2 | [GPT-6.1 Sol](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.5 · 196 reported samples | OpenAI | 58.30% |
| 3 | [GPT-6 Astra](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.1 · 198 reported samples | OpenAI | 54.50% |
| 4 | [Fable 5.1](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.6 · 198 reported samples | Anthropic | 53.00% |
| 5 | [Sonnet 5.5](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.9 · 174 reported samples | Anthropic | 52.50% |
| 6 | [Opus 5](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.1 · 198 reported samples | Anthropic | 49.50% |
| 7 | [GPT-6 Sol](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.4 · 191 reported samples | OpenAI | 43.80% |
| 8 | [GLM-5.3](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±10.1 · 198 reported samples | Zhipu | 37.90% |
| 9 | [GPT-5.6 Sol (Pro)](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±8.8 · 198 reported samples | OpenAI | 32.30% |
| 10 | [GLM-5.3-Flash](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±8.8 · 197 reported samples | Zhipu | 27.30% |
| 11 | [GPT-5.6 Terra](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±8.3 · 198 reported samples | OpenAI | 27.30% |
| 12 | [Kimi K3](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±8.6 · 198 reported samples | Kimi | 26.30% |
| 13 | [Grok 4.6](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · xhigh reasoning · Published margin ±8.3 · 198 reported samples | xAI | 25.30% |
| 14 | [GPT-6 Luna](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±7 · 192 reported samples | OpenAI | 13.50% |
| 15 | [DeepSeek-V4-Pro-0813](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±5.3 · 198 reported samples | DeepSeek | 12.60% |
| 16 | [Gemini 3.8 Flash](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · high reasoning · Published margin ±7.1 · 198 reported samples | Google | 12.10% |
| 17 | [GPT-5.6 Luna](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · max reasoning · Published margin ±5.8 · 198 reported samples | OpenAI | 11.60% |
| 18 | [Grok 4.7](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · xhigh reasoning · Published margin ±4.5 · 198 reported samples | xAI | 6.60% |
| 19 | [MiniMax-M3](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · high reasoning · Published margin ±1.3 · 197 reported samples | MiniMax | 1.00% |
| 20 | [Gemini 3.5 Flash Lite](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · high reasoning · Published margin ±0.8 · 198 reported samples | Google | 0.50% |
| 21 | [Inkling](https://www.mercor.com/apex/oss-benchmarks/oss-terminal-bench-4-0-leaderboard/) | mini-swe-agent (1000 max steps, 8 hour time limit) · high reasoning · Published margin ±0.8 · 198 reported samples | Thinking Machines | 0.50% |

## FAQ

### What does Terminal-Bench 4.0 (Mercor) measure?

Completing command-line tasks on the Terminal-Bench 4.0 set. This table shows Mercor-run configurations for reference and is excluded from model rankings.

### Which model leads the published Terminal-Bench 4.0 (Mercor) snapshot?

Opus 5.5 currently leads the published Terminal-Bench 4.0 (Mercor) snapshot with a score of 58.30%.

### How many models are evaluated on Terminal-Bench 4.0 (Mercor)?

The October 1, 2026 capture contains 21 source configurations.

### Does Terminal-Bench 4.0 (Mercor) affect BenchLM's overall score?

Not directly. Terminal-Bench 4.0 (Mercor) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
