Terminal-Bench 2.1 Extended — Mercor evaluation (Terminal-Bench 2.1 Extended (Mercor))
We show this table for reference; we do not rank on it.
Completing tasks in command-line environments. This table shows Mercor-run configurations for reference and is excluded from model rankings.
Evaluation results published by Mercor. Published evaluation aggregates only; task contents and dataset license grants are not included.
Pass@1 on Terminal-Bench 2.1 Extended (Mercor) — October 1, 2026 capture
We mirror the published pass@1 view for Terminal-Bench 2.1 Extended (Mercor). Opus 5.5 leads the public snapshot at 47.50%, followed by GPT-6.1 Sol (45.80%) and Grok 4.6 (45.50%). We do not use these results to rank models overall.
Opus 5.5
Anthropic
Mercor extended-set run · max reasoning
Published margin ±8.9 · Original set 83.5% ±7.7 · 297 reported samples
GPT-6.1 Sol
OpenAI
Mercor extended-set run · max reasoning
Published margin ±8.9 · Original set 85.4% ±6.4 · 297 reported samples
Grok 4.6
xAI
Mercor extended-set run · xhigh reasoning
Published margin ±8.8 · Original set 84.6% ±5.8 · 297 reported samples
35 configurationsAgenticDated Mercor configurationsDisplay onlyUpdated October 1, 2026 capture
Pass@1 table (35 configurations)
ScoreHow to read this Mercor evaluation
We captured Mercor's Terminal-Bench 2.1 Extended results on October 1, 2026. The table preserves 35 published configurations, their source identifiers, reasoning settings, and disclosed pass@1 values. These are results from Mercor's evaluation setup; we did not rerun them.
The headline uses the 99 Mercor-built tasks. The paired original-set values refer to the separately listed 89-task public benchmark. Missing baselines remain unreported. We keep the extension and original set separate from their upstream benchmark-owner and provider-run tables.
The headline is the source page's selected Pass@1 metric. The task set, harness, tools, and grader belong to this Mercor run and may differ from another evaluator's results.
The row labels retain published error margins and sample counts when supplied. Missing margins and counts remain unreported; a sample count is not used to reconstruct a task denominator. Except where the dedicated source defines an interval, the confidence level is unreported. Equal displayed scores and overlapping margins do not establish statistical ties. Capture time does not establish individual evaluation dates.
The general methodology lists programmatic unit tests (all must pass) and a judge of None. Its run-count label is k = 3. Those notes describe the original public benchmark; extension-specific changes are not established by that reference.
Snapshot
The published Terminal-Bench 2.1 Extended (Mercor) snapshot places Opus 5.5 first at 47.50%. The third row is 2.00 points behind. The broader top-10 range is 10.10 points, so the table still separates the published systems.
35 configurations are shown for Terminal-Bench 2.1 Extended (Mercor). The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Terminal-Bench 2.1 Extended (Mercor) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Terminal-Bench 2.1 Extended (Mercor)
Tasks
99 Mercor tasks; 89 tasks in the separately reported original set
Format
Pass@1
Difficulty
Source-specific evaluation
The selected headline is Pass@1. Source model identifiers, effort settings, published margins, and reported sample counts remain attached to the source configuration. Public-set results and Mercor task extensions occupy separate tables.
Freshness and provenance
Version
Terminal-Bench 2.1 Extended; Mercor Pass@1
Refresh cadence
Reviewed dated source capture
Staleness state
Dated Mercor configurations
Question availability
Upstream task access and terms; only aggregate results mirrored
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Terminal-Bench 2.1 Extended (Mercor) measure?
Completing tasks in command-line environments. This table shows Mercor-run configurations for reference and is excluded from model rankings.
Which model leads the published Terminal-Bench 2.1 Extended (Mercor) snapshot?
Opus 5.5 currently leads the published Terminal-Bench 2.1 Extended (Mercor) snapshot with 47.50% pass@1. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Terminal-Bench 2.1 Extended (Mercor)?
The October 1, 2026 capture snapshot contains 35 source configurations.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.