Skip to main content

Benchmark profile

FrontierBench v0.1

A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor.

How BenchLM uses FrontierBench

BenchLM mirrors the FrontierBench 0.1 launch table: 74 professional computer-work tasks across 7 domains. FrontierBench was built by the team behind Terminal-Bench and Harbor and is hosted by Harbor and the Laude Institute.

The visible leaderboard remains a model-and-agent system comparison because every row names a harness such as Codex, Claude Code, or Cursor CLI. Separately, BenchLM uses the normalized FrontierBench result as one independent agentic consensus family; the ranking pipeline still requires corroboration from another benchmark family and evaluator before the category can affect a model score.

8 launch rows74 tasks7 domainsv0.1Agentic consensus input

Tasks completed on FrontierBench v0.1 — FrontierBench v0.1 launch

BenchLM mirrors the published tasks completed view for FrontierBench v0.1. GPT-5.6 Sol leads the public snapshot at 34.4% , followed by Claude Fable 5 (33.8%) and Claude Opus 4.8 (21.1%). BenchLM does not use these results to rank models overall.

8 modelsAgenticCurrentDisplay onlyUpdated FrontierBench v0.1 launch

Tasks completed table (8 models)

Score
1
GPT-5.6 SolOpenAI · Closed
34.4%
2
Claude Fable 5Anthropic · Closed
33.8%
3
Claude Opus 4.8Anthropic · Closed
21.1%
4
GPT-5.6 TerraOpenAI · Closed
20.8%
5
Grok 4.5xAI · Closed
17.8%
6
Claude Sonnet 5Anthropic · Closed
14.6%
7
GPT-5.6 LunaOpenAI · Closed
14.3%
8
GLM-5.2Z.AI · Open weight
5.1%

The published FrontierBench v0.1 snapshot places GPT-5.6 Sol first at 34.4%. The third row is 13.3 points behind. The broader top-10 range is 29.3 points, so the table still separates the published systems.

8 models have been evaluated on FrontierBench v0.1. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. FrontierBench v0.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About FrontierBench v0.1

Year

2026

Tasks

74 professional computer-work tasks across 7 domains

Format

Task completion rate

Difficulty

Frontier autonomous knowledge work

The v0.1 launch contains 74 tasks across seven domains. Public rows combine a model with an agent harness. BenchLM displays the raw leaderboard and also treats its normalized result as one independent external agentic benchmark family; corroboration rules still prevent FrontierBench from making a category eligible by itself.

BenchLM freshness & provenance

Version

FrontierBench v0.1 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierBench v0.1 measure?

A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor.

Which model leads the published FrontierBench v0.1 snapshot?

GPT-5.6 Sol currently leads the published FrontierBench v0.1 snapshot with 34.4% tasks completed. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on FrontierBench v0.1?

8 AI models are included in BenchLM's mirrored FrontierBench v0.1 snapshot, based on the public leaderboard captured on FrontierBench v0.1 launch.

Last updated: FrontierBench v0.1 launch · mirrored from the public benchmark leaderboard