Benchmark profile
FrontierBench v0.1
A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor.
How BenchLM uses FrontierBench
BenchLM mirrors the FrontierBench 0.1 launch table: 74 professional computer-work tasks across 7 domains. FrontierBench was built by the team behind Terminal-Bench and Harbor and is hosted by Harbor and the Laude Institute.
The visible leaderboard remains a model-and-agent system comparison because every row names a harness such as Codex, Claude Code, or Cursor CLI. Separately, BenchLM uses the normalized FrontierBench result as one independent agentic consensus family; the ranking pipeline still requires corroboration from another benchmark family and evaluator before the category can affect a model score.
Tasks completed on FrontierBench v0.1 — FrontierBench v0.1 launch
BenchLM mirrors the published tasks completed view for FrontierBench v0.1. GPT-5.6 Sol leads the public snapshot at 34.4% , followed by Claude Fable 5 (33.8%) and Claude Opus 4.8 (21.1%). BenchLM does not use these results to rank models overall.
GPT-5.6 Sol
OpenAI
openai/gpt-5.6-sol:codex
Claude Fable 5
Anthropic
anthropic/claude-fable-5:claude-code
Claude Opus 4.8
Anthropic
anthropic/claude-opus-4-8:claude-code
Tasks completed table (8 models)
ScoreThe published FrontierBench v0.1 snapshot places GPT-5.6 Sol first at 34.4%. The third row is 13.3 points behind. The broader top-10 range is 29.3 points, so the table still separates the published systems.
8 models have been evaluated on FrontierBench v0.1. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. FrontierBench v0.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About FrontierBench v0.1
Year
2026
Tasks
74 professional computer-work tasks across 7 domains
Format
Task completion rate
Difficulty
Frontier autonomous knowledge work
The v0.1 launch contains 74 tasks across seven domains. Public rows combine a model with an agent harness. BenchLM displays the raw leaderboard and also treats its normalized result as one independent external agentic benchmark family; corroboration rules still prevent FrontierBench from making a category eligible by itself.
BenchLM freshness & provenance
Version
FrontierBench v0.1 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does FrontierBench v0.1 measure?
A continuously maintained professional computer-work benchmark from the team behind Terminal-Bench and Harbor.
Which model leads the published FrontierBench v0.1 snapshot?
GPT-5.6 Sol currently leads the published FrontierBench v0.1 snapshot with 34.4% tasks completed. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on FrontierBench v0.1?
8 AI models are included in BenchLM's mirrored FrontierBench v0.1 snapshot, based on the public leaderboard captured on FrontierBench v0.1 launch.