Terminal-Bench-Science 0.1
A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.
How to read this leaderboard
Operator receipt: 9 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 at 30.0%.
Honest limit: Each result combines a model, agent harness, and science-specific environment. The 70-task release is deliberately difficult and continuously maintained, so read the table as a versioned system evaluation rather than a base-model capability score.
How we show Terminal-Bench-Science 0.1
We mirror the official Terminal-Bench-Science 0.1 leaderboard across 70 expert-curated research workflows in the life, physical, Earth, mathematical, and engineering sciences. Each system runs 3 independent trials per task and is graded on concrete artifacts such as analyses, simulations, proofs, code, and data products.
Each row combines a model with an agent harness. The benchmark has its own task set and protocol, so we keep it as a standalone display-only surface outside the external-consensus scoring pipeline.
Snapshot
Resolution rate on Terminal-Bench-Science 0.1 — August 29, 2026 snapshot
We mirror the published resolution rate view for Terminal-Bench-Science 0.1. Claude Opus 5 leads the public snapshot at 30.0%, followed by GPT-5.6 Sol (22.4%) and Claude Fable 5 (21.4%). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
Claude Code · max reasoning
GPT-5.6 Sol
OpenAI
Codex · max reasoning
Claude Fable 5
Anthropic
Claude Code · max reasoning
Resolution rate table (9 models)
ScoreThe published Terminal-Bench-Science 0.1 snapshot places Claude Opus 5 first at 30.0%. The third row is 8.6 points behind. The broader top-10 range is 26.7 points, so the table still separates the published systems.
9 models have been evaluated on Terminal-Bench-Science 0.1. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Terminal-Bench-Science 0.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Terminal-Bench-Science 0.1
Year
2026
Tasks
70 expert-curated scientific research workflows
Format
Resolution rate across 3 trials per task
Difficulty
Frontier scientific research workflows
The first release contains 70 workflows contributed and reviewed by researchers across five scientific domains. Each model-and-agent system runs three independent trials per task and produces concrete artifacts graded by reproducible, task-specific tests. We mirror the official system leaderboard as a standalone display-only benchmark.
BenchLM freshness & provenance
Version
Terminal-Bench-Science 0.1.0
Refresh cadence
Continuous
Staleness state
Current
Question availability
70 public task definitions
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Terminal-Bench-Science 0.1 measure?
A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.
Which model leads the published Terminal-Bench-Science 0.1 snapshot?
Claude Opus 5 currently leads the published Terminal-Bench-Science 0.1 snapshot with 30.0% resolution rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Terminal-Bench-Science 0.1?
The August 29, 2026 snapshot contains 9 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.