Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief

Terminal-Bench-Science 0.1

A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.

How to read this leaderboard

Operator receipt: 9 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5 at 30.0%.

Honest limit: Each result combines a model, agent harness, and science-specific environment. The 70-task release is deliberately difficult and continuously maintained, so read the table as a versioned system evaluation rather than a base-model capability score.

How we show Terminal-Bench-Science 0.1

We mirror the official Terminal-Bench-Science 0.1 leaderboard across 70 expert-curated research workflows in the life, physical, Earth, mathematical, and engineering sciences. Each system runs 3 independent trials per task and is graded on concrete artifacts such as analyses, simulations, proofs, code, and data products.

Each row combines a model with an agent harness. The benchmark has its own task set and protocol, so we keep it as a standalone display-only surface outside the external-consensus scoring pipeline.

Resolution rate on Terminal-Bench-Science 0.1 — August 29, 2026 snapshot

We mirror the published resolution rate view for Terminal-Bench-Science 0.1. Claude Opus 5 leads the public snapshot at 30.0%, followed by GPT-5.6 Sol (22.4%) and Claude Fable 5 (21.4%). We do not use these results to rank models overall.

9 modelsAgenticCurrentDisplay onlyUpdated August 29, 2026 snapshot

Resolution rate table (9 models)

Score
1
Claude Opus 5Anthropic · ClosedClaude Code · max reasoning
30.0%
2
GPT-5.6 SolOpenAI · ClosedCodex · max reasoning
22.4%
3
Claude Fable 5Anthropic · ClosedClaude Code · max reasoning
21.4%
4
Claude Opus 4.8Anthropic · ClosedClaude Code · max reasoning
10.5%
5
GPT-5.6 TerraOpenAI · ClosedCodex · max reasoning
8.6%
6
GLM-5.3Z.AI · Open weightClaude Code · max reasoning
8.1%
7
Kimi K3Moonshot AI · ClosedClaude Code · max reasoning
7.1%
8
Grok 4.6xAI · ClosedGrok Build · xhigh reasoning
7.1%
9
GPT-5.6 LunaOpenAI · ClosedCodex · max reasoning
3.3%

The published Terminal-Bench-Science 0.1 snapshot places Claude Opus 5 first at 30.0%. The third row is 8.6 points behind. The broader top-10 range is 26.7 points, so the table still separates the published systems.

9 models have been evaluated on Terminal-Bench-Science 0.1. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. Terminal-Bench-Science 0.1 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Terminal-Bench-Science 0.1

Year

2026

Tasks

70 expert-curated scientific research workflows

Format

Resolution rate across 3 trials per task

Difficulty

Frontier scientific research workflows

The first release contains 70 workflows contributed and reviewed by researchers across five scientific domains. Each model-and-agent system runs three independent trials per task and produces concrete artifacts graded by reproducible, task-specific tests. We mirror the official system leaderboard as a standalone display-only benchmark.

BenchLM freshness & provenance

Version

Terminal-Bench-Science 0.1.0

Refresh cadence

Continuous

Staleness state

Current

Question availability

70 public task definitions

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does Terminal-Bench-Science 0.1 measure?

A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.

Which model leads the published Terminal-Bench-Science 0.1 snapshot?

Claude Opus 5 currently leads the published Terminal-Bench-Science 0.1 snapshot with 30.0% resolution rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Terminal-Bench-Science 0.1?

The August 29, 2026 snapshot contains 9 AI models.

Last updated: August 29, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.