ResearchClawBench
We show this table for reference; we do not rank on it.
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
RADS average (Pass@1) on ResearchClawBench — 2026-06-29 snapshot
We mirror the published rads average (pass@1) view for ResearchClawBench. Qiushi Engine leads the public snapshot at 38.6%, followed by GPT-5.5 (34.9%) and Ligase (34.8%). We do not use these results to rank models overall.
Qiushi Engine
InternScience
GPT-5.5
OpenAI
Ligase
InternScience
36 modelsAgenticCurrentDisplay onlyUpdated 2026-06-29 snapshot
RADS average (Pass@1) table (36 models)
ScoreHow ResearchClawBench is shown here
BenchLM mirrors the official ResearchClawBench Pass@1 leaderboard snapshot. The source benchmark contains 40 tasks across 10 scientific domains and uses RADS average (Pass@1) as the primary metric.
ResearchClawBench gives agents related literature and raw data, hides the target paper, and grades how much of the scientific result they rediscover. The RADS scale treats 50 as matching the original paper and 70+ as surpassing it.
ResearchClawBench is display only on BenchLM. The rows combine a model, a research harness, execution budget, and long-horizon scientific workflow, so BenchLM does not use them as weighted base-model ranking inputs.
Snapshot
The published ResearchClawBench snapshot places Qiushi Engine first at 38.6%. The third row is 3.8 points behind. The broader top-10 range is 17.9 points, so the table still separates the published systems.
36 models have been evaluated on ResearchClawBench. The benchmark falls in the Agentic category. ResearchClawBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ResearchClawBench
Year
2026
Tasks
40 tasks across 10 scientific domains
Format
End-to-end autonomous research evaluation with RADS scoring
Difficulty
Scientific research re-discovery
ResearchClawBench grades scientific agents with RADS, a rubric where 50 indicates matching the target paper and 70+ indicates surpassing it. BenchLM mirrors the official Pass@1 leaderboard as display-only because rows reflect a research-agent harness and long-horizon scientific workflow, not a normalized base-model-only comparison.
Freshness and provenance
Version
ResearchClawBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ResearchClawBench measure?
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
Which model leads the published ResearchClawBench snapshot?
Qiushi Engine currently leads the published ResearchClawBench snapshot with 38.6% rads average (pass@1). BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on ResearchClawBench?
The 2026-06-29 snapshot snapshot contains 36 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.