# ResearchClawBench

> An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

Canonical page: https://benchlm.ai/benchmarks/researchclawbench

- Category: [Agentic](/agentic)
- Last updated: 2026-06-29 snapshot

## About ResearchClawBench

- Year: 2026
- Tasks: 40 tasks across 10 scientific domains
- Format: End-to-end autonomous research evaluation with RADS scoring
- Difficulty: Scientific research re-discovery
- Paper: [ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research](https://arxiv.org/abs/2606.07591)

ResearchClawBench grades scientific agents with RADS, a rubric where 50 indicates matching the target paper and 70+ indicates surpassing it. BenchLM mirrors the official Pass@1 leaderboard as display-only because rows reflect a research-agent harness and long-horizon scientific workflow, not a normalized base-model-only comparison.

ResearchClawBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (36 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qiushi Engine](https://internscience.github.io/ResearchClawBench-Home/) | InternScience | 38.6% |
| 2 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 34.9% |
| 3 | [Ligase](https://internscience.github.io/ResearchClawBench-Home/) | InternScience | 34.8% |
| 4 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 33.2% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 32.7% |
| 6 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 28.8% |
| 7 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 22.8% |
| 8 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 21.5% |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 21.1% |
| 10 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 20.7% |
| 11 | [GLM-5.2](/models/glm-5-2) | Z.AI | 20.7% |
| 12 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 19.9% |
| 13 | [MiniMax M3](/models/minimax-m3) | MiniMax | 19.8% |
| 14 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 18.8% |
| 15 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 18.7% |
| 16 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 18.4% |
| 17 | [GLM-5.1](/models/glm-5-1) | Z.AI | 18.2% |
| 18 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 18.0% |
| 19 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 18.0% |
| 20 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 17.9% |
| 21 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 17.1% |
| 22 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 17.0% |
| 23 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 16.9% |
| 24 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 16.6% |
| 25 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 16.3% |
| 26 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 15.5% |
| 27 | [MiMo-V2-Pro](/models/mimo-v2-pro) | Xiaomi | 15.3% |
| 28 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 15.3% |
| 29 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 14.2% |
| 30 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 14.0% |
| 31 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 13.6% |
| 32 | [Grok 4.1](/models/grok-4-1) | xAI | 13.5% |
| 33 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 13.3% |
| 34 | [Hy3 Preview](/models/hy3-preview) | Tencent | 12.9% |
| 35 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 12.8% |
| 36 | [Grok 4.3](/models/grok-4-3) | xAI | 12.4% |

## FAQ

### What does ResearchClawBench measure?

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

### Which model leads the published ResearchClawBench snapshot?

Qiushi Engine currently leads the published ResearchClawBench snapshot with a score of 38.6%.

### How many models are evaluated on ResearchClawBench?

The 2026-06-29 snapshot contains 36 AI models.

### Does ResearchClawBench affect BenchLM's overall score?

Not directly. ResearchClawBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
