# Scientific Code Benchmark (SciCode)

> SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.

Canonical page: https://benchlm.ai/benchmarks/scicode

- Category: [Coding](/coding)
- Last updated: September 15, 2026

## About SciCode

- Year: 2024
- Tasks: 80

SciCode is currently weighted in BenchLM's scoring formula. The Coding category carries 20% of the overall score, and SciCode contributes 10% of that category score.

## Leaderboard (26 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Sakana Fugu](/models/sakana-fugu) | Sakana AI | 60.1% |
| 2 | [Sakana Fugu-Ultra](/models/sakana-fugu-ultra) | Sakana AI | 58.7% |
| 3 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 53.5% |
| 4 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 53.1% |
| 5 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 52.2% |
| 6 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 51.3% |
| 7 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 48.7% |
| 8 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 48.7% |
| 9 | [Grok 4.3](/models/grok-4-3) | xAI | 47.3% |
| 10 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 47% |
| 11 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 44.6% |
| 12 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 43.6% |
| 13 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 41.2% |
| 14 | [Hy3 Preview](/models/hy3-preview) | Tencent | 41.2% |
| 15 | [A.X K2](/models/a-x-k2) | SK Telecom | 41% |
| 16 | [Ling 3.0 Flash FP8](/models/ling-3-0-flash-fp8) | InclusionAI | 40.4% |
| 17 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 38.8% |
| 18 | [Mercury 2.5](/models/mercury-2-5) | Inception | 38% |
| 19 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 37.4% |
| 20 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 36.1% |
| 21 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 32% |
| 22 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 31.4% |
| 23 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 29.6% |
| 24 | [Ling 2.6 Flash](/models/ling-2-6-flash) | InclusionAI | 27% |
| 25 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 26.3% |
| 26 | [Granite 4.2 3B](/models/granite-4-2-3b) | IBM | 24.1% |

## FAQ

### What does SciCode measure?

SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.

### Which model scores highest on SciCode?

Sakana Fugu by Sakana AI currently leads with a score of 60.1% on SciCode.

### How many models are evaluated on SciCode?

26 AI models have been evaluated on SciCode on BenchLM.

### Does SciCode affect BenchLM's overall score?

Yes. SciCode is a weighted benchmark inside the Coding category, which carries 20% of BenchLM's overall score. SciCode itself contributes 10% of that category score.

## Compare Top Models on SciCode

- [Sakana Fugu vs Sakana Fugu-Ultra](/compare/sakana-fugu-vs-sakana-fugu-ultra)
- [Sakana Fugu-Ultra vs Qwen3.7 Max](/compare/qwen3-7-max-vs-sakana-fugu-ultra)
- [Qwen3.7 Max vs Gemini 3.5 Flash](/compare/gemini-3-5-flash-vs-qwen3-7-max)
- [Gemini 3.5 Flash vs Kimi K2.6](/compare/gemini-3-5-flash-vs-kimi-2-6)
