# RuneBench / runescape-bench (RuneScape-Bench)

> An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.

Canonical page: https://benchlm.ai/benchmarks/runescapebench

- Category: [external](/external)
- Last updated: May 2026 snapshot

## About RuneScape-Bench

- Year: 2026
- Tasks: 16 RuneScape skill-training tasks
- Format: Average log XP-rate score
- Difficulty: Agentic gameplay automation
- Paper: [RuneBench](https://maxbittker.github.io/runebench/)

RuneBench evaluates gameplay automation and coding-agent strategy. BenchLM mirrors the public aggregate computed as average ln(1 + XP/min) across 16 skill-training tasks, while keeping the benchmark display-only because rows reflect agent harness and gameplay strategy.

RuneScape-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.5 xhigh](https://maxbittker.github.io/runebench/) | OpenAI | 5.7% |
| 2 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 5.4% |
| 3 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 5.3% |
| 4 | [Claude Opus 4.8 max](https://maxbittker.github.io/runebench/) | Anthropic | 5.1% |
| 5 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 5.0% |
| 6 | [Gemini 3.5 Flash high](https://maxbittker.github.io/runebench/) | Google | 4.9% |
| 7 | [Claude Opus 4.7 xhigh](https://maxbittker.github.io/runebench/) | Anthropic | 4.7% |
| 8 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 4.7% |
| 9 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 4.7% |
| 10 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 4.6% |
| 11 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 4.5% |
| 12 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 4.4% |
| 13 | [Codex CLI 5.3](https://maxbittker.github.io/runebench/) | OpenAI | 4.3% |
| 14 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 4.1% |
| 15 | [GPT-5.4 Mini](https://maxbittker.github.io/runebench/) | OpenAI | 4.1% |
| 16 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 3.8% |
| 17 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 3.2% |
| 18 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 3.2% |
| 19 | [GPT-5.4 Nano](https://maxbittker.github.io/runebench/) | OpenAI | 2.3% |
| 20 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 2.1% |
| 21 | [GLM 5](https://maxbittker.github.io/runebench/) | Z.AI | 1.9% |
| 22 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 1.6% |
| 23 | [Qwen3 Max](/models/qwen3-max) | Alibaba | 1.4% |
| 24 | [Qwen3 Coder Next](https://maxbittker.github.io/runebench/) | Alibaba | 1.2% |
| 25 | [Qwen3.5 35B](https://maxbittker.github.io/runebench/) | Alibaba | 0.7% |

## FAQ

### What does RuneScape-Bench measure?

An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.

### Which model leads the published RuneScape-Bench snapshot?

GPT-5.5 xhigh currently leads the published RuneScape-Bench snapshot with a score of 5.7%.

### How many models are evaluated on RuneScape-Bench?

The May 2026 snapshot contains 25 AI models.

### Does RuneScape-Bench affect BenchLM's overall score?

Not directly. RuneScape-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
