# GBA-Eval

> An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.

Canonical page: https://benchlm.ai/benchmarks/gbaeval

- Category: [Coding](/coding)
- Last updated: May 30, 2026

## About GBA-Eval

- Year: 2026
- Tasks: 27 emulator test cases
- Format: Overall emulator score
- Difficulty: Long-horizon systems programming
- Paper: [GBA-Eval](https://gbaeval.com/)

GBA-Eval evaluates long-horizon coding agents by having them implement a working GBA emulator. The public leaderboard reports overall scores across 27 test cases with token usage and checkpoints preserved in the source JSON feed.

GBA-Eval is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 70.9% |
| 2 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 53.2% |
| 3 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 48.8% |
| 4 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 44.1% |
| 5 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 43.8% |
| 6 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 31.6% |
| 7 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 6.7% |
| 8 | [Grok Build 0.1](/models/grok-build-0-1) | xAI | 2.4% |
| 9 | [MiniMax M3](/models/minimax-m3) | MiniMax | 0.9% |
| 10 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 0.9% |
| 11 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 0.8% |
| 12 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 0.4% |
| 13 | [GLM-5.1](/models/glm-5-1) | Z.AI | 0.0% |
| 14 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 0.0% |

## FAQ

### What does GBA-Eval measure?

An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.

### Which model leads the published GBA-Eval snapshot?

Claude Opus 4.8 currently leads the published GBA-Eval snapshot with a score of 70.9%.

### How many models are evaluated on GBA-Eval?

The May 30, 2026 contains 14 AI models.

### Does GBA-Eval affect BenchLM's overall score?

Not directly. GBA-Eval is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
