# ZeroBench

> A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.

Canonical page: https://benchlm.ai/benchmarks/zerobench

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 10, 2026

## About ZeroBench

- Year: 2026
- Tasks: 100 visual reasoning questions
- Format: Multi-step visual reasoning
- Difficulty: Tool-augmented visual reasoning
- Paper: [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology)

Meta evaluates ZeroBench on 100 main questions and reports pass@5, using an LLM judge to compare free-form answers against references. BenchLM stores it as a display-only multimodal reasoning benchmark.

ZeroBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 41.0% |
| 2 | [Muse Spark](/models/muse-spark) | Meta | 33.0% |
| 3 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 29.0% |
| 4 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 24.0% |
| 5 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 23.0% |
| 6 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 19.0% |
| 7 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 18.0% |
| 8 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 11.0% |

## FAQ

### What does ZeroBench measure?

A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.

### Which model scores highest on ZeroBench?

GPT-5.4 by OpenAI currently leads with a score of 41.0% on ZeroBench.

### How many models are evaluated on ZeroBench?

8 AI models have been evaluated on ZeroBench on BenchLM.

### Does ZeroBench affect BenchLM's overall score?

Not directly. ZeroBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on ZeroBench

- [GPT-5.4 vs Muse Spark](/compare/gpt-5-4-vs-muse-spark)
- [Muse Spark vs Gemini 3.1 Pro](/compare/gemini-3-1-pro-vs-muse-spark)
- [Gemini 3.1 Pro vs Qwen3.8 Max](/compare/gemini-3-1-pro-vs-qwen3-8-max)
- [Qwen3.8 Max vs Kimi K3](/compare/kimi-k3-vs-qwen3-8-max)
