# ZeroBench_main with Python (ZeroBench w/ Python)

> A Python-assisted ZeroBench_main evaluation reported as pass@5.

Canonical page: https://benchlm.ai/benchmarks/zerobenchpython

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About ZeroBench w/ Python

- Year: 2026
- Tasks: Visual reasoning questions with Python
- Format: Pass@5
- Difficulty: Tool-augmented visual reasoning
- Paper: [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3)

Moonshot follows the official ZeroBench setting and runs it five times. BenchLM keeps this tool-assisted result separate from the non-Python ZeroBench key.

ZeroBench w/ Python is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 49.0% |
| 2 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 49.0% |
| 3 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 41.0% |

## FAQ

### What does ZeroBench w/ Python measure?

A Python-assisted ZeroBench_main evaluation reported as pass@5.

### Which model scores highest on ZeroBench w/ Python?

Qwen3.8 Max by Alibaba currently leads with a score of 49.0% on ZeroBench w/ Python.

### How many models are evaluated on ZeroBench w/ Python?

3 AI models have been evaluated on ZeroBench w/ Python on BenchLM.

### Does ZeroBench w/ Python affect BenchLM's overall score?

Not directly. ZeroBench w/ Python is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on ZeroBench w/ Python

- [Qwen3.8 Max vs DeepSeek V4.1 Flash](/compare/deepseek-v4-1-flash-vs-qwen3-8-max)
- [DeepSeek V4.1 Flash vs Kimi K3](/compare/deepseek-v4-1-flash-vs-kimi-k3)
