# CountBench

> A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.

Canonical page: https://benchlm.ai/benchmarks/countbench

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 22, 2026

## About CountBench

- Year: 2026
- Tasks: Visual counting tasks
- Format: Image-grounded counting
- Difficulty: Fine-grained visual perception
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

Counting failures are a common multimodal weakness even in otherwise strong models. CountBench isolates that skill and makes it easy to compare raw perception accuracy across models.

CountBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (4 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 97.8% |
| 2 | [ZAYA1-VL-8B](/models/zaya1-vl-8b) | Zyphra | 88.1% |
| 3 | [LFM2.5-VL-3B](/models/lfm2-5-vl-3b) | LiquidAI | 87.3% |
| 4 | [LFM2.5-VL-450M](/models/lfm2-5-vl-450m) | LiquidAI | 73.3% |

## FAQ

### What does CountBench measure?

A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.

### Which model scores highest on CountBench?

Qwen3.6-27B by Alibaba currently leads with a score of 97.8% on CountBench.

### How many models are evaluated on CountBench?

4 AI models have been evaluated on CountBench on BenchLM.

### Does CountBench affect BenchLM's overall score?

Not directly. CountBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on CountBench

- [Qwen3.6-27B vs ZAYA1-VL-8B](/compare/qwen3-6-27b-vs-zaya1-vl-8b)
- [ZAYA1-VL-8B vs LFM2.5-VL-3B](/compare/lfm2-5-vl-3b-vs-zaya1-vl-8b)
- [LFM2.5-VL-3B vs LFM2.5-VL-450M](/compare/lfm2-5-vl-3b-vs-lfm2-5-vl-450m)
