# BabyVision

> A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

Canonical page: https://benchlm.ai/benchmarks/babyvision

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About BabyVision

- Year: 2026
- Tasks: Visual perception tasks
- Format: Multimodal visual reasoning
- Difficulty: Fine-grained visual perception
- Paper: [Muse Spark 1.1 Evaluation Report](https://ai.meta.com/static-resource/muse-spark-1-1-evaluation-report)

BenchLM tracks BabyVision as a display-only multimodal perception benchmark when providers publish exact comparison values. It complements chart-heavy benchmarks such as CharXiv by emphasizing grounded visual understanding.

BabyVision is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (7 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 82.0% |
| 2 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 76.3% |
| 3 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 73.7% |
| 4 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 65.7% |
| 5 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 62.9% |
| 6 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 53.4% |
| 7 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 50.0% |

## FAQ

### What does BabyVision measure?

A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

### Which model scores highest on BabyVision?

Qwen3.8 Max by Alibaba currently leads with a score of 82.0% on BabyVision.

### How many models are evaluated on BabyVision?

7 AI models have been evaluated on BabyVision on BenchLM.

### Does BabyVision affect BenchLM's overall score?

Not directly. BabyVision is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on BabyVision

- [Qwen3.8 Max vs Muse Spark 1.1](/compare/muse-spark-1-1-vs-qwen3-8-max)
- [Muse Spark 1.1 vs Seed 2.1 Pro](/compare/muse-spark-1-1-vs-seed-2-1-pro)
- [Seed 2.1 Pro vs Qwen3.8-27B](/compare/qwen3-8-27b-vs-seed-2-1-pro)
- [Qwen3.8-27B vs Seed 2.1 Turbo](/compare/qwen3-8-27b-vs-seed-2-1-turbo)
