Skip to main content
BenchLM

RefCOCO average (RefCOCO (avg))

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.

Benchmark score on RefCOCO (avg) — September 27, 2026

We compile the RefCOCO (avg) rows from provider self-reports. Qwen3.6-27B leads the table at 92.5%, followed by Qwen3.6-35B-A3B (92.0%) and Nemotron 3 Nano Omni 30B A3B (90.5%). We do not use these results to rank models overall.

6 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (6 models)

Score
1
Qwen3.6-27BAlibaba · Open weight
92.5%
2
Qwen3.6-35B-A3BAlibaba · Open weight
92.0%
3
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
90.5%
4
LFM2.5-VL-3BLiquidAI · Open weight
87.9%
5
ZAYA1-VL-8BZyphra · Open weight
84.3%
6
Interfaze BetaInterfaze · Closed
82.1%

Among the reported RefCOCO (avg) rows, Qwen3.6-27B is first at 92.5%. The third row is 2.0 points behind. The broader top-10 range is 10.4 points, so the table still separates the published systems.

6 models have been evaluated on RefCOCO (avg). The benchmark falls in the Multimodal & Grounded category. RefCOCO (avg) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About RefCOCO (avg)

Year

2026

Tasks

Referring-expression grounding

Format

Grounded visual localization

Difficulty

Fine-grained visual grounding

RefCOCO-style tasks matter for grounding-heavy assistants because they measure whether the model can map language to specific objects or regions instead of only answering abstract questions. BenchLM stores provider-reported aggregate RefCOCO values as a display-only grounding row.

Freshness and provenance

Version

RefCOCO (avg) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does RefCOCO (avg) measure?

A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.

Which model scores highest on RefCOCO (avg)?

Qwen3.6-27B by Alibaba currently leads with a score of 92.5% on RefCOCO (avg).

How many models are evaluated on RefCOCO (avg)?

6 AI models have been evaluated on RefCOCO (avg) on BenchLM.

Last updated: September 27, 2026 · BenchLM version RefCOCO (avg) 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.