RefCOCO average (RefCOCO (avg))
We show this table for reference; we do not rank on it.
A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.
Benchmark score on RefCOCO (avg) — September 27, 2026
We compile the RefCOCO (avg) rows from provider self-reports. Qwen3.6-27B leads the table at 92.5%, followed by Qwen3.6-35B-A3B (92.0%) and Nemotron 3 Nano Omni 30B A3B (90.5%). We do not use these results to rank models overall.
Qwen3.6-27B
Alibaba
Qwen3.6-35B-A3B
Alibaba
Nemotron 3 Nano Omni 30B A3B
NVIDIA
6 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 27, 2026
Benchmark score table (6 models)
ScoreAmong the reported RefCOCO (avg) rows, Qwen3.6-27B is first at 92.5%. The third row is 2.0 points behind. The broader top-10 range is 10.4 points, so the table still separates the published systems.
6 models have been evaluated on RefCOCO (avg). The benchmark falls in the Multimodal & Grounded category. RefCOCO (avg) is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About RefCOCO (avg)
Year
2026
Tasks
Referring-expression grounding
Format
Grounded visual localization
Difficulty
Fine-grained visual grounding
RefCOCO-style tasks matter for grounding-heavy assistants because they measure whether the model can map language to specific objects or regions instead of only answering abstract questions. BenchLM stores provider-reported aggregate RefCOCO values as a display-only grounding row.
Freshness and provenance
Version
RefCOCO (avg) 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does RefCOCO (avg) measure?
A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.
Which model scores highest on RefCOCO (avg)?
Qwen3.6-27B by Alibaba currently leads with a score of 92.5% on RefCOCO (avg).
How many models are evaluated on RefCOCO (avg)?
6 AI models have been evaluated on RefCOCO (avg) on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.