NOVA-63
We show this table for reference; we do not rank on it.
A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.
Benchmark score on NOVA-63 — October 7, 2026
We compile the NOVA-63 rows from provider self-reports and secondary reports. Qwen3.5 397B leads the table at 59.1%, followed by Qwen3.7 Max (59.0%) and Qwen3.7 Plus (58.8%). We do not use these results to rank models overall.
Qwen3.5 397B
Alibaba
Qwen3.7 Max
Alibaba
Qwen3.7 Plus
Alibaba
7 modelsMultilingualCurrentDisplay onlyUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Qwen3.5 397BAlibaba | 59.1% | Not reported | Open |
| 2 | Qwen3.7 MaxAlibaba | 59.0% | Not reported | Closed |
| 3 | Qwen3.7 PlusAlibaba | 58.8% | Not reported | Closed |
| 4 | Qwen3.6 PlusAlibaba | 57.9% | Not reported | Closed |
| 5 | Claude Opus 4.5Anthropic | 56.7% | Not reported | Closed |
| 6 | Kimi K2.5Moonshot AI | 56.0% | Not reported | Open |
| 7 | GLM-5Z.AI | 55.1% | Not reported | Open |
Among the reported NOVA-63 rows, Qwen3.5 397B is first at 59.1%. The third row is 0.3 points behind. The broader top-10 range is 4.0 points, so many of the published results sit in a relatively narrow band.
7 models have been evaluated on NOVA-63. The benchmark falls in the Multilingual category. NOVA-63 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About NOVA-63
Year
2026
Tasks
Broad multilingual evaluation
Format
Cross-lingual benchmark
Difficulty
Broad multilingual capability
NOVA-63 appears in multilingual comparison tables as a harder broad-language benchmark than standard translated math. BenchLM tracks it as a display-only multilingual capability signal until a cleaner public benchmark specification is available.
Freshness and provenance
Version
NOVA-63 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does NOVA-63 measure?
A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.
Which model scores highest on NOVA-63?
Qwen3.5 397B by Alibaba currently leads with a score of 59.1% on NOVA-63.
How many models are evaluated on NOVA-63?
7 AI models have published results on NOVA-63 in the BenchLM catalog.
Compare top models on NOVA-63
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.