CharXiv Reasoning (CharXiv)
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Top models on CharXiv — October 7, 2026
As of October 7, 2026, Qwen3.8 Max leads the CharXiv leaderboard with 93.5% , followed by Claude Mythos 5 (93.5%) and Qwen3.8-Omni-Flash (91.4%).
Qwen3.8 Max
Alibaba
Claude Mythos 5
Anthropic
Qwen3.8-Omni-Flash
Alibaba
39 modelsMultimodal & Grounded20% of category scoreRefreshingUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Qwen3.8 MaxAlibaba | 93.5% | Not reported | Open |
| 2 | Claude Mythos 5Anthropic | 93.5% | Not reported | Closed |
| 3 | Qwen3.8-Omni-FlashAlibaba | 91.4% | Not reported | Closed |
| 4 | Kimi K3Moonshot AI | 91.3% | Not reported | Pending |
| 5 | Claude Opus 4.7 (Adaptive)Anthropic | 91% | Not reported | Closed |
| 6 | Qwen3.8-Flash-NextAlibaba | 90.6% | Not reported | Open |
| 7 | Qwen3.8-27BAlibaba | 90.2% | Not reported | Open |
| 8 | Claude Opus 4.8Anthropic | 89.9% | Not reported | Closed |
| 9 | GLM-5.3-FlashZ.AI | 89.4% | Not reported | Open |
| 10 | Gemini 3.7 FlashGoogle | 88.7% | Not reported | Closed |
| 11 | Muse Spark 1.1Meta | 88.4% | Not reported | Closed |
| 12 | Claude Sonnet 5Anthropic | 88.3% | Not reported | Closed |
| 13 | Sakana Fugu-UltraSakana AI | 86.6% | Not reported | Closed |
| 14 | Muse SparkMeta | 86.4% | Not reported | Closed |
| 15 | Qwen3.7 PlusAlibaba | 85.9% | Not reported | Closed |
| 16 | Seed 2.1 ProByteDance | 85.4% | Not reported | Closed |
| 17 | Sakana FuguSakana AI | 85.1% | Not reported | Closed |
| 18 | Gemini 3.5 FlashGoogle | 84.2% | Not reported | Closed |
| 19 | GPT-5.4OpenAI | 82.8% | Not reported | Closed |
| 20 | Seed 2.1 TurboByteDance | 82.5% | Not reported | Closed |
| 21 | GPT-5.2OpenAI | 82.1% | Not reported | Closed |
| 22 | InklingThinking Machines Lab | 82% | Not reported | Open |
| 23 | Qwen3.6 PlusAlibaba | 81.5% | Not reported | Closed |
| 24 | Gemini 3 ProGoogle | 81.4% | Not reported | Closed |
| 25 | Inkling-SmallThinking Machines Lab | 81.3% | Not reported | Open |
| 26 | MiMo-V2.5Xiaomi | 81% | Not reported | Closed |
| 27 | Qwen3.5 397BAlibaba | 80.8% | Not reported | Open |
| 28 | Kimi K2.6Moonshot AI | 80.4% | Not reported | Open |
| 29 | Gemini 3.1 ProGoogle | 80.2% | Not reported | Closed |
| 30 | Muse Glimmer 30BMeta | 78.8% | Not reported | Open |
| 31 | Qwen3.6-27BAlibaba | 78.4% | Not reported | Open |
| 32 | Qwen3.6-35B-A3BAlibaba | 78% | Not reported | Open |
| 33 | Claude Sonnet 4.6Anthropic | 77.4% | Not reported | Closed |
| 34 | Qwen3.5-122B-A10BAlibaba | 77.2% | Not reported | Open |
| 35 | Nemotron 3 Nano Omni 30B A3BNVIDIA | 76.3% | Not reported | Open |
| 36 | Gemini 3.1 Flash-LiteGoogle | 73.2% | Not reported | Closed |
| 37 | Claude Opus 4.5Anthropic | 68.5% | Not reported | Closed |
| 38 | Grok 4.20xAI | 60.9% | Not reported | Closed |
| 39 | Command A+Cohere | 52.7% | Not reported | Open |
According to BenchLM.ai, Qwen3.8 Max leads the CharXiv benchmark with a score of 93.5%, followed by Claude Mythos 5 (93.5%) and Qwen3.8-Omni-Flash (91.4%). The top models are clustered within 2.1 points, suggesting this benchmark is nearing saturation for frontier models.
39 models have been evaluated on CharXiv. The benchmark falls in the Multimodal & Grounded category. The Multimodal & Grounded leaderboard ranks models by a weighted category score, and CharXiv contributes 20% of it. It does not enter the overall BenchAlign v5.8 ranking.
About CharXiv
Year
2024
Tasks
Scientific chart reasoning
Format
Chart understanding and reasoning
Difficulty
Scientific visualization reasoning
CharXiv evaluates a model's ability to reason about real-world scientific charts rather than simple visual QA. With-tools and without-tools variants isolate raw visual reasoning from tool-augmented performance.
Freshness and provenance
Version
CharXiv 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CharXiv measure?
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Which model scores highest on CharXiv?
Qwen3.8 Max by Alibaba currently leads with a score of 93.5% on CharXiv.
How many models are evaluated on CharXiv?
39 AI models have published results on CharXiv in the BenchLM catalog.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.