Skip to main content
BenchLM
Data

CharXiv Reasoning (CharXiv)

Data verified 38 confirmed releases in the last 30 daysFollow model changes

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Top models on CharXiv — October 7, 2026

As of October 7, 2026, Qwen3.8 Max leads the CharXiv leaderboard with 93.5% , followed by Claude Mythos 5 (93.5%) and Qwen3.8-Omni-Flash (91.4%).

39 modelsMultimodal & Grounded20% of category scoreRefreshingUpdated October 7, 2026

CharXiv leaderboard
RankModel / configurationScoreParameters (B)Open / closed
1Qwen3.8 MaxAlibaba
93.5%
Not reportedOpen
2Claude Mythos 5Anthropic
93.5%
Not reportedClosed
3Qwen3.8-Omni-FlashAlibaba
91.4%
Not reportedClosed
4Kimi K3Moonshot AI
91.3%
Not reportedPending
5Claude Opus 4.7 (Adaptive)Anthropic
91%
Not reportedClosed
6Qwen3.8-Flash-NextAlibaba
90.6%
Not reportedOpen
7Qwen3.8-27BAlibaba
90.2%
Not reportedOpen
8Claude Opus 4.8Anthropic
89.9%
Not reportedClosed
9GLM-5.3-FlashZ.AI
89.4%
Not reportedOpen
10Gemini 3.7 FlashGoogle
88.7%
Not reportedClosed
11Muse Spark 1.1Meta
88.4%
Not reportedClosed
12Claude Sonnet 5Anthropic
88.3%
Not reportedClosed
13Sakana Fugu-UltraSakana AI
86.6%
Not reportedClosed
14Muse SparkMeta
86.4%
Not reportedClosed
15Qwen3.7 PlusAlibaba
85.9%
Not reportedClosed
16Seed 2.1 ProByteDance
85.4%
Not reportedClosed
17Sakana FuguSakana AI
85.1%
Not reportedClosed
18Gemini 3.5 FlashGoogle
84.2%
Not reportedClosed
19GPT-5.4OpenAI
82.8%
Not reportedClosed
20Seed 2.1 TurboByteDance
82.5%
Not reportedClosed
21GPT-5.2OpenAI
82.1%
Not reportedClosed
22InklingThinking Machines Lab
82%
Not reportedOpen
23Qwen3.6 PlusAlibaba
81.5%
Not reportedClosed
24Gemini 3 ProGoogle
81.4%
Not reportedClosed
25Inkling-SmallThinking Machines Lab
81.3%
Not reportedOpen
26MiMo-V2.5Xiaomi
81%
Not reportedClosed
27Qwen3.5 397BAlibaba
80.8%
Not reportedOpen
28Kimi K2.6Moonshot AI
80.4%
Not reportedOpen
29Gemini 3.1 ProGoogle
80.2%
Not reportedClosed
30Muse Glimmer 30BMeta
78.8%
Not reportedOpen
31Qwen3.6-27BAlibaba
78.4%
Not reportedOpen
32Qwen3.6-35B-A3BAlibaba
78%
Not reportedOpen
33Claude Sonnet 4.6Anthropic
77.4%
Not reportedClosed
34Qwen3.5-122B-A10BAlibaba
77.2%
Not reportedOpen
35Nemotron 3 Nano Omni 30B A3BNVIDIA
76.3%
Not reportedOpen
36Gemini 3.1 Flash-LiteGoogle
73.2%
Not reportedClosed
37Claude Opus 4.5Anthropic
68.5%
Not reportedClosed
38Grok 4.20xAI
60.9%
Not reportedClosed
39Command A+Cohere
52.7%
Not reportedOpen

According to BenchLM.ai, Qwen3.8 Max leads the CharXiv benchmark with a score of 93.5%, followed by Claude Mythos 5 (93.5%) and Qwen3.8-Omni-Flash (91.4%). The top models are clustered within 2.1 points, suggesting this benchmark is nearing saturation for frontier models.

39 models have been evaluated on CharXiv. The benchmark falls in the Multimodal & Grounded category. The Multimodal & Grounded leaderboard ranks models by a weighted category score, and CharXiv contributes 20% of it. It does not enter the overall BenchAlign v5.8 ranking.

About CharXiv

Year

2024

Tasks

Scientific chart reasoning

Format

Chart understanding and reasoning

Difficulty

Scientific visualization reasoning

CharXiv evaluates a model's ability to reason about real-world scientific charts rather than simple visual QA. With-tools and without-tools variants isolate raw visual reasoning from tool-augmented performance.

Freshness and provenance

Version

CharXiv 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CharXiv measure?

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Which model scores highest on CharXiv?

Qwen3.8 Max by Alibaba currently leads with a score of 93.5% on CharXiv.

How many models are evaluated on CharXiv?

39 AI models have published results on CharXiv in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version CharXiv 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.