Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

CharXiv Reasoning (CharXiv)

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Data verified 25 confirmed releases in the last 30 daysStart free brief

Top models on CharXiv — August 21, 2026

As of August 21, 2026, Claude Mythos 5 leads the CharXiv leaderboard with 93.5% , followed by Qwen3.8 Max (93.5%) and Kimi K3 (91.3%).

33 modelsMultimodal & Grounded25% of category scoreRefreshingUpdated August 21, 2026

Leaderboard (33 models)

Score
1
Claude Mythos 5Anthropic · Closed
93.5%
2
Qwen3.8 MaxAlibaba · Open weight
93.5%
3
Kimi K3Moonshot AI · Closed
91.3%
4
Claude Opus 4.7 (Adaptive)Anthropic · Closed
91%
5
Qwen3.8-27BAlibaba · Open weight
90.2%
6
Claude Opus 4.8Anthropic · Closed
89.9%
7
Muse Spark 1.1Meta · Closed
88.4%
8
Claude Sonnet 5Anthropic · Closed
88.3%
9
Sakana Fugu-UltraSakana AI · Closed
86.6%
10
Muse SparkMeta · Closed
86.4%
11
Qwen3.7 PlusAlibaba · Closed
85.9%
12
Sakana FuguSakana AI · Closed
85.1%
13
Gemini 3.5 FlashGoogle · Closed
84.2%
14
GPT-5.4OpenAI · Closed
82.8%
15
GPT-5.2OpenAI · Closed
82.1%
16
InklingThinking Machines Lab · Open weight
82%
17
Qwen3.6 PlusAlibaba · Closed
81.5%
18
Gemini 3 ProGoogle · Closed
81.4%
19
Inkling-SmallThinking Machines Lab · Open weight
81.3%
20
MiMo-V2.5Xiaomi · Closed
81%
21
Qwen3.5 397BAlibaba · Open weight
80.8%
22
Kimi K2.6Moonshot AI · Open weight
80.4%
23
Gemini 3.1 ProGoogle · Closed
80.2%
24
Muse Glimmer 30BMeta · Open weight
78.8%
25
Qwen3.6-27BAlibaba · Open weight
78.4%
26
Qwen3.6-35B-A3BAlibaba · Open weight
78%
27
Claude Sonnet 4.6Anthropic · Closed
77.4%
28
Qwen3.5-122B-A10BAlibaba · Open weight
77.2%
29
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
76.3%
30
Gemini 3.1 Flash-LiteGoogle · Closed
73.2%
31
Claude Opus 4.5Anthropic · Closed
68.5%
32
Grok 4.20xAI · Closed
60.9%
33
Command A+Cohere · Open weight
52.7%

According to BenchLM.ai, Claude Mythos 5 leads the CharXiv benchmark with a score of 93.5%, followed by Qwen3.8 Max (93.5%) and Kimi K3 (91.3%). The top models are clustered within 2.2 points, suggesting this benchmark is nearing saturation for frontier models.

33 models have been evaluated on CharXiv. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, CharXiv contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.

About CharXiv

Year

2024

Tasks

Scientific chart reasoning

Format

Chart understanding and reasoning

Difficulty

Scientific visualization reasoning

CharXiv evaluates a model's ability to reason about real-world scientific charts rather than simple visual QA. With-tools and without-tools variants isolate raw visual reasoning from tool-augmented performance.

BenchLM freshness & provenance

Version

CharXiv 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does CharXiv measure?

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Which model scores highest on CharXiv?

Claude Mythos 5 by Anthropic currently leads with a score of 93.5% on CharXiv.

How many models are evaluated on CharXiv?

33 AI models have been evaluated on CharXiv on BenchLM.

Last updated: August 21, 2026 · BenchLM version CharXiv 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.