CharXiv Reasoning (CharXiv)
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Data verified 25 confirmed releases in the last 30 daysStart free briefTop models on CharXiv — August 21, 2026
As of August 21, 2026, Claude Mythos 5 leads the CharXiv leaderboard with 93.5% , followed by Qwen3.8 Max (93.5%) and Kimi K3 (91.3%).
Claude Mythos 5
Anthropic
Qwen3.8 Max
Alibaba
Kimi K3
Moonshot AI
Leaderboard (33 models)
ScoreAccording to BenchLM.ai, Claude Mythos 5 leads the CharXiv benchmark with a score of 93.5%, followed by Qwen3.8 Max (93.5%) and Kimi K3 (91.3%). The top models are clustered within 2.2 points, suggesting this benchmark is nearing saturation for frontier models.
33 models have been evaluated on CharXiv. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, CharXiv contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.
About CharXiv
Year
2024
Tasks
Scientific chart reasoning
Format
Chart understanding and reasoning
Difficulty
Scientific visualization reasoning
CharXiv evaluates a model's ability to reason about real-world scientific charts rather than simple visual QA. With-tools and without-tools variants isolate raw visual reasoning from tool-augmented performance.
BenchLM freshness & provenance
Version
CharXiv 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does CharXiv measure?
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Which model scores highest on CharXiv?
Claude Mythos 5 by Anthropic currently leads with a score of 93.5% on CharXiv.
How many models are evaluated on CharXiv?
33 AI models have been evaluated on CharXiv on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.