CharXiv Descriptive and Reasoning Combined (CharXiv (overall))
We show this table for reference; we do not rank on it.
A scientific chart understanding benchmark scored across both the descriptive and reasoning splits rather than the reasoning split alone.
Benchmark score on CharXiv (overall) — September 27, 2026
We compile the CharXiv (overall) rows from provider self-reports. Ternary Bonsai 2 27B leads the table at 80.0%. We do not use these results to rank models overall.
1 modelMultimodal & GroundedRefreshingDisplay onlyUpdated September 27, 2026
Benchmark score table (1 model)
ScoreAbout CharXiv (overall)
Year
2024
Tasks
Scientific chart description and reasoning
Format
Chart understanding and reasoning
Difficulty
Scientific visualization reasoning
BenchLM keeps the combined descriptive-plus-reasoning score on its own display-only key because the descriptive split is markedly easier than the reasoning split. Mixing a combined score into the weighted CharXiv reasoning lane would overstate chart reasoning.
Freshness and provenance
Version
CharXiv (overall) 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does CharXiv (overall) measure?
A scientific chart understanding benchmark scored across both the descriptive and reasoning splits rather than the reasoning split alone.
Which model scores highest on CharXiv (overall)?
Ternary Bonsai 2 27B by Prism ML currently leads with a score of 80.0% on CharXiv (overall).
How many models are evaluated on CharXiv (overall)?
1 AI models have been evaluated on CharXiv (overall) on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.