# CharXiv Descriptive and Reasoning Combined (CharXiv (overall))

> A scientific chart understanding benchmark scored across both the descriptive and reasoning splits rather than the reasoning split alone.

Canonical page: https://benchlm.ai/benchmarks/charxivoverall

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About CharXiv (overall)

- Year: 2024
- Tasks: Scientific chart description and reasoning
- Format: Chart understanding and reasoning
- Difficulty: Scientific visualization reasoning
- Paper: [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/)

BenchLM keeps the combined descriptive-plus-reasoning score on its own display-only key because the descriptive split is markedly easier than the reasoning split. Mixing a combined score into the weighted CharXiv reasoning lane would overstate chart reasoning.

CharXiv (overall) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Prism ML | 80.0% |

## FAQ

### What does CharXiv (overall) measure?

A scientific chart understanding benchmark scored across both the descriptive and reasoning splits rather than the reasoning split alone.

### Which model scores highest on CharXiv (overall)?

Ternary Bonsai 2 27B by Prism ML currently leads with a score of 80.0% on CharXiv (overall).

### How many models are evaluated on CharXiv (overall)?

1 AI models have been evaluated on CharXiv (overall) on BenchLM.

### Does CharXiv (overall) affect BenchLM's overall score?

Not directly. CharXiv (overall) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
