# CharXiv Reasoning without tools (CharXiv w/o tools)

> Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

Canonical page: https://benchlm.ai/benchmarks/charxivnotools

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About CharXiv w/o tools

- Year: 2024
- Tasks: Scientific chart reasoning (tool-free)
- Format: Chart understanding without tools
- Difficulty: Scientific visualization reasoning
- Paper: [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/)

The tool-free CharXiv variant measures pure multimodal reasoning. Mythos Preview scores 86.1% without tools vs 93.2% with tools, demonstrating strong baseline chart reasoning.

CharXiv w/o tools is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (14 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 88.9% |
| 2 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 88.4% |
| 3 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 86.2% |
| 4 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 84.8% |
| 5 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 84.6% |
| 6 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 84.5% |
| 7 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 83.7% |
| 8 | [Qwen3.8-Omni-Flash](/models/qwen3-8-omni-flash) | Alibaba | 83.5% |
| 9 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 83.1% |
| 10 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 82.1% |
| 11 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 80.5% |
| 12 | [Inkling](/models/inkling) | Thinking Machines Lab | 78.1% |
| 13 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 77.4% |
| 14 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 77% |

## FAQ

### What does CharXiv w/o tools measure?

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

### Which model scores highest on CharXiv w/o tools?

Claude Mythos 5 by Anthropic currently leads with a score of 88.9% on CharXiv w/o tools.

### How many models are evaluated on CharXiv w/o tools?

14 AI models have been evaluated on CharXiv w/o tools on BenchLM.

### Does CharXiv w/o tools affect BenchLM's overall score?

Not directly. CharXiv w/o tools is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on CharXiv w/o tools

- [Claude Mythos 5 vs Qwen3.8 Max](/compare/claude-mythos-5-vs-qwen3-8-max)
- [Qwen3.8 Max vs Gemini 3.8 Flash](/compare/gemini-3-8-flash-vs-qwen3-8-max)
- [Gemini 3.8 Flash vs Kimi K3](/compare/gemini-3-8-flash-vs-kimi-k3)
- [Kimi K3 vs Qwen3.8-Flash-Next](/compare/kimi-k3-vs-qwen3-8-flash-next)
