Skip to main content
BenchLM

CharXiv Reasoning without tools (CharXiv w/o tools)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

Benchmark score on CharXiv w/o tools — September 27, 2026

We compile the CharXiv w/o tools rows from provider self-reports. Claude Mythos 5 leads the table at 88.9%, followed by Qwen3.8 Max (88.4%) and Gemini 3.8 Flash (86.2%). We do not use these results to rank models overall.

14 modelsMultimodal & GroundedRefreshingDisplay onlyUpdated September 27, 2026

Benchmark score table (14 models)

Score
1
Claude Mythos 5Anthropic · Closed
88.9%
2
Qwen3.8 MaxAlibaba · Open weight
88.4%
3
Gemini 3.8 FlashGoogle · Closed
86.2%
4
Kimi K3Moonshot AI · Closed
84.8%
5
Qwen3.8-Flash-NextAlibaba · Open weight
84.6%
6
Gemini 3.7 FlashGoogle · Closed
84.5%
7
Qwen3.8-27BAlibaba · Open weight
83.7%
8
Qwen3.8-Omni-FlashAlibaba · Closed
83.5%
9
dots3-note PreviewDots Studio · Open weight
83.1%
10
Claude Opus 4.7 (Adaptive)Anthropic · Closed
82.1%
11
Claude Opus 4.8Anthropic · Closed
80.5%
12
InklingThinking Machines Lab · Open weight
78.1%
13
Inkling-SmallThinking Machines Lab · Open weight
77.4%
14
Claude Sonnet 5Anthropic · Closed
77%

Among the reported CharXiv w/o tools rows, Claude Mythos 5 is first at 88.9%. The third row is 2.7 points behind. The broader top-10 range is 6.8 points, so many of the published results sit in a relatively narrow band.

14 models have been evaluated on CharXiv w/o tools. The benchmark falls in the Multimodal & Grounded category. CharXiv w/o tools is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About CharXiv w/o tools

Year

2024

Tasks

Scientific chart reasoning (tool-free)

Format

Chart understanding without tools

Difficulty

Scientific visualization reasoning

The tool-free CharXiv variant measures pure multimodal reasoning. Mythos Preview scores 86.1% without tools vs 93.2% with tools, demonstrating strong baseline chart reasoning.

Freshness and provenance

Version

CharXiv w/o tools 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does CharXiv w/o tools measure?

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

Which model scores highest on CharXiv w/o tools?

Claude Mythos 5 by Anthropic currently leads with a score of 88.9% on CharXiv w/o tools.

How many models are evaluated on CharXiv w/o tools?

14 AI models have been evaluated on CharXiv w/o tools on BenchLM.

Last updated: September 27, 2026 · BenchLM version CharXiv w/o tools 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.