Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief
BenchLM recommendation

Best Multimodal LLMs in 2026

Data verified

As of August 27, 2026, the top model in best multimodal llms on the BenchLM leaderboard is Qwen3.8 Max with a score of 87.1.

Bottom line: Gemini 3 Pro Deep Think leads multimodal understanding at 95, with Claude Mythos 5 and GPT-5.1 behind it. For image generation, look at dedicated image models — this table measures visual understanding.

About this ranking

Last verified: August 27, 2026

This page ranks models by BenchLM's multimodal & grounded category: image understanding, document and chart reading, and visually grounded reasoning. Note the scope: these scores measure how well a model understands images — not how well it generates them. Image-generation systems are a different product class and are not ranked here.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Qwen3.8 Max leads this ranking with a score of 87.1, followed by Kimi K3 (86.2) and Claude Opus 5 (85.9). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Qwen3.8 Max (ranked #1 with a score of 87.1). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional weighted averages across the scoring benchmarks in multimodalGrounded. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

Gemini 3 Pro Deep Think leads multimodal & grounded understanding at 95.

Claude Mythos 5 strong #2 at 92.8 on grounded visual reasoning.

GPT-5.1 still a top-3 multimodal row at 91.8 — and cheap at $1.25/$10.

How to choose

Full Rankings (33 models)

1
Qwen3.8 Max
Alibaba·Open Weight·1M

87.1

prov. avg

2
Kimi K3
Moonshot AI·Pending·1.05M

86.2

prov. avg

3
Claude Opus 5
Anthropic·Proprietary·

85.9

prov. avg

4
Claude Opus 4.8
Anthropic·Proprietary·1M

85.4

prov. avg

5
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

82.7

prov. avg

6
Gemini 3.5 Flash
Google·Proprietary·1M

80.6

prov. avg

7
GLM-5.3-Flash
Z.AI·Open Weight·1M

77.8

prov. avg

8
Gemini 3.1 Pro
Google·Proprietary·1M

76.8

prov. avg

9
Qwen3.8-27B
Alibaba·Open Weight·262K

76.5

prov. avg

10
Gemini 3.7 Flash
Google·Pending·

73.6

prov. avg

11
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

71.8

prov. avg

12
Muse Spark
Meta·Proprietary·262K

71.4

prov. avg

13
Gemini 3 Pro
Google·Proprietary·2M

67.6

prov. avg

14
Qwen3.7 Plus
Alibaba·Proprietary·1M

65.7

prov. avg

15
GPT-5.5
OpenAI·Proprietary·1M

65.3

prov. avg

16
GPT-5.4
OpenAI·Proprietary·1.05M

63.5

prov. avg

17
GPT-5.2
OpenAI·Proprietary·400K

62.9

prov. avg

18
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

60.8

prov. avg

19
Kimi K2.6
Moonshot AI·Open Weight·256K

60.5

prov. avg

20
Qwen3.6 Plus
Alibaba·Proprietary·1M

59.6

prov. avg

21
Qwen3.5 397B
Alibaba·Open Weight·128K

59.5

prov. avg

22
Gemini 3.5 Flash-Lite
Google·Proprietary·1M

59

prov. avg

23
MiMo-V2.5
Xiaomi·Proprietary·1M

55.7

prov. avg

24
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

46

prov. avg

25
Qwen3.6-27B
Alibaba·Open Weight·262K

44.7

prov. avg

26
MiniMax M3
MiniMax·Open Weight·1M

42.8

prov. avg

27
Inkling
Thinking Machines Lab·Open Weight·1M

42.6

prov. avg

28
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

42.4

prov. avg

29
Inkling-Small
Thinking Machines Lab·Open Weight·1M

41.6

prov. avg

30
Muse Glimmer 30B
Meta·Open Weight·131K

38.6

prov. avg

31
Grok 4.20
xAI·Proprietary·2M

30.5

prov. avg

32
Claude Opus 4.5
Anthropic·Proprietary·200K

13.5

prov. avg

33
Command A+
Cohere·Open Weight·128K

8

prov. avg

Key Takeaways

The top model is Qwen3.8 Max by Alibaba with a provisional score of 87.1.

The best open-weight model is Qwen3.8 Max at position #1.

33 models are included in this ranking.

Score in Context

What these scores mean

The multimodal & grounded score blends image understanding, document/chart reading, and visually grounded reasoning benchmarks. Higher means the model more reliably extracts and reasons over visual content.

Known limitations

These are understanding scores — they say nothing about image generation quality. OCR-heavy production workloads should also test resolution limits and page-count behavior, which benchmarks only partially cover.

Best Multimodal LLMs FAQ

What is the best multimodal LLM?

Gemini 3 Pro Deep Think currently leads BenchLM's multimodal & grounded category at 95, ahead of Claude Mythos 5 (92.8) and GPT-5.1 (91.8). The table above recomputes with every data refresh. For most teams GPT-5.1 is the value pick — top-3 understanding at $1.25/$10 per million tokens.

What is the best LLM for image generation?

This page ranks image understanding, not generation — LLM leaderboards measure how well models read and reason over images. Image generation is a separate product class (diffusion and autoregressive image models) that BenchLM's categories do not currently score. Treat any "LLM" ranking of image generation with skepticism.

Which LLM is best for reading documents and PDFs?

Document understanding tracks the multimodal & grounded category closely: Gemini 3 Pro Deep Think leads, and long-context support matters as much as vision quality for multi-hundred-page work — see the large-context rankings for models pairing 1M-token windows with strong vision.

Can these models understand charts and tables?

The top rows handle standard charts and tables well — chart-reading benchmarks feed this category. Reliability drops on dense dashboards, small text, and unusual chart types, so test your hardest real documents rather than trusting a single aggregate score.

Last updated: August 27, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.