Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Multimodal & Grounded benchmark report

Best LLMs for Multimodal & GroundedSeptember 2026 Leaderboard

Data refreshed:

Vision, document, and grounded enterprise workflow benchmarks

As of September 2026, the top multimodal & grounded model on the BenchLM leaderboard is Kimi K3 with a weighted multimodal & grounded score of 89.5.

Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.

Data refreshed
September 8, 2026
Provisional-ranked
48 of 418 models
Verified-ranked
45 of 418 models
Weighted evidence
2 of 24 benchmarks
24 tracked benchmarks

MMMU-Pro, OCRBench V2, olmOCR, VoxPopuli WER, OfficeQA Pro, MMMU-Pro w/ Python, OmniDocBench 1.5, Liquid Extract JSON Validity, Liquid Extract F1, Liquid Extract VLM Judge, GDPval-AA, Blueprint-Bench 2, MedXpertQA (MM), ZeroBench, Design2Code, Flame-VLM-Code, Vision2Web, ImageMining, MMSearch, MMSearch-Plus, SimpleVQA, Facts-VLM, V*, BabyVision

Scope: Vision, Document-office, GUI/web, Video

Best Multimodal & Grounded picks

BenchLM summaries for multimodal & grounded plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Multimodal & Grounded Leaderboard

Primary score: weighted multimodal & grounded score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.

Filters
Provisional-ranked mode includes source-unverified non-generated benchmark evidence.P = provisional benchmark row
Rank / modelWeighted Multimodal & Grounded
1
Kimi K3Moonshot AI · Closed
89.5%
2
Claude Opus 5Anthropic · Closed
88.7%
3
Claude Opus 4.8Anthropic · Closed
87.8%
4
GPT-5.6 SolOpenAI · Closed
87.5%
5
Qwen3.8 MaxAlibaba · Open weight
87.4%
6
Gemini 3.5 FlashGoogle · Closed
86.9%
7
Qwen3.8-Flash-NextAlibaba · Open weight
83.3%
8
Gemini 3.8 FlashGoogle · Closed
82.8%
9
Gemini 3.7 FlashGoogle · Closed
82.7%
10
GLM-5.3-FlashZ.AI · Open weight
80.5%
11
Qwen3.8-27BAlibaba · Open weight
80%
12
Gemini 3.1 ProGoogle · Closed
79.2%
13
Claude Sonnet 5Anthropic · Closed
77.5%
14
Muse SparkMeta · Closed
77.5%
15
GPT-5.6 TerraOpenAI · Closed
77.1%
16
Gemini 3.5 Flash-LiteGoogle · Closed
76.3%
17
Gemini 3 ProGoogle · Closed
74.2%
18
Qwen3.7 PlusAlibaba · Closed
72.5%
19
GPT-5.5OpenAI · Closed
71.3%
20
GPT-5.4OpenAI · Closed
69.3%
21
GPT-5.6 LunaOpenAI · Closed
67.1%
22
Qwen3.6 PlusAlibaba · Closed
66.3%
23
GPT-5.2OpenAI · Closed
66.3%
24
Kimi K2.5Moonshot AI · Open weight
65.7%
25
Grok 4.3xAI · Closed
65.6%

Top AI Models for Multimodal & GroundedSeptember 2026

As of September 2026, Kimi K3 leads the provisional multimodal & grounded leaderboard with a score of 89.5%, followed by Claude Opus 5 (88.7%) and Claude Opus 4.8 (87.8%). BenchLM is currently showing 48 provisional-ranked models and 45 verified-ranked models in this category.

What changed

Kimi K3 leads the current multimodal table across the available grounded evidence.

GPT-5.4 close behind with strong OfficeQA Pro and MMMU-Pro results.

Claude Opus 4.7 adds official CharXiv visual reasoning coverage.

Top models by benchmark

Frontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems(40% of category score)

RankModelReported score
5Kimi K381.6

Score in Context

What these scores mean

Multimodal & Grounded carries a 12% weight in overall scoring. The weighted score blends MMMU-Pro (academic multimodal reasoning), OfficeQA Pro (enterprise document understanding), and CharXiv (visual chart and figure reasoning). A model can know facts in text and still fail when the information is in a chart, screenshot, or spreadsheet — this category measures that gap.

Known limitations

Not all models support image input — text-only models are excluded from this category entirely. OfficeQA Pro and CharXiv coverage is still building, so rankings should be read as a blend of available public evidence rather than a complete visual capability profile. Enterprise-specific document formats (scanned PDFs, handwritten notes) remain under-tested by all benchmarks.

How we weight

Multimodal & Grounded carries a 12% weight in BenchLM.ai's overall scoring. It remains important for enterprise copilots and document-heavy workflows where models need to interpret visuals, screenshots, and scanned artifacts.

This category tests whether a model can read the world as it actually appears in products: screenshots, charts, scanned documents, and mixed visual-text artifacts. See the multimodal leaderboard or compare with knowledge benchmarks.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Multimodal & Grounded benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
MMMU-Pro40%WeightedFrontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems
OCRBench V2Display onlyA native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.
olmOCRDisplay onlyAn end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.
VoxPopuli WERDisplay onlyA speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.
OfficeQA Pro25%WeightedGrounded office and enterprise document benchmark
MMMU-Pro w/ PythonDisplay onlyTool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning
OmniDocBench 1.5Display onlyDocument understanding benchmark measured by edit distance on complex document extraction tasks
Liquid Extract JSON ValidityDisplay onlyA display-only Liquid AI extraction metric measuring strict JSON parseability.
Liquid Extract F1Display onlyA display-only Liquid AI extraction metric measuring requested-field agreement.
Liquid Extract VLM JudgeDisplay onlyA display-only Liquid AI extraction metric measuring judged agreement with the source image.
GDPval-AADisplay onlyAn evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.
Blueprint-Bench 2Display onlyAgentic spatial reasoning benchmark reported as a normalized score.
MedXpertQA (MM)Display onlyA clinically grounded multimodal medical multiple-choice benchmark with image inputs.
ZeroBenchDisplay onlyA multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.
Design2CodeDisplay onlyMultimodal coding benchmark for turning visual designs into working frontend implementations.
Flame-VLM-CodeDisplay onlyVision-language coding benchmark for generating correct code from visual and multimodal inputs.
Vision2WebDisplay onlyBenchmark for converting visual references into functional web implementations.
ImageMiningDisplay onlyMultimodal retrieval and extraction benchmark over image-heavy task settings.
MMSearchDisplay onlyMultimodal search benchmark for retrieval and grounded answering across mixed-media inputs.
MMSearch-PlusDisplay onlyA harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.
SimpleVQADisplay onlyVisual question answering benchmark focused on straightforward image-grounded understanding.
Facts-VLMDisplay onlyGrounded multimodal factuality benchmark for evidence-linked answer correctness.
V*Display onlyVision-centric benchmark for high-level multimodal reasoning and perception quality.
BabyVisionDisplay onlyA multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

About Multimodal & Grounded Benchmarks

Frontier multimodal reasoning benchmark spanning charts, diagrams, tables, and visual academic problems

Common questions

What do multimodal and grounded LLM benchmarks measure?

They measure whether models can reason over images, charts, documents, and office artifacts instead of only plain-text prompts.

Which benchmarks matter most here?

MMMU-Pro is a strong frontier multimodal reasoning benchmark, OfficeQA Pro is useful for grounded document and office-style enterprise workflows, and CharXiv tests chart and figure reasoning.

Why is this category separate from knowledge?

A model can know facts in text and still be weak when the information is embedded in a chart, screenshot, PDF, or spreadsheet. This category measures that gap directly.

Multimodal benchmark updates

Multimodal rankings are heating up. Get the weekly update.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Related