# Best LLMs for Multimodal & Grounded — September 2026 Leaderboard

> As of September 2026, Kimi K3 leads BenchLM's multimodal & grounded leaderboard with a weighted score of 89.5.

- **Last verified:** September 8, 2026
- Canonical page: https://benchlm.ai/multimodal-grounded
- **Ranking coverage:** 48 category-ranked models from 412 tracked models
- **Category weight:** 12% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 89.5 | 15 | 52 total |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 88.7 | 10 | 84 total |
| 3 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 87.8 | 5 | 44 total |
| 4 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 87.5 | 3 | 47 total |
| 5 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 87.4 | 24 | 55 total |
| 6 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 86.9 | 5 | 38 total |
| 7 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 83.3 | 9 | 29 total |
| 8 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 82.8 | 4 | 28 total |
| 9 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 82.7 | 5 | 30 total |
| 10 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 80.5 | 5 | 25 total |
| 11 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 80 | 11 | 42 total |
| 12 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 79.2 | 9 | 25 total |
| 13 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 77.5 | 4 | 32 total |
| 14 | [Muse Spark](/models/muse-spark) | Meta | 77.5 | 8 | 30 total |
| 15 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 77.1 | 3 | 44 total |
| 16 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 76.3 | 1 | 19 total |
| 17 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 74.2 | 7 | 18 total |
| 18 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 72.5 | 17 | 61 total |
| 19 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 71.3 | 5 | 44 total |
| 20 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 69.3 | 11 | 40 total |
| 21 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 67.1 | 3 | 39 total |
| 22 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 66.3 | 9 | 52 total |
| 23 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 66.3 | 5 | 20 total |
| 24 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 65.7 | 6 | 48 total |
| 25 | [Grok 4.3](/models/grok-4-3) | xAI | 65.6 | 3 | 20 total |
| 26 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 64 | 7 | 34 total |
| 27 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 63 | 7 | 35 total |
| 28 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 59.5 | 6 | 37 total |
| 29 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 59.2 | 4 | 13 total |
| 30 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 58.4 | 2 | 15 total |
| 31 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 57.2 | 3 | 26 total |
| 32 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 57 | 6 | 22 total |
| 33 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 54.1 | 3 | 30 total |
| 34 | [MiniMax M3](/models/minimax-m3) | MiniMax | 52.2 | 7 | 32 total |
| 35 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 51.9 | 16 | 46 total |
| 36 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 50.3 | 4 | 20 total |
| 37 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 50.1 | 15 | 49 total |
| 38 | [Inkling](/models/inkling) | Thinking Machines Lab | 49.2 | 5 | 31 total |
| 39 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 48.8 | 4 | 32 total |
| 40 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 46.4 | 5 | 23 total |
| 41 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 44.1 | 2 | 12 total |
| 42 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 39.3 | 8 | 22 total |
| 43 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 34.6 | 6 | 22 total |
| 44 | [LFM2.5-VL-1.6B-Extract](/models/lfm2-5-vl-1-6b-extract) | LiquidAI | 26.5 | 4 | 3 total |
| 45 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 26.1 | 4 | 19 total |
| 46 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 23.8 | 3 | 25 total |
| 47 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 23.5 | 8 | 49 total |
| 48 | [Command A+](/models/command-a-plus) | Cohere | 15.1 | 4 | 12 total |

## Decision-ready shortlist

- #1 [Kimi K3](/models/kimi-k3) — 89.5 weighted score, Pending, 1.05M context.
- #2 [Claude Opus 5](/models/claude-opus-5) — 88.7 weighted score, Proprietary, null context.
- #3 [Claude Opus 4.8](/models/claude-opus-4-8) — 87.8 weighted score, Proprietary, 1M context.
- #4 [GPT-5.6 Sol](/models/gpt-5-6-sol) — 87.5 weighted score, Proprietary, 1.05M context.
- #5 [Qwen3.8 Max](/models/qwen3-8-max) — 87.4 weighted score, Open Weight, 1M context.

## Benchmarks in this category

### [MMMU-Pro](/benchmarks/mmmu-pro) (Massive Multi-discipline Multimodal Understanding Pro)

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

- Ranking status: Weighted (40% of this category)
- Year: 2024
- Format: Image + text question answering
- Difficulty: Frontier multimodal

### [OCRBench V2](/benchmarks/ocrbenchv2) (OCRBench V2)

A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.

- Ranking status: Display only
- Year: 2025
- Format: Accuracy
- Difficulty: Native visual text understanding

### [olmOCR](/benchmarks/olmocr) (olmOCR-Bench)

An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.

- Ranking status: Display only
- Year: 2025
- Format: Mean accuracy
- Difficulty: Complex document processing

### [VoxPopuli WER](/benchmarks/voxpopuliwer) (VoxPopuli-Cleaned-AA Word Error Rate)

A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.

- Ranking status: Display only
- Year: 2026
- Format: Word error rate
- Difficulty: Audio speech recognition

### [OfficeQA Pro](/benchmarks/officeqapro) (OfficeQA Pro)

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

- Ranking status: Weighted (25% of this category)
- Year: 2026
- Format: Grounded QA over office artifacts
- Difficulty: Enterprise grounded reasoning

### [MMMU-Pro w/ Python](/benchmarks/mmmupropython) (MMMU-Pro with Python)

Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning.

- Ranking status: Display only
- Year: 2026
- Format: Image + text question answering with Python
- Difficulty: Frontier multimodal

### [OmniDocBench 1.5](/benchmarks/omnidocbench15) (OmniDocBench 1.5)

A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.

- Ranking status: Display only
- Year: 2026
- Format: Document understanding benchmark
- Difficulty: Grounded document reasoning

### [Liquid Extract JSON Validity](/benchmarks/liquidextractjsonvalidity) (Liquid image-to-JSON extraction JSON validity)

A display-only Liquid AI extraction metric measuring the share of image-to-JSON outputs that parse as strict JSON.

- Ranking status: Display only
- Year: 2026
- Format: Strict JSON parseability rate
- Difficulty: Structured visual extraction

### [Liquid Extract F1](/benchmarks/liquidextractschemaf1) (Liquid image-to-JSON extraction schema consistency F1)

A display-only Liquid AI extraction metric measuring field-name agreement between requested schema fields and extracted JSON fields.

- Ranking status: Display only
- Year: 2026
- Format: Schema field F1
- Difficulty: Structured visual extraction

### [Liquid Extract VLM Judge](/benchmarks/liquidextractvlmjudge) (Liquid image-to-JSON extraction VLM judge score)

A display-only Liquid AI extraction metric measuring judged agreement between extracted values and the source image.

- Ranking status: Display only
- Year: 2026
- Format: VLM-judged extraction accuracy
- Difficulty: Structured visual extraction

### [GDPval-AA](/benchmarks/gdpvalaa) (GDPval-AA)

An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.

- Ranking status: Display only
- Year: 2026
- Format: ELO-style office benchmark
- Difficulty: Professional knowledge work

### [Blueprint-Bench 2](/benchmarks/blueprintbench2) (Blueprint-Bench 2)

An agentic spatial reasoning benchmark reported as a normalized score.

- Ranking status: Display only
- Year: 2026
- Format: Normalized score
- Difficulty: Agentic spatial reasoning

### [MedXpertQA (MM)](/benchmarks/medxpertqamm) (MedXpertQA Multimodal)

A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

- Ranking status: Display only
- Year: 2026
- Format: Medical visual MCQ
- Difficulty: Clinical multimodal reasoning

### [ZeroBench](/benchmarks/zerobench) (ZeroBench)

A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.

- Ranking status: Display only
- Year: 2026
- Format: Multi-step visual reasoning
- Difficulty: Tool-augmented visual reasoning

### [Design2Code](/benchmarks/design2code) (Design2Code)

A multimodal coding benchmark for turning visual designs into working frontend implementations.

- Ranking status: Display only
- Year: 2026
- Format: Visual input to frontend implementation
- Difficulty: Multimodal coding

### [Flame-VLM-Code](/benchmarks/flamevlmcode) (Flame-VLM-Code)

A vision-language coding benchmark for generating correct code from visual and multimodal inputs.

- Ranking status: Display only
- Year: 2026
- Format: Vision-language code generation
- Difficulty: Multimodal coding

### [Vision2Web](/benchmarks/vision2web) (Vision2Web)

A benchmark for converting visual references into functional web implementations.

- Ranking status: Display only
- Year: 2026
- Format: Visual reference to web implementation
- Difficulty: Multimodal web generation

### [ImageMining](/benchmarks/imagemining) (ImageMining)

A multimodal retrieval and extraction benchmark over image-heavy task settings.

- Ranking status: Display only
- Year: 2026
- Format: Image-grounded retrieval and extraction
- Difficulty: Multimodal retrieval

### [MMSearch](/benchmarks/mmsearch) (MMSearch)

A multimodal search benchmark for retrieval and grounded answering across mixed-media inputs.

- Ranking status: Display only
- Year: 2026
- Format: Mixed-media retrieval and grounded answering
- Difficulty: Multimodal search

### [MMSearch-Plus](/benchmarks/mmsearchplus) (MMSearch-Plus)

A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.

- Ranking status: Display only
- Year: 2026
- Format: Advanced mixed-media retrieval benchmark
- Difficulty: Advanced multimodal search

### [SimpleVQA](/benchmarks/simplevqa) (SimpleVQA)

A visual question answering benchmark focused on straightforward image-grounded understanding.

- Ranking status: Display only
- Year: 2026
- Format: Image-grounded question answering
- Difficulty: General visual understanding

### [Facts-VLM](/benchmarks/factsvlm) (Facts-VLM)

A grounded multimodal factuality benchmark for evidence-linked answer correctness.

- Ranking status: Display only
- Year: 2026
- Format: Evidence-linked multimodal factuality
- Difficulty: Grounded multimodal factuality

### [V*](/benchmarks/vstar) (V*)

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

- Ranking status: Display only
- Year: 2026
- Format: Vision-centric reasoning benchmark
- Difficulty: Frontier multimodal

### [BabyVision](/benchmarks/babyvision) (BabyVision)

A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

- Ranking status: Display only
- Year: 2026
- Format: Multimodal visual reasoning
- Difficulty: Fine-grained visual perception
