# SWE-bench Multimodal

> A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.

Canonical page: https://benchlm.ai/benchmarks/swe-bench-multimodal

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 18, 2026

## About SWE-bench Multimodal

- Year: 2025
- Tasks: Multimodal software engineering tasks
- Format: Code patch generation with visual context
- Difficulty: Frontier multimodal coding
- Paper: [SWE-bench Multimodal](https://www.swebench.com/multimodal)

SWE-bench Multimodal is important because real-world software engineering increasingly involves visual inputs like UI mockups, error screenshots, and design specifications. Scores tend to be much lower than text-only SWE-bench variants.

SWE-bench Multimodal is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 54.9% |

## FAQ

### What does SWE-bench Multimodal measure?

A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.

### Which model scores highest on SWE-bench Multimodal?

Claude Mythos 5 by Anthropic currently leads with a score of 54.9% on SWE-bench Multimodal.

### How many models are evaluated on SWE-bench Multimodal?

1 AI models have been evaluated on SWE-bench Multimodal on BenchLM.

### Does SWE-bench Multimodal affect BenchLM's overall score?

Not directly. SWE-bench Multimodal is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
