# IMOAnswerBench

> A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

Canonical page: https://benchlm.ai/benchmarks/imoanswerbench

- Category: [Mathematics](/math)
- Last updated: September 15, 2026

## About IMOAnswerBench

- Year: 2026
- Tasks: Advanced mathematical answer generation
- Format: Pass@1 math benchmark
- Difficulty: Olympiad-level mathematics
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores IMOAnswerBench as a display-only provider-table row when exact values are published for frontier math comparisons.

IMOAnswerBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (9 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 90.9% |
| 2 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 90.0% |
| 3 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 89.8% |
| 4 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 88.4% |
| 5 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 86.0% |
| 6 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 83.7% |
| 7 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 78.6% |
| 8 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 59.3% |
| 9 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 49.4% |

## FAQ

### What does IMOAnswerBench measure?

A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

### Which model scores highest on IMOAnswerBench?

dots3-note Preview by Dots Studio currently leads with a score of 90.9% on IMOAnswerBench.

### How many models are evaluated on IMOAnswerBench?

9 AI models have been evaluated on IMOAnswerBench on BenchLM.

### Does IMOAnswerBench affect BenchLM's overall score?

Not directly. IMOAnswerBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on IMOAnswerBench

- [dots3-note Preview vs Qwen3.7 Max](/compare/dots3-note-preview-vs-qwen3-7-max)
- [Qwen3.7 Max vs DeepSeek V4 Pro 0813](/compare/deepseek-v4-pro-0813-vs-qwen3-7-max)
- [DeepSeek V4 Pro 0813 vs DeepSeek V4 Flash 0731](/compare/deepseek-v4-flash-0731-vs-deepseek-v4-pro-0813)
- [DeepSeek V4 Flash 0731 vs Qwen3.7 Plus](/compare/deepseek-v4-flash-0731-vs-qwen3-7-plus)
