# Apex

> A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

Canonical page: https://benchlm.ai/benchmarks/apex

- Category: [Mathematics](/math)
- Last updated: September 15, 2026

## About Apex

- Year: 2026
- Tasks: Advanced mathematical reasoning
- Format: Pass@1 math benchmark
- Difficulty: Frontier math reasoning
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores Apex as a display-only provider-table reference for exact first-party model comparisons.

Apex is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Hy4 preview](/models/hy4-preview) | Tencent | 74.2% |
| 2 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 65.6% |
| 3 | [A.X K2](/models/a-x-k2) | SK Telecom | 45.8% |
| 4 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 44.5% |
| 5 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 38.3% |
| 6 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 33.0% |
| 7 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 32.2% |
| 8 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 22.7% |

## FAQ

### What does Apex measure?

A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

### Which model scores highest on Apex?

Hy4 preview by Tencent currently leads with a score of 74.2% on Apex.

### How many models are evaluated on Apex?

8 AI models have been evaluated on Apex on BenchLM.

### Does Apex affect BenchLM's overall score?

Not directly. Apex is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on Apex

- [Hy4 preview vs DeepSeek V4.1 Flash](/compare/deepseek-v4-1-flash-vs-hy4-preview)
- [DeepSeek V4.1 Flash vs A.X K2](/compare/a-x-k2-vs-deepseek-v4-1-flash)
- [A.X K2 vs Qwen3.7 Max](/compare/a-x-k2-vs-qwen3-7-max)
- [Qwen3.7 Max vs DeepSeek V4 Pro 0813](/compare/deepseek-v4-pro-0813-vs-qwen3-7-max)
