# Vals IOI v1 archived leaderboard (IOI v1 (Vals, archived))

> Archived Vals AI results on 2024 and 2025 International Olympiad in Informatics programming tasks.

Canonical page: https://benchlm.ai/benchmarks/vals-ioi-v1

- Category: [Coding](/coding)
- Last updated: August 9, 2026

## About IOI v1 (Vals, archived)

- Year: 2026
- Tasks: IOI 2024 and 2025 programming tasks
- Format: Accuracy score
- Difficulty: Olympiad programming
- Paper: [Vals IOI v1](https://www.vals.ai/benchmarks/ioi-v1)

Vals stopped running IOI v1 on new model releases and preserves its prior results. This display-only version-1 table stays separate from the current IOI v2 task set, which also includes IOI 2026.

IOI v1 (Vals, archived) is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (62 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 91.67% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 86.67% |
| 3 | [Qwen3.8 Max](/models/qwen3-8-max) | — | Alibaba | 73.00% |
| 4 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 72.92% |
| 5 | [Claude Fable 5](/models/claude-fable) | — | Anthropic | 72.25% |
| 6 | [GPT-5.4](/models/gpt-5-4) | xhigh reasoning | OpenAI | 67.83% |
| 7 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | xhigh reasoning | OpenAI | 65.25% |
| 8 | [GPT-5.2](/models/gpt-5-2) | xhigh reasoning | OpenAI | 54.83% |
| 9 | [Muse Spark 1.2](/models/muse-spark-1-2) | — | Meta | 49.50% |
| 10 | [Claude Opus 4.7](/models/claude-opus-4-7) | — | Anthropic | 47.08% |
| 11 | [Qwen3.7 Max](/models/qwen3-7-max) | — | Alibaba | 46.75% |
| 12 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | xhigh reasoning | OpenAI | 43.83% |
| 13 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | high reasoning | Google | 39.08% |
| 14 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | high reasoning | Google | 38.83% |
| 15 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 35.83% |
| 16 | [Grok 4.20 0309 Reasoning](https://www.vals.ai/models/grok_grok-4.20-0309-reasoning) | — | xAI | 30.17% |
| 17 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | high reasoning | Google | 26.17% |
| 18 | [Grok 4 0709](https://www.vals.ai/models/grok_grok-4-0709) | — | xAI | 26.17% |
| 19 | [Claude Opus 4.5](/models/claude-opus-4-5) | — | Anthropic | 23.58% |
| 20 | [GLM 5 Thinking](https://www.vals.ai/models/zai_glm-5-thinking) | — | Zhipu AI | 22.00% |
| 21 | [GPT-5.1](/models/gpt-5-1) | high reasoning | OpenAI | 21.50% |
| 22 | [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) | high reasoning | OpenAI | 21.42% |
| 23 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | — | Anthropic | 20.25% |
| 24 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | high reasoning | OpenAI | 20.00% |
| 25 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | — | Anthropic | 18.33% |
| 26 | [Kimi K2.5 Thinking](https://www.vals.ai/models/kimi_kimi-k2.5-thinking) | — | Moonshot AI | 17.67% |
| 27 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | — | Google | 17.08% |
| 28 | [Qwen3 Max](/models/qwen3-max) | — | Alibaba | 15.67% |
| 29 | [Grok 4.3](/models/grok-4-3) | high reasoning | xAI | 15.33% |
| 30 | [GPT-5.4 nano](/models/gpt-5-4-nano) | high reasoning | OpenAI | 15.25% |
| 31 | [DeepSeek V3p2](https://www.vals.ai/models/fireworks_deepseek-v3p2) | none reasoning | Fireworks AI | 14.42% |
| 32 | [Qwen3 Max](/models/qwen3-max) | — | Alibaba | 13.75% |
| 33 | [Claude Opus 4.1](https://www.vals.ai/models/anthropic_claude-opus-4-1-20250805) | — | Anthropic | 12.52% |
| 34 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | — | xAI | 11.50% |
| 35 | [DeepSeek V3p2 Thinking](https://www.vals.ai/models/fireworks_deepseek-v3p2-thinking) | high reasoning | Fireworks AI | 10.67% |
| 36 | [GPT-5 Codex](https://www.vals.ai/models/openai_gpt-5-codex) | high reasoning | OpenAI | 9.75% |
| 37 | [Qwen3 Max Preview](https://www.vals.ai/models/alibaba_qwen3-max-preview) | — | Alibaba | 7.75% |
| 38 | [Grok 4.1 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-1-fast-non-reasoning) | — | xAI | 7.67% |
| 39 | [GLM-4.7](/models/glm-4-7) | — | Z.AI | 7.58% |
| 40 | [GPT-5 mini](/models/gpt-5-mini) | high reasoning | OpenAI | 6.75% |
| 41 | [MiniMax M2.5](/models/minimax-m2-5) | — | MiniMax | 6.67% |
| 42 | [Claude Sonnet 4](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514) | — | Anthropic | 6.50% |
| 43 | [GPT-5.4 mini](/models/gpt-5-4-mini) | xhigh reasoning | OpenAI | 6.42% |
| 44 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | — | Anthropic | 6.17% |
| 45 | [MiniMax M2.7](/models/minimax-m2-7) | — | MiniMax | 4.92% |
| 46 | [O4 Mini](https://www.vals.ai/models/openai_o4-mini-2025-04-16) | high reasoning | OpenAI | 4.83% |
| 47 | [Claude Sonnet 4 20250514 Thinking](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514-thinking) | — | Anthropic | 4.58% |
| 48 | [GLM-4.6](/models/glm-4-6) | — | Z.AI | 4.33% |
| 49 | [Grok Code Fast 1](/models/grok-code-fast-1) | — | xAI | 4.33% |
| 50 | [Mistral Large 2512](https://www.vals.ai/models/mistralai_mistral-large-2512) | — | Mistral AI | 4.00% |
| 51 | [Grok 4 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-fast-non-reasoning) | — | xAI | 3.83% |
| 52 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | high reasoning | OpenAI | 3.67% |
| 53 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | — | xAI | 3.08% |
| 54 | [GLM-4.5](/models/glm-4-5) | — | Z.AI | 2.92% |
| 55 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | — | Google | 2.61% |
| 56 | [Labs Devstral Small 2512](https://www.vals.ai/models/mistralai_labs-devstral-small-2512) | — | Mistral AI | 2.50% |
| 57 | [MiniMax M2.1](https://www.vals.ai/models/minimax_MiniMax-M2.1) | — | MiniMax | 2.33% |
| 58 | [DeepSeek V3 0324](https://www.vals.ai/models/fireworks_deepseek-v3-0324) | — | Fireworks AI | 1.67% |
| 59 | [Moonshotai Kimi K2 Instruct](https://www.vals.ai/models/together_moonshotai/Kimi-K2-Instruct) | — | Together AI | 1.25% |
| 60 | [Magistral Medium 2509](https://www.vals.ai/models/mistralai_magistral-medium-2509) | — | Mistral AI | 0.67% |
| 61 | [Devstral 2512](https://www.vals.ai/models/mistralai_devstral-2512) | — | Mistral AI | 0.67% |
| 62 | [Qwen3 235b A22b](https://www.vals.ai/models/fireworks_qwen3-235b-a22b) | — | Fireworks AI | 0.00% |

## FAQ

### What does IOI v1 (Vals, archived) measure?

Archived Vals AI results on 2024 and 2025 International Olympiad in Informatics programming tasks.

### Which model leads the published IOI v1 (Vals, archived) snapshot?

Claude Opus 5 currently leads the published IOI v1 (Vals, archived) snapshot with a score of 91.67%.

### How many models are evaluated on IOI v1 (Vals, archived)?

The August 9, 2026 contains 62 AI models.

### Does IOI v1 (Vals, archived) affect BenchLM's overall score?

Not directly. IOI v1 (Vals, archived) is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
