Skip to main content
BenchLM

Vals IOI (IOI)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Based on the International Olympiad in Informatics

IOI score on IOI — September 27, 2026

We mirror the published ioi score view for IOI. GPT-6 Astra leads the public snapshot at 100.00%, followed by Claude Opus 5.5 (95.06%) and GPT-5.6 Sol (91.17%). We do not use these results to rank models overall.

36 modelsCodingCurrentDisplay onlyUpdated September 27, 2026

IOI score table (36 models)

Score
1
GPT-6 AstraOpenAI · Closedmax reasoning
100.00%
2
Claude Opus 5.5Anthropic · Closed
95.06%
3
GPT-5.6 SolOpenAI · Closedmax reasoning
91.17%
4
Claude Fable 5.1Anthropic · Closed
90.78%
5
GPT-5.6 TerraOpenAI · Closedmax reasoning
87.61%
6
Claude Opus 5Anthropic · Closed
84.33%
7
Claude Sonnet 5.5Anthropic · Closed
83.06%
8
GPT-6 SolOpenAI · Closedmax reasoning
82.61%
9
Qwen3.8 MaxAlibaba · Open weight
68.89%
10
GLM-5.3Z.AI · Open weightmax reasoning
68.44%
11
Gemini 3.7 FlashGoogle · Closedhigh reasoning
67.83%
12
GPT-5.6 LunaOpenAI · Closedmax reasoning
61.78%
13
Hy4 previewTencent · Open weight
59.33%
14
Grok 4.7xAI · Closedxhigh reasoning
57.72%
15
Gemini 3.8 FlashGoogle · Closedhigh reasoning
56.94%
16
Muse Spark 1.3 MaxMetamax reasoning
56.56%
17
GPT-6 LunaOpenAI · Closedmax reasoning
55.56%
18
GPT-5.3 CodexOpenAI · Closedxhigh reasoning
53.83%
19
GLM-5.3-FlashZ.AI · Open weightmax reasoning
52.50%
20
Gemini 3.1 Pro PreviewGooglehigh reasoning
51.83%
21
DeepSeek V4 Pro 0813DeepSeek · Open weightmax reasoning
51.61%
22
Kimi K3Moonshot AI · Closedmax reasoning
48.94%
23
MiMo-V2.6-FlashXiaomi · Open weight
47.67%
24
Grok 4.6xAI · Closedhigh reasoning
47.61%
25
Claude Sonnet 5Anthropic · Closed
45.00%
26
Muse Spark 1.3Meta · Closedxhigh reasoning
43.94%
27
Grok 4.5xAI · Closedhigh reasoning
40.56%
28
DeepSeek V4.1 FlashDeepSeek · Open weighthigh reasoning
40.28%
29
MiMo-V2.6-ProXiaomi · Open weight
39.33%
30
Qwen3.8-27BAlibaba · Open weightxhigh reasoning
39.06%
31
Gemini 3.6 FlashGoogle · Closedhigh reasoning
35.06%
32
DeepSeek V4 Flash 0731DeepSeek · Open weighthigh reasoning
32.72%
33
Muse Spark 1.2Meta · Closedxhigh reasoning
21.78%
34
InklingThinking Machines Lab · Open weight0.99 reasoning
14.94%
35
Inkling-SmallThinking Machines Lab · Open weight0.99 reasoning
9.33%
36
Mercury 2.5Inception · Closedhigh reasoning
2.44%

How IOI is shown here

BenchLM mirrors the public Vals AI IOI leaderboard captured from https://www.vals.ai/benchmarks/ioi and updated by Vals on September 27, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

IOI is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

36 Vals rows4 task viewspublic datasetTasks: Overall, IOI 2024, IOI 2025, IOI 2026Display only

The published IOI snapshot places GPT-6 Astra first at 100.00%. The third row is 8.83 points behind. The broader top-10 range is 31.56 points, so the table still separates the published systems.

36 models have been evaluated on IOI. The benchmark falls in the Coding category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. IOI is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About IOI

Year

2026

Tasks

International Olympiad in Informatics-style programming tasks

Format

Accuracy score

Difficulty

Olympiad programming

BenchLM mirrors the public Vals AI IOI leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Freshness and provenance

Version

IOI 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does IOI measure?

Based on the International Olympiad in Informatics

Which model leads the published IOI snapshot?

GPT-6 Astra currently leads the published IOI snapshot with 100.00% ioi score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on IOI?

The September 27, 2026 snapshot contains 36 AI models.

Last updated: September 27, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.