# WeirdML v2 (WeirdML)

> A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.

Canonical page: https://benchlm.ai/benchmarks/weirdml

- Category: [Coding](/coding)
- Last updated: WeirdML v2

## About WeirdML

- Year: 2026
- Tasks: 17 novel ML engineering tasks
- Format: Average accuracy across tasks
- Difficulty: Novel dataset modeling and iterative debugging
- Paper: [WeirdML](https://htihle.github.io/weirdml.html)

WeirdML v2 evaluates models on 17 unusual ML tasks and reports average accuracy across tasks from the official CSV. BenchLM mirrors the top official rows as display-only ML-agent evidence.

WeirdML is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (25 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 84.91% |
| 2 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 83.90% |
| 3 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 82.89% |
| 4 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 77.95% |
| 5 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 77.90% |
| 6 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 77.70% |
| 7 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 76.44% |
| 8 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 76.40% |
| 9 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 76.04% |
| 10 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 75.45% |
| 11 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 72.19% |
| 12 | [gemini-3.1-pro-preview (high)](https://htihle.github.io/weirdml.html) | Google | 72.07% |
| 13 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 70.45% |
| 14 | [gemini-3-pro-preview (high)](https://htihle.github.io/weirdml.html) | Google | 69.93% |
| 15 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 67.15% |
| 16 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 66.07% |
| 17 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 65.87% |
| 18 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 63.74% |
| 19 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 63.44% |
| 20 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 62.64% |
| 21 | [gemini-3-flash-preview (high)](https://htihle.github.io/weirdml.html) | Google | 61.60% |
| 22 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 60.77% |
| 23 | [gpt-5 (high)](https://htihle.github.io/weirdml.html) | OpenAI | 60.70% |
| 24 | [gpt-5-pro (high)](https://htihle.github.io/weirdml.html) | OpenAI | 60.39% |
| 25 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 60.30% |

## FAQ

### What does WeirdML measure?

A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.

### Which model leads the published WeirdML snapshot?

GPT-5.5 currently leads the published WeirdML snapshot with a score of 84.91%.

### How many models are evaluated on WeirdML?

The WeirdML v2 contains 25 AI models.

### Does WeirdML affect BenchLM's overall score?

Not directly. WeirdML is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
