# Vals MATH 500 (MATH 500)

> Academic math benchmark on probability, algebra, and trigonometry

Canonical page: https://benchlm.ai/benchmarks/valsmath500

- Category: [Mathematics](/math)
- Last updated: January 9, 2026

## About MATH 500

- Year: 2026
- Tasks: MATH 500 academic math problems
- Format: Accuracy score
- Difficulty: Advanced academic math
- Paper: [MATH 500](https://www.vals.ai/benchmarks/math500)

BenchLM mirrors the public Vals AI MATH 500 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

MATH 500 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (60 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | Google | 96.40% |
| 2 | [Grok 4 0709](https://www.vals.ai/models/grok_grok-4-0709) | xAI | 96.20% |
| 3 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | OpenAI | 96.00% |
| 4 | [Claude Opus 4.1 20250805 Thinking](https://www.vals.ai/models/anthropic_claude-opus-4-1-20250805-thinking) | Anthropic | 95.40% |
| 5 | [Gemini 2.5 Pro Exp 03 25](https://www.vals.ai/models/google_gemini-2.5-pro-exp-03-25) | Google | 95.20% |
| 6 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 94.80% |
| 7 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 94.80% |
| 8 | [o3](/models/o3) | OpenAI | 94.60% |
| 9 | [Qwen3 235b A22b](https://www.vals.ai/models/fireworks_qwen3-235b-a22b) | Fireworks | 94.60% |
| 10 | [O4 Mini](https://www.vals.ai/models/openai_o4-mini-2025-04-16) | OpenAI | 94.20% |
| 11 | [Grok 3 Mini Fast High Reasoning](https://www.vals.ai/models/grok_grok-3-mini-fast-high-reasoning) | xAI | 94.20% |
| 12 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 94.20% |
| 13 | [Moonshotai Kimi K2 Instruct](https://www.vals.ai/models/together_moonshotai/Kimi-K2-Instruct) | Together AI | 94.20% |
| 14 | [GLM-4.5](/models/glm-4-5) | Z.AI | 94.00% |
| 15 | [GPT-5 nano](/models/gpt-5-nano) | OpenAI | 93.80% |
| 16 | [Claude Sonnet 4 20250514 Thinking](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514-thinking) | Anthropic | 93.80% |
| 17 | [Claude Opus 4.1](https://www.vals.ai/models/anthropic_claude-opus-4-1-20250805) | Anthropic | 93.00% |
| 18 | [DeepSeek-R1](/models/deepseek-r1) | DeepSeek | 92.20% |
| 19 | [o3-mini](/models/o3-mini) | OpenAI | 91.80% |
| 20 | [Gemini 2.5 Flash Preview 04 17 Thinking](https://www.vals.ai/models/google_gemini-2.5-flash-preview-04-17-thinking) | Google | 91.80% |
| 21 | [Gemini 2.5 Flash Preview 04 17](https://www.vals.ai/models/google_gemini-2.5-flash-preview-04-17) | Google | 91.60% |
| 22 | [Claude 3.7 Sonnet 20250219 Thinking](https://www.vals.ai/models/anthropic_claude-3-7-sonnet-20250219-thinking) | Anthropic | 91.60% |
| 23 | [Langston Nim Nvidia Llama 3.3 Nemotron Super 49b V1 42e84561 Thinking](https://www.vals.ai/models/together_langston/nim/nvidia/llama-3.3-nemotron-super-49b-v1-42e84561-thinking) | Together AI | 91.40% |
| 24 | [Claude Opus 4](https://www.vals.ai/models/anthropic_claude-opus-4-20250514) | Anthropic | 90.40% |
| 25 | [o1](/models/o1) | OpenAI | 90.40% |
| 26 | [Claude Sonnet 4](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514) | Anthropic | 90.32% |
| 27 | [Grok 3](https://www.vals.ai/models/grok_grok-3) | xAI | 89.80% |
| 28 | [Gemini 2.0 Flash Exp](https://www.vals.ai/models/google_gemini-2.0-flash-exp) | Google | 89.00% |
| 29 | [MiniMax M2.1](https://www.vals.ai/models/minimax_MiniMax-M2.1) | MiniMax | 89.00% |
| 30 | [DeepSeek V3 0324](https://www.vals.ai/models/fireworks_deepseek-v3-0324) | Fireworks | 88.60% |
| 31 | [Gemini 2.0 Flash 001](https://www.vals.ai/models/google_gemini-2.0-flash-001) | Google | 88.00% |
| 32 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 88.00% |
| 33 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 87.20% |
| 34 | [Mistral Medium 2505](https://www.vals.ai/models/mistralai_mistral-medium-2505) | Mistral | 87.00% |
| 35 | [Llama4 Maverick Instruct Basic](https://www.vals.ai/models/fireworks_llama4-maverick-instruct-basic) | Fireworks | 85.20% |
| 36 | [Gemini 2.0 Flash Thinking Exp 01 21](https://www.vals.ai/models/google_gemini-2.0-flash-thinking-exp-01-21) | Google | 84.60% |
| 37 | [Gemini 1.5 Pro 002](https://www.vals.ai/models/google_gemini-1.5-pro-002) | Google | 82.80% |
| 38 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 80.40% |
| 39 | [GPT-4.1 nano](/models/gpt-4-1-nano) | OpenAI | 80.20% |
| 40 | [Meta Llama Llama 4 Scout 17B 16E Instruct](https://www.vals.ai/models/together_meta-llama/Llama-4-Scout-17B-16E-Instruct) | Together AI | 79.20% |
| 41 | [Gemini 1.5 Flash 002](https://www.vals.ai/models/google_gemini-1.5-flash-002) | Google | 78.80% |
| 42 | [Grok 2 1212](https://www.vals.ai/models/grok_grok-2-1212) | xAI | 78.40% |
| 43 | [Claude 3.7 Sonnet](https://www.vals.ai/models/anthropic_claude-3-7-sonnet-20250219) | Anthropic | 76.80% |
| 44 | [Command A 03 2025](https://www.vals.ai/models/cohere_command-a-03-2025) | Cohere | 76.20% |
| 45 | [GPT-4o](/models/gpt-4o) | OpenAI | 75.20% |
| 46 | [Mistral Large 2411](https://www.vals.ai/models/mistralai_mistral-large-2411) | Mistral | 74.40% |
| 47 | [GPT-4o](/models/gpt-4o) | OpenAI | 74.00% |
| 48 | [Meta Llama Llama 3.3 70B Instruct Turbo](https://www.vals.ai/models/together_meta-llama/Llama-3.3-70B-Instruct-Turbo) | Together AI | 73.40% |
| 49 | [GPT-4o mini](/models/gpt-4o-mini) | OpenAI | 72.60% |
| 50 | [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) | Anthropic | 72.40% |
| 51 | [Meta Llama Meta Llama 3.1 405B Instruct Turbo](https://www.vals.ai/models/together_meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo) | Together AI | 71.40% |
| 52 | [Langston Nim Nvidia Llama 3.3 Nemotron Super 49b V1 42e84561](https://www.vals.ai/models/together_langston/nim/nvidia/llama-3.3-nemotron-super-49b-v1-42e84561) | Together AI | 71.20% |
| 53 | [Mistral Small 2402](https://www.vals.ai/models/mistralai_mistral-small-2402) | Mistral | 70.60% |
| 54 | [Grok 3 Mini Fast Low Reasoning](https://www.vals.ai/models/grok_grok-3-mini-fast-low-reasoning) | xAI | 70.20% |
| 55 | [Mistral Small 2503](https://www.vals.ai/models/mistralai_mistral-small-2503) | Mistral | 68.40% |
| 56 | [Meta Llama Meta Llama 3.1 70B Instruct Turbo](https://www.vals.ai/models/together_meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo) | Together AI | 65.00% |
| 57 | [Claude 3.5 Haiku](https://www.vals.ai/models/anthropic_claude-3-5-haiku-20241022) | Anthropic | 64.20% |
| 58 | [Jamba Large 1.6](https://www.vals.ai/models/ai21labs_jamba-large-1.6) | AI21 Labs | 54.80% |
| 59 | [Meta Llama Meta Llama 3.1 8B Instruct Turbo](https://www.vals.ai/models/together_meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo) | Together AI | 44.40% |
| 60 | [Jamba Mini 1.6](https://www.vals.ai/models/ai21labs_jamba-mini-1.6) | AI21 Labs | 25.40% |

## FAQ

### What does MATH 500 measure?

Academic math benchmark on probability, algebra, and trigonometry

### Which model leads the published MATH 500 snapshot?

Gemini 3 Pro Preview currently leads the published MATH 500 snapshot with a score of 96.40%.

### How many models are evaluated on MATH 500?

The January 9, 2026 contains 60 AI models.

### Does MATH 500 affect BenchLM's overall score?

Not directly. MATH 500 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
