# Vals-hosted GPQA Diamond mirror (Vals GPQA Diamond mirror)

> Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.

Canonical page: https://benchlm.ai/benchmarks/valsgpqadiamond

- Category: [external](/external)
- Last updated: September 1, 2026

## About Vals GPQA Diamond mirror

- Year: 2026
- Tasks: GPQA Diamond task splits
- Format: Accuracy score
- Difficulty: Graduate science reasoning
- Paper: [Vals GPQA Diamond](https://www.vals.ai/benchmarks/gpqa)

BenchLM keeps this Vals-hosted GPQA Diamond table separate from canonical GPQA source records.

Vals GPQA Diamond mirror is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (138 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | high reasoning | Google | 95.45% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 95.20% |
| 3 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 94.70% |
| 4 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | high reasoning | Google | 94.44% |
| 5 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | high reasoning | Google | 93.94% |
| 6 | [Qwen3.8 Max](/models/qwen3-8-max) | — | Alibaba | 93.69% |
| 7 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | high reasoning | Google | 93.43% |
| 8 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 93.43% |
| 9 | [Claude Fable 5.1](/models/claude-fable-5-1) | — | Anthropic | 93.43% |
| 10 | [Claude Fable 5](/models/claude-fable) | — | Anthropic | 93.18% |
| 11 | [GPT-5.5](/models/gpt-5-5) | xhigh reasoning | OpenAI | 93.18% |
| 12 | [Kimi K3](/models/kimi-k3) | — | Moonshot AI | 92.93% |
| 13 | [Grok 4.5](/models/grok-4-5) | high reasoning | xAI | 92.93% |
| 14 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | high reasoning | Google | 92.68% |
| 15 | [MiniMax M3](/models/minimax-m3) | — | MiniMax | 92.68% |
| 16 | [Claude Opus 4.8](/models/claude-opus-4-8) | — | Anthropic | 92.42% |
| 17 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 92.42% |
| 18 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 91.67% |
| 19 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | high reasoning | Google | 91.67% |
| 20 | [GPT-5.2](/models/gpt-5-2) | xhigh reasoning | OpenAI | 91.67% |
| 21 | [GPT-5.4](/models/gpt-5-4) | xhigh reasoning | OpenAI | 91.67% |
| 22 | [Grok 4.3](/models/grok-4-3) | — | xAI | 91.41% |
| 23 | [Muse Spark 1.1](/models/muse-spark-1-1) | xhigh reasoning | Meta | 91.16% |
| 24 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | xhigh reasoning | OpenAI | 90.91% |
| 25 | [Qwen3.7 Max](/models/qwen3-7-max) | — | Alibaba | 90.15% |
| 26 | [Claude Opus 4.7](/models/claude-opus-4-7) | — | Anthropic | 90.15% |
| 27 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | high reasoning | DeepSeek | 89.90% |
| 28 | [Muse Spark](/models/muse-spark) | — | Meta | 89.65% |
| 29 | [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) | — | Anthropic | 89.65% |
| 30 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 89.39% |
| 31 | [Kimi K2.6](/models/kimi-2-6) | — | Moonshot AI | 89.14% |
| 32 | [Claude Sonnet 5](/models/claude-sonnet-5) | — | Anthropic | 88.89% |
| 33 | [Qwen3.8-27B](/models/qwen3-8-27b) | xhigh reasoning | Alibaba | 88.89% |
| 34 | [Grok 4.20 0309 Reasoning](https://www.vals.ai/models/grok_grok-4.20-0309-reasoning) | — | xAI | 88.64% |
| 35 | [Grok 4 0709](https://www.vals.ai/models/grok_grok-4-0709) | — | xAI | 88.13% |
| 36 | [GLM-5.3](/models/glm-5-3) | max reasoning | Z.AI | 88.13% |
| 37 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | high reasoning | Google | 87.88% |
| 38 | [Qwen3.5 Plus Thinking](https://www.vals.ai/models/alibaba_qwen3.5-plus-thinking) | — | Alibaba | 87.37% |
| 39 | [Qwen3.6 Plus](/models/qwen3-6-plus) | — | Alibaba | 87.37% |
| 40 | [Inkling](/models/inkling) | 0.99 reasoning | Thinking Machines Lab | 87.12% |
| 41 | [GPT-5.1](/models/gpt-5-1) | high reasoning | OpenAI | 86.62% |
| 42 | [MiniMax M2.7](/models/minimax-m2-7) | — | MiniMax | 86.62% |
| 43 | [GLM-5.3-Flash](/models/glm-5-3-flash) | max reasoning | Z.AI | 86.36% |
| 44 | [Nemotron 3 Ultra 550b A55b](https://www.vals.ai/models/nvidia_nemotron-3-ultra-550b-a55b) | — | Nvidia | 86.11% |
| 45 | [Claude Opus 4.5 Thinking](/models/claude-opus-4-5-thinking) | — | Anthropic | 85.86% |
| 46 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | high reasoning | OpenAI | 85.61% |
| 47 | [GLM-5.2](/models/glm-5-2) | — | Z.AI | 85.61% |
| 48 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | — | Anthropic | 85.61% |
| 49 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | — | xAI | 85.35% |
| 50 | [Ling 3.0 Flash 2607](https://www.vals.ai/models/ant_ling-3.0-flash-2607) | — | Ant | 84.85% |
| 51 | [Qwen3 Max](/models/qwen3-max) | — | Alibaba | 84.85% |
| 52 | [GLM-5.1](/models/glm-5-1) | — | Z.AI | 84.52% |
| 53 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | — | xAI | 84.34% |
| 54 | [o3](/models/o3) | high reasoning | OpenAI | 84.09% |
| 55 | [Kimi K2.5 Thinking](https://www.vals.ai/models/kimi_kimi-k2.5-thinking) | — | Moonshot AI | 84.09% |
| 56 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | high reasoning | Google | 83.84% |
| 57 | [Inkling-Small](/models/inkling-small) | 0.99 reasoning | Thinking Machines Lab | 83.59% |
| 58 | [GLM 5 Thinking](https://www.vals.ai/models/zai_glm-5-thinking) | — | Zhipu AI | 83.33% |
| 59 | [GPT-5.4 mini](/models/gpt-5-4-mini) | xhigh reasoning | OpenAI | 83.08% |
| 60 | [Qwen3.5 Flash](/models/qwen3-5-flash) | — | Alibaba | 82.83% |
| 61 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | — | Xiaomi | 82.58% |
| 62 | [MiniMax M2.5](/models/minimax-m2-5) | — | MiniMax | 82.07% |
| 63 | [Claude Sonnet 4.5 Thinking](/models/claude-sonnet-4-5-thinking) | — | Anthropic | 81.63% |
| 64 | [Gemini 2.5 Flash Preview 09 2025](https://www.vals.ai/models/google_gemini-2.5-flash-preview-09-2025) | — | Google | 81.57% |
| 65 | [MiMo-V2.5](/models/mimo-v2-5) | — | Xiaomi | 81.57% |
| 66 | [Gemini 3.1 Flash Lite Preview](https://www.vals.ai/models/google_gemini-3.1-flash-lite-preview) | high reasoning | Google | 81.06% |
| 67 | [Gemini 2.5 Pro Exp 03 25](https://www.vals.ai/models/google_gemini-2.5-pro-exp-03-25) | — | Google | 80.81% |
| 68 | [GPT-5 mini](/models/gpt-5-mini) | high reasoning | OpenAI | 80.30% |
| 69 | [DeepSeek V3p2 Thinking](https://www.vals.ai/models/fireworks_deepseek-v3p2-thinking) | high reasoning | Fireworks AI | 80.30% |
| 70 | [GLM-4.7](/models/glm-4-7) | — | Z.AI | 80.05% |
| 71 | [Claude Opus 4.5](/models/claude-opus-4-5) | — | Anthropic | 79.55% |
| 72 | [Qwen3 Max](/models/qwen3-max) | — | Alibaba | 79.55% |
| 73 | [Grok 3 Mini Fast High Reasoning](https://www.vals.ai/models/grok_grok-3-mini-fast-high-reasoning) | high reasoning | xAI | 79.29% |
| 74 | [GPT-OSS 120B](/models/gpt-oss-120b) | — | OpenAI | 78.54% |
| 75 | [MiniMax M2.1](https://www.vals.ai/models/minimax_MiniMax-M2.1) | — | MiniMax | 78.54% |
| 76 | [Kimi K2 Thinking](https://www.vals.ai/models/kimi_kimi-k2-thinking) | — | Moonshot AI | 78.54% |
| 77 | [Qwen3 Max Preview](https://www.vals.ai/models/alibaba_qwen3-max-preview) | — | Alibaba | 77.78% |
| 78 | [GPT-5.4 nano](/models/gpt-5-4-nano) | high reasoning | OpenAI | 77.53% |
| 79 | [Gemini 2.5 Flash Preview 09 2025 Thinking](https://www.vals.ai/models/google_gemini-2.5-flash-preview-09-2025-thinking) | — | Google | 76.52% |
| 80 | [Claude Opus 4.1 20250805 Thinking](https://www.vals.ai/models/anthropic_claude-opus-4-1-20250805-thinking) | — | Anthropic | 76.26% |
| 81 | [DeepSeek V3p2](https://www.vals.ai/models/fireworks_deepseek-v3p2) | none reasoning | Fireworks AI | 76.26% |
| 82 | [o3-mini](/models/o3-mini) | high reasoning | OpenAI | 75.50% |
| 83 | [Claude 3.7 Sonnet 20250219 Thinking](https://www.vals.ai/models/anthropic_claude-3-7-sonnet-20250219-thinking) | — | Anthropic | 75.25% |
| 84 | [Claude Sonnet 4 20250514 Thinking](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514-thinking) | — | Anthropic | 75.00% |
| 85 | [O4 Mini](https://www.vals.ai/models/openai_o4-mini-2025-04-16) | high reasoning | OpenAI | 74.50% |
| 86 | [GLM-4.6](/models/glm-4-6) | — | Z.AI | 74.50% |
| 87 | [Grok 3](https://www.vals.ai/models/grok_grok-3) | — | xAI | 74.24% |
| 88 | [o1](/models/o1) | high reasoning | OpenAI | 73.23% |
| 89 | [Grok 3 Mini Fast Low Reasoning](https://www.vals.ai/models/grok_grok-3-mini-fast-low-reasoning) | low reasoning | xAI | 72.98% |
| 90 | [Claude Haiku 4.5 Thinking](/models/claude-haiku-4-5-thinking) | — | Anthropic | 72.22% |
| 91 | [GLM-4.5](/models/glm-4-5) | — | Z.AI | 72.22% |
| 92 | [Claude Opus 4](https://www.vals.ai/models/anthropic_claude-opus-4-20250514) | — | Anthropic | 71.72% |
| 93 | [Moonshotai Kimi K2 Instruct](https://www.vals.ai/models/together_moonshotai/Kimi-K2-Instruct) | — | Together AI | 71.46% |
| 94 | [Gemini 2.5 Flash Lite Preview 09 2025 Thinking](https://www.vals.ai/models/google_gemini-2.5-flash-lite-preview-09-2025-thinking) | — | Google | 70.20% |
| 95 | [Qwen3 235b A22b](https://www.vals.ai/models/fireworks_qwen3-235b-a22b) | — | Fireworks AI | 70.20% |
| 96 | [Claude Opus 4.1](https://www.vals.ai/models/anthropic_claude-opus-4-1-20250805) | — | Anthropic | 69.95% |
| 97 | [Llama4 Maverick Instruct Basic](https://www.vals.ai/models/fireworks_llama4-maverick-instruct-basic) | — | Fireworks AI | 69.44% |
| 98 | [Claude Sonnet 4](https://www.vals.ai/models/anthropic_claude-sonnet-4-20250514) | — | Anthropic | 69.44% |
| 99 | [GPT-OSS 20B](/models/gpt-oss-20b) | — | OpenAI | 68.94% |
| 100 | [Mistral Large 2512](https://www.vals.ai/models/mistralai_mistral-large-2512) | — | Mistral AI | 68.43% |
| 101 | [GPT-4.1 mini](/models/gpt-4-1-mini) | high reasoning | OpenAI | 67.93% |
| 102 | [Claude 3.7 Sonnet](https://www.vals.ai/models/anthropic_claude-3-7-sonnet-20250219) | — | Anthropic | 67.42% |
| 103 | [GPT-4.1](/models/gpt-4-1) | high reasoning | OpenAI | 65.40% |
| 104 | [Gemini 2.0 Flash 001](https://www.vals.ai/models/google_gemini-2.0-flash-001) | — | Google | 65.15% |
| 105 | [Grok 4.1 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-1-fast-non-reasoning) | — | xAI | 65.15% |
| 106 | [Gemini 2.5 Flash Lite Preview 09 2025](https://www.vals.ai/models/google_gemini-2.5-flash-lite-preview-09-2025) | — | Google | 64.14% |
| 107 | [GPT-5 nano](/models/gpt-5-nano) | high reasoning | OpenAI | 63.38% |
| 108 | [Langston Nim Nvidia Llama 3.3 Nemotron Super 49b V1 42e84561 Thinking](https://www.vals.ai/models/together_langston/nim/nvidia/llama-3.3-nemotron-super-49b-v1-42e84561-thinking) | — | Together AI | 62.37% |
| 109 | [Magistral Medium 2509](https://www.vals.ai/models/mistralai_magistral-medium-2509) | — | Mistral AI | 62.37% |
| 110 | [Grok 4 Fast Non Reasoning](https://www.vals.ai/models/grok_grok-4-fast-non-reasoning) | — | xAI | 62.12% |
| 111 | [DeepSeek V3 0324](https://www.vals.ai/models/fireworks_deepseek-v3-0324) | — | Fireworks AI | 61.62% |
| 112 | [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) | — | Anthropic | 59.34% |
| 113 | [MiMo-V2-Flash](/models/mimo-v2-flash) | — | Xiaomi | 59.34% |
| 114 | [Gemini 1.5 Pro 002](https://www.vals.ai/models/google_gemini-1.5-pro-002) | — | Google | 58.33% |
| 115 | [Magistral Small 2509](https://www.vals.ai/models/mistralai_magistral-small-2509) | — | Mistral AI | 58.33% |
| 116 | [Gemini 2.5 Flash Preview 04 17](https://www.vals.ai/models/google_gemini-2.5-flash-preview-04-17) | — | Google | 57.58% |
| 117 | [Laguna XS.2](/models/laguna-xs-2) | — | Poolside | 55.05% |
| 118 | [DeepSeek V3](/models/deepseek-v3) | — | DeepSeek | 54.55% |
| 119 | [GPT-4o](/models/gpt-4o) | high reasoning | OpenAI | 53.79% |
| 120 | [GPT-4.1 nano](/models/gpt-4-1-nano) | high reasoning | OpenAI | 50.76% |
| 121 | [Mistral Small 2402](https://www.vals.ai/models/mistralai_mistral-small-2402) | — | Mistral AI | 50.76% |
| 122 | [Grok 2 1212](https://www.vals.ai/models/grok_grok-2-1212) | — | xAI | 50.76% |
| 123 | [GPT-4o](/models/gpt-4o) | high reasoning | OpenAI | 50.25% |
| 124 | [Meta Llama Llama 3.3 70B Instruct Turbo](https://www.vals.ai/models/together_meta-llama/Llama-3.3-70B-Instruct-Turbo) | — | Together AI | 50.00% |
| 125 | [Command A 03 2025](https://www.vals.ai/models/cohere_command-a-03-2025) | — | Cohere | 48.48% |
| 126 | [Mistral Large 2411](https://www.vals.ai/models/mistralai_mistral-large-2411) | — | Mistral AI | 47.73% |
| 127 | [Langston Nim Nvidia Llama 3.3 Nemotron Super 49b V1 42e84561](https://www.vals.ai/models/together_langston/nim/nvidia/llama-3.3-nemotron-super-49b-v1-42e84561) | — | Together AI | 47.73% |
| 128 | [Meta Llama Llama 4 Scout 17B 16E Instruct](https://www.vals.ai/models/together_meta-llama/Llama-4-Scout-17B-16E-Instruct) | — | Together AI | 46.97% |
| 129 | [Gemini 1.5 Flash 002](https://www.vals.ai/models/google_gemini-1.5-flash-002) | — | Google | 45.96% |
| 130 | [Mistral Small 2503](https://www.vals.ai/models/mistralai_mistral-small-2503) | — | Mistral AI | 44.19% |
| 131 | [GPT-4o mini](/models/gpt-4o-mini) | high reasoning | OpenAI | 44.19% |
| 132 | [Claude 3.5 Haiku](https://www.vals.ai/models/anthropic_claude-3-5-haiku-20241022) | — | Anthropic | 37.88% |
| 133 | [Jamba Large 1.6](https://www.vals.ai/models/ai21labs_jamba-large-1.6) | — | AI21 Labs | 36.11% |
| 134 | [Mistral Medium 3.5](https://www.vals.ai/models/mistralai_mistral-medium-3.5) | high reasoning | Mistral AI | 34.85% |
| 135 | [Jamba Mini 1.6](https://www.vals.ai/models/ai21labs_jamba-mini-1.6) | — | AI21 Labs | 32.07% |
| 136 | [Command R Plus](https://www.vals.ai/models/cohere_command-r-plus) | — | Cohere | 31.06% |
| 137 | [GPT-3.5 Turbo](https://www.vals.ai/models/openai_gpt-3.5-turbo) | high reasoning | OpenAI | 30.56% |
| 138 | [Laguna M.1](/models/laguna-m-1) | — | Poolside | 27.02% |

## FAQ

### What does Vals GPQA Diamond mirror measure?

Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.

### Which model leads the published Vals GPQA Diamond mirror snapshot?

Gemini 3.1 Pro Preview currently leads the published Vals GPQA Diamond mirror snapshot with a score of 95.45%.

### How many models are evaluated on Vals GPQA Diamond mirror?

The September 1, 2026 contains 138 AI models.

### Does Vals GPQA Diamond mirror affect BenchLM's overall score?

Not directly. Vals GPQA Diamond mirror is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
