# Conceptual Reasoning Benchmark (Conceptual Reasoning)

> Tests whether model judgments rank argumentative critiques in the same order as expert human ratings across philosophy, AI alignment, and other concept-heavy texts.

Canonical page: https://benchlm.ai/benchmarks/conceptual-reasoning

- Category: [Reasoning](/reasoning)
- Last updated: September 23, 2026

## About Conceptual Reasoning

- Tasks: 224 texts and 608 within-text critique pairs
- Format: Average pairwise-ranking loss against expert ratings
- Difficulty: Fuzzy, expert-rated argumentative reasoning
- Paper: [Conceptual Reasoning Benchmark Results](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html)

The benchmark asks models to score critiques of argumentative texts, then measures whether the induced within-text ordering agrees with expert ratings. A reversed pair incurs loss proportional to the expert rating gap. We mirror the 22 published endpoint rows and their 95% confidence intervals as display-only evidence.

Conceptual Reasoning is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (22 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Sonnet 4 (2025-05-14)](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.017 · claude-sonnet-4-20250514 | Anthropic | 0.072 |
| 2 | [Claude Opus 4 (2025-05-14)](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.018 · claude-opus-4-20250514 | Anthropic | 0.073 |
| 3 | [o3-pro](/models/o3-pro) | 95% CI ±0.019 · o3-pro-2025-06-10 | OpenAI | 0.081 |
| 4 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | 95% CI ±0.018 · gemini-2.5-pro | Google | 0.081 |
| 5 | [Gemini 2.5 Flash](/models/gemini-2-5-flash) | 95% CI ±0.019 · gemini-2.5-flash | Google | 0.081 |
| 6 | [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) | 95% CI ±0.019 · claude-3-5-sonnet-20241022 | Anthropic | 0.082 |
| 7 | [o4-mini (2025-04-16)](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.019 · o4-mini-2025-04-16 | OpenAI | 0.087 |
| 8 | [GPT-4.1](/models/gpt-4-1) | 95% CI ±0.019 · gpt-4.1-2025-04-14 | OpenAI | 0.091 |
| 9 | [o3](/models/o3) | 95% CI ±0.022 · o3-2025-04-16 | OpenAI | 0.093 |
| 10 | [Magistral Medium 2506](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.022 · magistral-medium-2506 | Mistral AI | 0.102 |
| 11 | [o1](/models/o1) | 95% CI ±0.022 · o1-2024-12-17 | OpenAI | 0.105 |
| 12 | [GPT-4o](/models/gpt-4o) | 95% CI ±0.022 · gpt-4o-2024-08-06 | OpenAI | 0.112 |
| 13 | [Gemini 1.5 Pro](/models/gemini-1-5-pro) | 95% CI ±0.023 · gemini-1.5-pro-002 | Google | 0.114 |
| 14 | [Claude 3 Opus](/models/claude-3-opus) | 95% CI ±0.025 · claude-3-opus-20240229 | Anthropic | 0.118 |
| 15 | [Gemini 2.5 Flash-Lite](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.025 · gemini-2.5-flash-lite | Google | 0.123 |
| 16 | [Claude 3.5 Sonnet](/models/claude-3-5-sonnet) | 95% CI ±0.027 · claude-3-5-sonnet-20240620 | Anthropic | 0.126 |
| 17 | [GPT-4.1 mini](/models/gpt-4-1-mini) | 95% CI ±0.027 · gpt-4.1-mini-2025-04-14 | OpenAI | 0.128 |
| 18 | [Magistral Small 2506](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.027 · magistral-small-2506 | Mistral AI | 0.137 |
| 19 | [Claude 3 Haiku](/models/claude-3-haiku) | 95% CI ±0.027 · claude-3-haiku-20240307 | Anthropic | 0.139 |
| 20 | [Claude 3.5 Haiku (2024-10-22)](https://www.andrew.cmu.edu/user/coesterh/conceptual_reasoning_benchmark.html) | 95% CI ±0.029 · claude-3-5-haiku-20241022 | Anthropic | 0.146 |
| 21 | [GPT-4o mini](/models/gpt-4o-mini) | 95% CI ±0.030 · gpt-4o-mini-2024-07-18 | OpenAI | 0.150 |
| 22 | [GPT-4.1 nano](/models/gpt-4-1-nano) | 95% CI ±0.028 · gpt-4.1-nano-2025-04-14 | OpenAI | 0.157 |

## FAQ

### What does the Conceptual Reasoning benchmark measure?

It measures whether an AI model ranks critiques of the same argumentative text in the same order as an expert human rater. The benchmark focuses on concept-heavy arguments rather than factual recall or mechanically verified answers.

### Which model has the lowest published loss?

Claude Sonnet 4 dated May 14, 2025 has the lowest published average loss at 0.072, with a reported 95% confidence interval of ±0.017. Claude Opus 4 is nearly tied at 0.073 ±0.018.

### Does Conceptual Reasoning affect BenchLM rankings?

No. We keep the benchmark display only because the authors describe it as an early version and have not yet released the dataset or evaluation code.
