Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Reasoning benchmark report

Best LLMs for ReasoningSeptember 2026 Leaderboard

Data refreshed:

Logical reasoning and problem solving

As of September 2026, the top reasoning model on the BenchLM leaderboard is GPT-6 Astra with a weighted reasoning score of 89.5.

Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.

Data refreshed
September 15, 2026
Provisional-ranked
20 of 492 models
Verified-ranked
10 of 492 models
Weighted evidence
3 of 11 benchmarks
11 tracked benchmarks

MuSR, BBH, LisanBench, Pencil Puzzle Bench, LongBench v2, MRCRv2, MRCR v2 64K-128K, MRCR v2 128K-256K, Graphwalks BFS 128K, Graphwalks Parents 128K, ARC-AGI-2

Scope: Abstract reasoning, Long-context reasoning

Best Reasoning picks

BenchLM summaries for reasoning plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Reasoning Leaderboard

Primary score: weighted reasoning score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.

Filters
Provisional-ranked mode includes source-unverified non-generated benchmark evidence.P = provisional benchmark row
Rank / modelWeighted Reasoning
1
GPT-6 AstraOpenAI · Closed
89.5%
2
Claude Fable 5.1Anthropic · Closed
79.4%
3
Kimi K3Moonshot AI · Closed
78.5%
4
MiniMax M3MiniMax · Open weight
78%
5
Muse Spark 1.3Meta · Closed
78%
6
Claude Fable 5Anthropic · Closed
77.6%
7
Qwen3.8-27BAlibaba · Open weight
77.4%
8
Gemini 3.8 FlashGoogle · Closed
76.9%
9
Grok 4.6xAI · Closed
76.2%
10
GLM-5.3Z.AI · Open weight
75.8%
11
Claude Opus 5Anthropic · Closed
75.7%
12
GPT-5.6 SolOpenAI · Closed
70.4%
13
GPT-5.5OpenAI · Closed
64%
14
GPT-5.6 TerraOpenAI · Closed
63.7%
15
DeepSeek V4 Pro 0813DeepSeek · Closed
60.3%
16
Gemini 3.5 FlashGoogle · Closed
59.4%
17
GPT-5.4OpenAI · Closed
57.2%
18
Claude Opus 4.8Anthropic · Closed
55.5%
19
GPT-5.6 LunaOpenAI · Closed
51.2%
20
Grok 4.5xAI · Closed
44.8%

Top AI Models for ReasoningSeptember 2026

As of September 2026, GPT-6 Astra leads the provisional reasoning leaderboard with a score of 89.5%, followed by Claude Fable 5.1 (79.4%) and Kimi K3 (78.5%). BenchLM is currently showing 20 provisional-ranked models and 10 verified-ranked models in this category.

What changed

Qwen3.8 Max leads the current reasoning table on its source-backed weighted rows.

GPT-5.4 strong #2, with a notable edge on LongBench v2.

Claude Opus 4.6 holds #3, excelling on long-context reasoning tasks.

Top models by benchmark

Long-context reasoning and retrieval benchmark(25% of category score)

RankModelReported score

Score in Context

What these scores mean

Reasoning carries a 17% weight in overall scoring. The weighted score blends long-context comprehension (LongBench v2, MRCRv2), multi-step inference (MuSR), and novel problem solving (ARC-AGI-2). A 5-point gap here usually means the difference between a model that tracks complex argument chains reliably and one that loses the thread.

Known limitations

Models with explicit chain-of-thought (reasoning models) tend to outperform standard models by large margins, but at the cost of higher latency and token usage. ARC-AGI-2 is still early — coverage is uneven, and some models lack scores. MuSR is underrepresented because few providers run it.

How we weight

Reasoning carries a 17% weight in BenchLM.ai's overall scoring, reflecting the growing importance of long-context reasoning in production systems.

Models with explicit chain-of-thought capabilities tend to outperform standard models by significant margins on MuSR and long-context tasks, though at the cost of higher latency and token usage. See the reasoning leaderboard or try the LLM selector quiz.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Reasoning benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
MuSRDisplay onlyComplex multi-step reasoning problems
BBHDisplay only23 challenging tasks from BIG-Bench where language models previously underperformed humans
LisanBenchDisplay onlyWord-chain reasoning benchmark for planning, recall, and constraint following.
Pencil Puzzle BenchDisplay onlyMulti-step verifiable reasoning benchmark built from pencil puzzles with unique solutions.
LongBench v225%WeightedLong-context reasoning and retrieval benchmark
MRCRv220%WeightedMulti-round coreference and retrieval benchmark for long-context models
MRCR v2 64K-128KDisplay onlyLong-context retrieval benchmark slice focused on 64K-128K context lengths
MRCR v2 128K-256KDisplay onlyLong-context retrieval benchmark slice focused on 128K-256K context lengths
Graphwalks BFS 128KDisplay onlyLong-context graph traversal benchmark using breadth-first search tasks
Graphwalks Parents 128KDisplay onlyLong-context graph reasoning benchmark for parent-retrieval accuracy
ARC-AGI-225%WeightedAbstract reasoning

About Reasoning Benchmarks

Complex multi-step reasoning problems

Common questions

What is the best LLM for reasoning?

The top reasoning LLMs are ranked using benchmarks like MuSR, SimpleQA, LongBench v2, and MRCRv2, which test logical deduction, multi-step reasoning, factual accuracy, and long-context discipline.

How do reasoning benchmarks evaluate LLMs?

Reasoning benchmarks evaluate LLMs by presenting tasks that require multi-step logical deduction, causal inference, and complex problem solving beyond simple pattern matching.

What is the difference between reasoning and knowledge benchmarks?

Reasoning benchmarks test logical thinking and multi-step inference, while knowledge benchmarks focus on factual recall. A model can have strong reasoning but limited factual knowledge, or vice versa.

Reasoning benchmark updates

Reasoning benchmarks shift fast. Get the update before you commit to a model.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Related