# Best LLMs for Reasoning — September 2026 Leaderboard

> As of September 2026, GPT-6 Astra leads BenchLM's reasoning leaderboard with a weighted score of 89.5.

- **Last verified:** September 15, 2026
- Canonical page: https://benchlm.ai/reasoning
- **Ranking coverage:** 20 category-ranked models from 486 tracked models
- **Category weight:** 17% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 89.5 | 9 | 34 total |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 79.4 | 5 | 33 total |
| 3 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 78.5 | 3 | 52 total |
| 4 | [MiniMax M3](/models/minimax-m3) | MiniMax | 78 | 2 | 32 total |
| 5 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | 78 | 5 | 17 total |
| 6 | [Claude Fable 5](/models/claude-fable) | Anthropic | 77.6 | 3 | 31 total |
| 7 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 77.4 | 3 | 42 total |
| 8 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 76.9 | 3 | 29 total |
| 9 | [Grok 4.6](/models/grok-4-6) | xAI | 76.2 | 2 | 23 total |
| 10 | [GLM-5.3](/models/glm-5-3) | Z.AI | 75.8 | 3 | 32 total |
| 11 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 75.7 | 6 | 84 total |
| 12 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 70.4 | 6 | 47 total |
| 13 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 64 | 6 | 44 total |
| 14 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 63.7 | 5 | 44 total |
| 15 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 60.3 | 5 | 36 total |
| 16 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 59.4 | 5 | 38 total |
| 17 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 57.2 | 4 | 40 total |
| 18 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 55.5 | 4 | 44 total |
| 19 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 51.2 | 5 | 38 total |
| 20 | [Grok 4.5](/models/grok-4-5) | xAI | 44.8 | 4 | 24 total |

## Decision-ready shortlist

- #1 [GPT-6 Astra](/models/gpt-6-astra) — 89.5 weighted score, Proprietary, 1.05M context.
- #2 [Claude Fable 5.1](/models/claude-fable-5-1) — 79.4 weighted score, Proprietary, 1M context.
- #3 [Kimi K3](/models/kimi-k3) — 78.5 weighted score, Pending, 1.05M context.
- #4 [MiniMax M3](/models/minimax-m3) — 78 weighted score, Open Weight, 1M context.
- #5 [Muse Spark 1.3](/models/muse-spark-1-3) — 78 weighted score, Proprietary, 1M context.

## Benchmarks in this category

### [MuSR](/benchmarks/musr) (Testing the Limits of Chain-of-thought with Multistep Soft Reasoning)

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

- Ranking status: Display only
- Year: 2023
- Format: Narrative-based reasoning
- Difficulty: Complex reasoning tasks

### [BBH](/benchmarks/bbh) (BIG-Bench Hard)

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

- Ranking status: Display only
- Year: 2022
- Format: Mixed reasoning tasks
- Difficulty: Advanced reasoning

### [LisanBench](/benchmarks/lisanbench) (LisanBench)

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

- Ranking status: Display only
- Year: 2026
- Format: Difficulty-weighted word-chain reasoning
- Difficulty: Open-ended lexical planning

### [Pencil Puzzle Bench](/benchmarks/ppbench) (Pencil Puzzle Bench)

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

- Ranking status: Display only
- Year: 2026
- Format: Direct and agentic puzzle solve rate
- Difficulty: Multi-step verifiable reasoning

### [LongBench v2](/benchmarks/longbench-v2) (LongBench v2)

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

- Ranking status: Weighted (25% of this category)
- Year: 2025
- Format: Extended-context retrieval and reasoning
- Difficulty: Hard long-context

### [MRCRv2](/benchmarks/mrcrv2) (MRCRv2)

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

- Ranking status: Weighted (20% of this category)
- Year: 2025
- Format: Multi-round long-context evaluation
- Difficulty: Hard long-context

### [MRCR v2 64K-128K](/benchmarks/mrcr-v2-64k-128k) (OpenAI MRCR v2 8-needle 64K-128K)

MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.

- Ranking status: Display only
- Year: 2026
- Format: Long-context retrieval
- Difficulty: Long-context reasoning

### [MRCR v2 128K-256K](/benchmarks/mrcr-v2-128k-256k) (OpenAI MRCR v2 8-needle 128K-256K)

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

- Ranking status: Display only
- Year: 2026
- Format: Very-long-context retrieval
- Difficulty: Very long-context reasoning

### [Graphwalks BFS 128K](/benchmarks/graphwalksbfs128k) (Graphwalks BFS 0K-128K)

Long-context graph traversal benchmark using breadth-first search tasks.

- Ranking status: Display only
- Year: 2026
- Format: Long-context graph reasoning
- Difficulty: Algorithmic long-context reasoning

### [Graphwalks Parents 128K](/benchmarks/graphwalksparents128k) (Graphwalks parents 0-128K)

Long-context benchmark for recovering parent relationships inside graph tasks.

- Ranking status: Display only
- Year: 2026
- Format: Long-context graph reasoning
- Difficulty: Algorithmic long-context reasoning

### [ARC-AGI-2](/benchmarks/arc-agi-2) (Abstraction and Reasoning Corpus for AGI v2)

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

- Ranking status: Weighted (25% of this category)
- Year: 2025
- Format: Grid transformation puzzles with novel rules
- Difficulty: Expert-level — hardest public reasoning benchmark
