# Best Value LLM for Reasoning in 2026 — Cost-Adjusted Rankings

> Top AI models ranked by reasoning benchmark performance per dollar. Cost-adjusted rankings using ARC-AGI-2, LongBench v2, MRCRv2, and MuSR.

Reasoning models tend to be the most expensive tier — they use chain-of-thought, produce more output tokens, and are priced accordingly. This ranking divides each model's weighted reasoning score by output token price, revealing which models deliver the best abstract reasoning, long-context comprehension, and multi-step logic per dollar. For applications that need strong reasoning without frontier-model budgets, the value leaders here are worth serious consideration.

Canonical page: https://benchlm.ai/best/best-value-reasoning

Last updated: October 8, 2026

## Rankings

| Rank | Model | Creator | Type | Context | Score/$ | Score | Output $/1M |
|------|-------|---------|------|---------|---------|-------|-------------|
| 1 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | Proprietary | 1.05M | 111.8 | 55.9 | $0.5 |
| 2 | [MiniMax M3](/models/minimax-m3) | MiniMax | Open Weight | 1M | 66.17 | 79.4 | $1.2 |
| 3 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | Proprietary | 1.05M | 45.42 | 54.5 | $1.2 |
| 4 | [Mistral Large 4](/models/mistral-large-4) | Mistral | Pending | 1M | 37.42 | 78.2 | $2.09 |
| 5 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | Proprietary | 1M | 18.83 | 70.6 | $3.75 |
| 6 | [Muse Spark 1.3](/models/muse-spark-1-3) | Meta | Proprietary | 1M | 18.68 | 79.4 | $4.25 |
| 7 | [Inkling](/models/inkling) | Thinking Machines Lab | Open Weight | 1M | 16.13 | 75.5 | $4.68 |
| 8 | [Grok 4.7](/models/grok-4-7) | xAI | Proprietary | 500K | 12.5 | 75 | $6 |
| 9 | [Grok 4.6](/models/grok-4-6) | xAI | Proprietary | 500K | 9.62 | 57.7 | $6 |
| 10 | [GPT-6.1 Sol](/models/gpt-6-1-sol) | OpenAI | Proprietary | 1.05M | 8.68 | 86.8 | $10 |
| 11 | [Grok 4.5](/models/grok-4-5) | xAI | Proprietary | 500K | 8.4 | 50.4 | $6 |
| 12 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Anthropic | Proprietary | 1M | 7.92 | 79.2 | $10 |
| 13 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | Proprietary | 1M | 6.98 | 62.8 | $9 |
| 14 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | Proprietary | 1.05M | 6.95 | 69.5 | $10 |
| 15 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | Proprietary | 1.05M | 5.46 | 65.5 | $12 |
| 16 | [Kimi K3](/models/kimi-k3) | Moonshot AI | Pending | 1.05M | 4.41 | 66.2 | $15 |
| 17 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | Proprietary | 1M | 4.12 | 82.4 | $20 |
| 18 | [GPT-5.4](/models/gpt-5-4) | OpenAI | Proprietary | 1.05M | 4.04 | 60.6 | $15 |
| 19 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | Proprietary | 1.05M | 3.64 | 72.8 | $20 |
| 20 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | Proprietary | null | 3.09 | 77.3 | $25 |
| 21 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | Proprietary | 1M | 2.36 | 59.1 | $25 |
| 22 | [GPT-5.5](/models/gpt-5-5) | OpenAI | Proprietary | 1M | 2.22 | 66.7 | $30 |
| 23 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | Proprietary | 1.05M | 1.79 | 89.6 | $50 |
| 24 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | Proprietary | 1M | 1.63 | 81.4 | $50 |
| — | [Mistral Large 4](/models/mistral-large-4) | Mistral | Pending | 1M | — | — | — |

## Key Takeaways

- Top model: [GPT-6 Luna](/models/gpt-6-luna) with a Score/$ of 111.8 (score 55.9, $0.5/1M output tokens)
- Best open-weight option: [MiniMax M3](/models/minimax-m3) at #2
- Models included: 25 (24 ranked, 1 listed without a rank)

## Compare the Leaders

- [GPT-6 Luna vs MiniMax M3](/compare/gpt-6-luna-vs-minimax-m3)
