InclusionAI · Aug 4, 2026
Source confirmed
LLM leaderboard, August 2026
Compare frontier AI models by quality, cost, and context. 104 Supported and 111 Estimated models among 378 tracked LLMs — 381 benchmarks, real pricing, and runtime data in one place.
InclusionAI · Aug 4, 2026
Source confirmed
LiquidAI · Aug 4, 2026
Source confirmed
Alibaba · Aug 3, 2026
Source confirmed
New model, price change, API update, or outage: Radar sends what matters to your inbox with the source. Try it free for seven days; plans start at $4.99/month.
New models, price changes, benchmark moves, and the releases worth waiting for—one concise email each week.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.
Current models that clear the ranking, freshness, and evidence thresholds for each decision.
Benchmarks, pricing, runtime signals, and context window in one table. Filter state syncs to the URL so every view is shareable. Supported and Estimated labels show how much independent evidence backs each position.
1 Claude Mythos 5 AnthropicSupported | Anthropic | Closed | Current | Reasoning | 1M+ | $10.00 / $50.00 | Not measured | Not measured | 83.04 | 76 | 80 | — | 86 | 94 | — | — | — | — |
2 Claude Fable 5 AnthropicSupported | Anthropic | Closed | Current | Reasoning | 1M+ | $10.00 / $50.00 | 73 | 83.30s | 82.79 | 75 | 80 | — | 63 | 71 | — | — | — | 1508.58 |
3 Claude Opus 5 AnthropicSupported | Anthropic | Closed | Current | Reasoning | — | $5.00 / $25.00 | 56 | 63.43s | 82.59 | 81 | 77 | 91 | 89 | 94 | — | — | — | 1491.82 |
4 GPT-5.6 Sol OpenAISupported | OpenAI | Closed | Current | Reasoning | 1.05M | $5.00 / $30.00 | 71 | 131.39s | 81.48 | 68 | 78 | 93 | 85 | 83 | — | — | 97 | 1482.77 |
5 Kimi K3 Moonshot AISupported | Moonshot AI | Closed | Current | Reasoning | 1.05M | $3.00 / $15.00 | 36 | 58.35s | 79.89 | 74 | 77 | — | 88 | 85 | — | — | — | 1485.33 |
6 Claude Opus 4.8 AnthropicSupported | Anthropic | Closed | Superseded | Reasoning | 1M | $5.00 / $25.00 | 58 | 22.68s | 77.34 | 61 | 72 | 76 | 88 | 88 | — | — | 67 | 1483.64 |
7 Muse Spark 1.1 MetaSupported | Meta | Closed | Current | Reasoning | 1M | Not listed | 216 | 11.92s | 76.15 | 65 | 64 | — | 76 | 94 | — | — | — | 1489.67 |
8 Grok 4.5 xAISupported | xAI | Closed | Current | Reasoning | 500K | $2.00 / $6.00 | 61 | 10.38s | 75.38 | 61 | 57 | 60 | 68 | 69 | — | — | — | 1468.83 |
9 Gemini 3.6 Flash GoogleSupported | Closed | Current | Reasoning | 1M | $1.50 / $7.50 | 213 | 16.09s | 75.3 | 47 | 64 | — | 74 | 70 | — | — | — | 1482.68 | |
10 GPT-5.4 OpenAISupported | OpenAI | Closed | Superseded | Reasoning | 1.05M | $2.50 / $15.00 | 143 | 130.50s | 73.2 | 57 | 58 | 78 | 67 | 77 | — | — | 66 | 1465.37 |
11 GPT-5.5 OpenAIEstimated | OpenAI | Closed | Superseded | Reasoning | 1M | $5.00 / $30.00 | 87 | 65.57s | 72.28 | 64 | 71 | 87 | 68 | 79 | — | — | 71 | 1482.44 |
12 GPT-5.6 Terra OpenAIEstimated | OpenAI | Closed | Current | Reasoning | 1.05M | $2.00 / $12.00 | Not measured | Not measured | 72.26 | 63 | 65 | 86 | 75 | 82 | — | — | 97 | — |
13 Qwen3.7 Max AlibabaSupported | Alibaba | Closed | Superseded | Reasoning | 1M | Not listed | 205 | 14.07s | 71.8 | 40 | 54 | 87 | — | 72 | 100 | 93 | 82 | 1474.66 |
14 Claude Opus 4.7 (Adaptive) AnthropicEstimated | Anthropic | Closed | Superseded | Reasoning | 1M | $5.00 / $25.00 | Not measured | Not measured | 71.46 | 63 | 64 | 79 | 48 | 83 | — | — | — | 1501.5 |
15 Claude Opus 4.7 AnthropicSupported | Anthropic | Closed | Current | Standard | 1M | $5.00 / $25.00 | 52 | 22.82s | 71.21 | 56 | 67 | — | — | 64 | — | — | 62 | 1492.41 |
16 Muse Spark MetaSupported | Meta | Closed | Superseded | Reasoning | 262K | Not listed | Not measured | Not measured | 70.49 | 58 | 62 | 52 | 76 | 75 | — | — | 56 | 1487.88 |
17 MiMo-V2.5-Pro XiaomiSupported | Xiaomi | Closed | Current | Reasoning | 1M | Not listed | Not measured | Not measured | 69.21 | 42 | 55 | — | — | 71 | — | — | — | 1465.71 |
18 | MiniMax | Open | Current | Standard | 1M | $0.30 / $1.20 | 92 | 23.04s | 68.8 | 40 | 48 | — | 48 | 66 | — | — | — | 1445.27 |
| Tencent | Open | Current | Reasoning | 256K | $0.00 / $0.00 | 70 | 31.36s | 67.91 | 61 | 63 | — | — | — | — | — | — | 1456.27 | |
20 Claude Opus 4.6 AnthropicSupported | Anthropic | Closed | Superseded | Standard | 1M | $5.00 / $25.00 | 37 | 2.51s | 67.74 | 56 | 61 | — | 62 | 83 | — | — | 60 | 1497.19 |
21 Gemini 3 Pro GoogleSupported | Closed | Established | Standard | 2M | $2.00 / $12.00 | 109 | 32.65s | 66.96 | 58 | 61 | 42 | 73 | 69 | — | — | 56 | 1485.73 | |
| Z.AI | Open | Superseded | Reasoning | 203K | $1.40 / $4.40 | 73 | 53.82s | 66.9 | 49 | 55 | — | — | 78 | — | — | 65 | 1468.79 | |
23 GPT-5.6 Luna OpenAIEstimated | OpenAI | Closed | Current | Reasoning | 1.05M | $0.20 / $1.20 | Not measured | Not measured | 66.85 | 56 | 72 | 66 | 66 | 81 | — | — | 97 | — |
24 MiMo-V2-Pro XiaomiSupported | Xiaomi | Closed | Superseded | Reasoning | 1M | Not listed | Not measured | Not measured | 66.76 | — | 61 | — | — | — | — | — | — | 1448.24 |
| Thinking Machines Lab | Open | Current | Standard | 1M | $1.87 / $4.68 | 85 | 25.24s | 66.5 | 49 | 43 | — | 51 | 69 | — | 91 | 79 | 1440.8 |
The BenchLM LLM leaderboard 2026 ranks 215+ models and tracks 378+ large language models side by side across 381 benchmarks — from SWE-bench and LiveCodeBench for coding to GPQA Diamond and MMLU-Pro for knowledge and reasoning. Whether you need the best AI models 2026 has to offer for agentic workflows, math, multilingual tasks, or instruction following, our AI benchmark comparison tables make it easy to see how GPT-5, Claude, Gemini, DeepSeek, Llama, and dozens of other frontier and open-source models stack up on both benchmarks and operator tradeoffs like price and context. Supported and Estimated labels separate evidence strength from score, so incomplete reporting does not automatically remove a model from the comparison.
Dataset journal
Choose any two ranked models. The comparison opens with the decision, exact units, and evidence coverage.
The evidence map
Move from the overall ranking into the capability or workload that actually decides your model choice.
Six-month release record
Track which release led its month, how provider depth changed, and whether the crown moved.
Best ranked model released in Aug 2026
Qwen3.8 MaxAlibabaEach model's overall score is a normalized weighted average of category averages. Within each category, benchmark results are normalized to a common scale and combined using weights that favor harder, less-saturated evaluations.
Confidence remains separate from score. It shows how much public benchmark evidence supports a result, while display-only benchmarks stay visible for context without receiving ranking weight.
Data comes from OpenBench, official model papers, and public leaderboards. External consensus signals are bounded calibration inputs and are not exposed in exported data.
Read the complete methodologyTerminal-Bench 2.0 · BrowseComp · OSWorld-Verified
3 weighted · 79 context-only
SWE-bench Verified · SWE-Rebench · LiveCodeBench · SWE-bench Pro · SciCode
5 weighted · 49 context-only
LongBench v2 · MRCRv2 · ARC-AGI-2
3 weighted · 23 context-only
MMMU-Pro · OfficeQA Pro · CharXiv
3 weighted · 62 context-only
GPQA · SuperGPQA · MMLU-Pro · HLE · SimpleQA
5 weighted · 41 context-only
MMLU-ProX
1 weighted · 12 context-only
IFEval · IFBench
2 weighted · 2 context-only
AIME26 · HMMT Feb 2026 · FrontierMath v2 (Tiers 1-3) · FrontierMath v2 (Tier 4) · USAMO 2026
5 weighted · 27 context-only