Choose the failure first, then the model. A coding assistant that writes elegant prose but breaks builds is a bad coding assistant. An agent that scores well overall but needs constant intervention is a bad agent. A frontier API that costs more than the workflow saves is a bad business decision.
We use the overall ranking as a shortlist, not a purchase order. Use the relevant capability view, check evidence strength and access, then test the top two affordable candidates on your own work.
The 60-second answer
Names in this table re-render from the live BenchAlign v5.8 board on every build, so they show the current leaders, not the ones from the day this guide was written.
| Need | Start with | Why | Main caveat |
|---|---|---|---|
| Highest capability | Claude Opus 5.5 (86.4, Supported) | #1 overall | Check access before it becomes your default |
| Best generally available model | The highest overall row you can call today | Restricted rows cannot carry production traffic | Model pages carry the access notes |
| Coding | Claude Sonnet 5.5 (85, Supported) | #1 coding | Read the evidence label before the score |
| Agents and tool use | Claude Opus 5.5 (88.2, Supported) | #1 agentic | Task score hides intervention rate |
| Budget general use | MiMo-V2.6-Pro (74.1, Estimated) | Top overall row at or below $1.50 per 1M input tokens | Check its agentic row before long autonomous chains |
| Open-weight control | MiMo-V2.6-Pro (74.1, Estimated) | #1 open-weight overall | Price the serving stack, not just the weights |
| Open coding | MiMo-V2.6-Pro (63.8, Supported) | #1 open-weight coding | Validate on your repository before self-hosting |
Start with the task
Coding
Claude Sonnet 5.5 (85, Supported) leads the coding ranking on BenchAlign v5.8.
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 | Reasoning | Proprietary | 85 | Supported |
| 2 | Claude Opus 5.5 | Reasoning | Proprietary | 83.5 | Supported |
| 3 | Claude Fable 5.1 | Reasoning | Proprietary | 77.8 | Supported |
| 4 | GPT-6 Astra | Reasoning | Proprietary | 76.5 | Supported |
| 5 | Gemini 4 Argon | Reasoning | Proprietary | 72.8 | Supported |
Read the Evidence column before the Score column. An Estimated row can sit at the top on thin evidence, and adjacent rows a point apart do not settle a repository decision. Put the two best affordable candidates against real issues from your codebase. Measure accepted patches, tests passed without repair, regressions, review time, and token cost. A model that needs one fewer repair loop can be cheaper even at a higher list price.
Agents and tool use
Claude Opus 5.5 (88.2, Supported) leads the agentic ranking.
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Reasoning | Proprietary | 88.2 | Supported |
| 2 | Claude Fable 5.1 | Reasoning | Proprietary | 79 | Supported |
| 3 | Claude Opus 5 | Reasoning | Proprietary | 77.5 | Supported |
| 4 | Gemini 4 Argon | Reasoning | Proprietary | 75 | Estimated |
| 5 | Claude Fable 5 | Reasoning | Proprietary | 74.5 | Supported |
Rank and list price do not move together. Price the top five on live pricing before you name a value candidate.
Do not optimize an agent for task score alone. Track intervention rate, recovery after a bad tool call, permission errors, time to completion, and cost per successful run. The best model is the one that completes your workflow safely, not the one that begins it most impressively.
General knowledge work
Use the overall ranking to build the shortlist.
| Rank | Model | Type | License | Score | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Reasoning | Proprietary | 86.4 | Supported |
| 2 | GPT-6 Astra | Reasoning | Proprietary | 84.8 | Supported |
| 3 | Claude Sonnet 5.5 | Reasoning | Proprietary | 83.9 | Supported |
| 4 | Gemini 4 Argon | Reasoning | Proprietary | 81.8 | Estimated |
| 5 | Claude Fable 5.1 | Reasoning | Proprietary | 81.7 | Supported |
Then strike what you cannot use. Access removes a restricted model whatever its rank: in the July 14 snapshot the leader was Claude Mythos 5, which most teams could not call, and Mythos no longer has a public row. Price may remove the next row. What survives, plus the rest of the top ten, is the real candidate list for the workload.
High-volume summarization, extraction, research preparation, and document workflows deserve a look at the cheap tier. Among rows priced at or below $1.50 per million input tokens, MiMo-V2.6-Pro (74.1, Estimated) ranks highest overall. Check its agentic row and keep long autonomous chains out of the default plan unless your evaluation proves otherwise.
Writing and conversation
The current catalog does not justify a universal creative-writing winner. Public writing data is thinner, more preference-sensitive, and easier to contaminate than coding or agentic evidence. Claude may still be the right first candidate, but that is a hypothesis to test.
Create a blind set of briefs, rewrites, edits, and difficult tone corrections. Have the people who own the voice grade adherence, originality, edit distance, and factual discipline. Do not borrow a coding rank as proof of writing quality.
Math, multilingual, and multimodal work
Treat these as lenses. Select the models with direct evidence for the exact modality or language, then evaluate on representative inputs. A text-only configuration should not lose the core text-model rank because it cannot see an image, but modality support should be explicit in a use-case recommendation.
For multilingual work, test your actual language pairs and domains. An aggregate can hide a model that is excellent in French and weak in Japanese. For multimodal work, distinguish image understanding, document extraction, charts, video, and computer use; they are not interchangeable capabilities.
Choose the operating model
Proprietary API
Choose an API when peak capability, rapid upgrades, and low infrastructure overhead matter. The costs are vendor dependence, variable behavior after model updates, and less control over data handling. Keep an evaluation gate between a provider update and production traffic.
Open weight
Choose open weight when data must stay inside your environment, you need fine-tuning or serving control, or sustained volume can justify infrastructure. MiMo-V2.6-Pro (74.1, Estimated) leads the current open-weight overall ranking, and MiMo-V2.6-Pro (63.8, Supported) leads open-weight coding.
Open weight is not automatically cheaper. Include GPUs, idle capacity, engineering time, observability, safety controls, and upgrades. It can still be the only acceptable answer when control is part of the requirement.
→ Best open-weight models · Self-host calculator
Read the evidence label
Supported and Estimated are not separate leagues. They are confidence signals attached to one ranking.
- Supported means the row has enough independent evidence to carry normally.
- Estimated means the available evidence can place the model, but the uncertainty is wider.
An Estimated model can be excellent. Claude Opus 5.5 entered the overall ranking in first place on September 22 with an Estimated label, because a single external index, Artificial Analysis, placed it. Its 90 percent interval ran from 78.7 to 100. The label tells you to avoid false precision and prioritize a direct trial. Missing results do not become zeros, so a newly released model can rank without being punished for a different disclosure table.
Compare total cost, not token price
At September 22, 2026 list prices, one million input and 200,000 output tokens cost roughly $20 on Claude Fable 5, $8 on GPT-5.6 Sol, and $3.30 on Gemini 3.5 Flash. Sol's rate is a promotion OpenAI lists through at least November 21. The same $8 bought that volume on Claude Opus 5.5, which led the overall ranking that day. That spread matters. It is still incomplete.
Add retries, human review, failed runs, latency, caching, and the cost of an incorrect answer. If Fable cuts intervention enough, it can beat a cheaper model. If Gemini handles a high-volume extraction task correctly, paying for frontier agent capability is waste.
→ Live pricing · Cost calculator
Run a small evaluation before committing
- Collect 30 to 100 real examples, including failures and edge cases.
- Define acceptance before seeing model names.
- Compare two or three available models within budget.
- Blind the outputs where human preference is involved.
- Measure success rate, intervention, latency, and cost per successful task.
- Repeat after meaningful provider or prompt changes.
The live ranking reduces the search space. Your evaluation makes the decision.
Current defaults
For peak capability, start with Claude Opus 5.5 (86.4, Supported) if you can call it. If you cannot, move down the overall ranking to the first row with public API access. For coding, start with Claude Sonnet 5.5 (85, Supported). For agents, start with Claude Opus 5.5 (88.2, Supported). For budget-sensitive general workloads, start with MiMo-V2.6-Pro (74.1, Estimated). When open-weight control is required, start with MiMo-V2.6-Pro (74.1, Estimated).
On July 14, this section named Claude Mythos 5, Claude Fable 5, GPT-5.6 Sol, Gemini 3.5 Flash, and MiniMax M3. By the September 22 release, none of them still held its slot.
Those are starting points, not permanent endorsements. The models, prices, and evidence will keep changing. The selection process should not.
Frequently asked questions
01What is the best AI model for most people in 2026?
There is no defensible universal default. Start from the live overall ranking, then strike the rows you cannot use: restricted access, a list price above your budget, or an Estimated label on the category that matters to you. Test the two strongest survivors on your own tasks before you commit.
02How do I choose between Claude, GPT, and Gemini?
Start with the failure you cannot afford, then open the ranking that measures it: coding, agentic, or overall. Take the best Claude, GPT, and Gemini rows you can actually call, compare their list prices, and test the two strongest on your own tasks. The leader changes between releases. Your evaluation set should not.
03Should I use an open-weight model or a proprietary API?
Use an open-weight model when privacy, data residency, customization, or serving control outweighs the capability gap. The open-weight ranking shows which model leads today and how far it sits behind the proprietary frontier. Use a proprietary API when peak capability and low operational overhead matter more.
04What does Estimated mean on BenchLM?
Estimated means the model can be placed from the evidence available, but its uncertainty is wider. Supported means enough independent evidence exists for the row to carry normally. Neither label is a separate leaderboard, and missing benchmarks are not converted to zeros.
05Can I switch LLMs later?
Yes, if prompts, tool contracts, evaluation cases, and model-specific fallbacks are kept separate from the provider client. The hidden switching cost is behavioral: prompts and acceptance thresholds often need retuning for a new model.