Which LLM Should I Use?
Start with the job. Get a shortlist with the evidence, tradeoffs, and checks that matter for your choice.
What are you choosing?
How the LLM selector makes its recommendation
The selector covers 31 tasks for everyday users and builders. Each task uses a relevant evidence source, with its limitations stated. Budget, context, and local-deployment requirements narrow the candidates. Scores are never increased because a model is open weight or uses a particular reasoning style.
Voice, image generation, specialist research, and other unsupported tasks receive a clear evidence gap instead of a recommendation built from unrelated scores. Compare current API pricing, inspect visual and document evidence, or read the scoring methodology.
Once you have a shortlist, draft a task prompt or improve an existing prompt to use in your comparison. Running a model you may replace? The alternative finder shows the quality and price delta against it.
Method: bench-align-v5.5-2026-09-04.
Questions
How does the LLM selector work?
Choose a task and the constraints that matter. The selector uses a relevant category estimate or observed task result, filters for the stated budget, context, and deployment needs, and explains the evidence. It does not create a hidden weighted fit score.
Can I describe my project instead of answering the questions?
Yes. The description box turns your words into the same multiple-choice answers the quiz uses, and nothing else. A model classifies the task; it never names or ranks a model. Every suggested answer is shown and editable, and the shortlist is computed from the public dataset exactly as it would be after clicking.
Which LLM should I use for coding?
Choose Build or fix software, then the kind of work you do. Coding-agent results depend on the model, editor or agent harness, and effort setting. Compare the suggested models on representative repository tasks before switching.
What is the best free LLM?
A free assistant plan, an open-weight model, and free hosting are different things. For local deployment, the selector limits candidates to open weights. For assistant plans, model access and message limits must be checked with the provider; API prices are not subscription prices.
Why might the selector return no recommendation?
A task may lack suitable benchmark evidence, or no measured candidate may satisfy your constraints. Missing prices cannot pass a price limit. The selector keeps these gaps visible instead of silently relaxing your requirements.
Can I share or reuse a result?
Every answer is written into the page address, so the link on a results page reopens the same shortlist against the current dataset. The same answers work as query parameters on the /api/tools/recommend endpoint for scripts and agents.