# Which LLM Should I Use? — AI Model Selector

> Answer 5 quick questions and get a personalized LLM recommendation based on benchmark data from BenchLM.ai.

## How It Works

1. **Use case**: Coding, writing, analysis, math, or general purpose
2. **Budget**: Free/self-hosted, budget, mid-range, or unlimited
3. **Context needs**: Standard (<200K), large (200K+), or very large (1M+)
4. **Open-weight preference**: Self-host/fine-tune or API access
5. **Speed priority**: Fast responses, balanced, or quality first

Models are scored using weighted category averages that match your stated priorities across 411 models.

## Quick Recommendations by Use Case

- **Coding**: prioritize models with strong non-generated SWE-bench coverage before relying on LiveCodeBench or HumanEval tie-breakers
- **Agents**: treat Terminal-Bench, BrowseComp, and OSWorld rows cautiously when a model shows limited trusted coverage
- **Docs / Charts / Screenshots**: multimodal-grounded rankings are most trustworthy when MMMU-Pro and OfficeQA rows are source-backed for the model you care about
- **Math**: Most frontier models score 95-99% on competition math
- **Knowledge**: HLE, GPQA, and MMLU-Pro remain the cleanest frontier signals when generated rows have been excluded
- **Budget**: DeepSeek V3, Gemini Flash
- **Self-hosted**: Llama 4 Maverick, DeepSeek R1, Qwen3.5

Related tools: [Alternative Finder](/tools/alternative-finder), [Cost Calculator](/tools/cost-calculator)

Try the interactive tool: https://benchlm.ai/tools/llm-selector

Canonical page: https://benchlm.ai/tools/llm-selector
