# Best AI Models (2026)

> BenchLM best-of rankings grouped by model family, deployment type, context size, value, and task fit.

## Ranking Pages

- [Best Long Context AI Models](/best/long-context) ([md](/md/best/long-context.md)) - AI models ranked on sourced long-context and memory benchmarks including LongBench v2, MRCRv2, AI-Needle, Graphwalks, and MMLongBench-Doc.
- [Best Tool Use & Function Calling Models](/best/tool-use) ([md](/md/best/tool-use.md)) - AI models ranked on sourced tool-use benchmarks including BFCL v4, MCP Atlas, Toolathlon, Tau Bench, and related tool-calling evaluations.
- [Best AI Models for Web Research](/best/web-research) ([md](/md/best/web-research.md)) - AI models ranked on sourced browsing and research benchmarks including BrowseComp, WebArena, and WebVoyager.
- [Best Computer Use AI Models](/best/computer-use) ([md](/md/best/computer-use.md)) - AI models ranked on sourced computer-use and GUI benchmarks including OSWorld-Verified, ScreenSpot Pro, WebArena, WebVoyager, and Vision2Web.
- [Best Document AI Models](/best/document-ai) ([md](/md/best/document-ai.md)) - AI models ranked on sourced document AI and OCR benchmarks including OfficeQA Pro, OmniDocBench 1.5, CC-OCR, and MMLongBench-Doc.
- [Best Image Understanding Models](/best/image-understanding) ([md](/md/best/image-understanding.md)) - AI models ranked on sourced image-understanding benchmarks including MMMU-Pro, RealWorldQA, AI2D, CountBench, RefCOCO, and related grounding evaluations.
- [Best Frontend & App Dev Models](/best/frontend-app-dev) ([md](/md/best/frontend-app-dev.md)) - AI models ranked on sourced frontend and app-development benchmarks including React Native Evals, Design2Code, Vision2Web, and related web-task evaluations.
- [Best Factuality AI Models](/best/factuality) ([md](/md/best/factuality.md)) - AI models ranked on sourced factuality and hallucination-adjacent benchmarks including SimpleQA, HLE without tools, and Facts-VLM.
- [Best Open Source LLMs](/best/open-source) ([md](/md/best/open-source.md)) - Compare open-weight LLMs by benchmark score, license, size, context, quantization, and deployment needs.
- [Best Proprietary LLMs](/best/proprietary) ([md](/md/best/proprietary.md)) - Top proprietary/closed-source AI models ranked by benchmark performance.
- [Best Reasoning AI Models](/best/reasoning-models) ([md](/md/best/reasoning-models.md)) - Top AI models with dedicated reasoning capabilities, ranked by benchmark performance.
- [Best OpenAI Models](/best/openai-models) ([md](/md/best/openai-models.md)) - All OpenAI models ranked by benchmark performance — GPT-5, GPT-4o, o1, o3, and more.
- [Best Anthropic Models](/best/anthropic-models) ([md](/md/best/anthropic-models.md)) - All Anthropic Claude models ranked by benchmark performance.
- [Best Google AI Models](/best/google-models) ([md](/md/best/google-models.md)) - All Google Gemini and Gemma models ranked by benchmark performance.
- [Best Meta AI Models](/best/meta-models) ([md](/md/best/meta-models.md)) - All Meta Llama models ranked by benchmark performance.
- [Best DeepSeek Models](/best/deepseek-models) ([md](/md/best/deepseek-models.md)) - All DeepSeek models ranked by benchmark performance.
- [Best AI Models Overall](/best/overall) ([md](/md/best/overall.md)) - The top AI models ranked by overall benchmark performance across all categories.
- [Best Large Context Window LLMs](/best/large-context-window) ([md](/md/best/large-context-window.md)) - AI models with the largest context windows (200K+ tokens), ranked by benchmark performance.
- [Best Chinese AI Models](/best/chinese-models) ([md](/md/best/chinese-models.md)) - A live ranking of AI models from Chinese labs, using the same current public ranking contract as the overall leaderboard.
- [European AI Models](/best/european-models) ([md](/md/best/european-models.md)) - European AI models from Mistral, H Company, LightOn, and Aleph Alpha — ranked models first, then tracked sparse rows.
- [Best Non-Reasoning LLMs](/best/non-reasoning-models) ([md](/md/best/non-reasoning-models.md)) - Top standard AI models (no chain-of-thought reasoning) ranked by benchmark performance. Faster and cheaper than reasoning models.
- [Best Mistral Models](/best/mistral-models) ([md](/md/best/mistral-models.md)) - All Mistral AI models ranked by benchmark performance — Mistral Large, Mixtral, and more.
- [Best xAI Grok Models](/best/xai-models) ([md](/md/best/xai-models.md)) - All xAI Grok models ranked by benchmark performance.
- [Best Alibaba Qwen Models](/best/alibaba-models) ([md](/md/best/alibaba-models.md)) - All Alibaba Qwen models ranked by benchmark performance.
- [Best Value LLM for Coding](/best/best-value-coding) ([md](/md/best/best-value-coding.md)) - Top AI models ranked by coding benchmark performance per dollar. Find the most cost-effective LLM for coding tasks including SWE-bench, LiveCodeBench, and more.
- [Best Value Agentic AI Model](/best/best-value-agentic) ([md](/md/best/best-value-agentic.md)) - Top AI models ranked by agentic benchmark performance per dollar. Find the most cost-effective model for AI agents, tool use, and multi-step workflows.
- [Best Value LLM for Reasoning](/best/best-value-reasoning) ([md](/md/best/best-value-reasoning.md)) - Top AI models ranked by reasoning benchmark performance per dollar. Cost-adjusted rankings using ARC-AGI-2, LongBench v2, MRCRv2, and MuSR.
- [Best Value LLM for Knowledge](/best/best-value-knowledge) ([md](/md/best/best-value-knowledge.md)) - Top AI models ranked by knowledge benchmark performance per dollar. Cost-adjusted rankings using GPQA, MMLU-Pro, HLE, and more.
- [Best Value LLM for Math](/best/best-value-math) ([md](/md/best/best-value-math.md)) - Top AI models ranked by math benchmark performance per dollar. Cost-adjusted rankings using AIME 2025, BRUMO, and MATH-500.
- [Best Value Multimodal AI Model](/best/best-value-multimodal) ([md](/md/best/best-value-multimodal.md)) - Top AI models ranked by multimodal benchmark performance per dollar. Cost-adjusted rankings using MMMU-Pro and OfficeQA Pro.
- [Best Value LLM Overall](/best/best-value-overall) ([md](/md/best/best-value-overall.md)) - Top AI models ranked by overall benchmark performance per dollar. Find the most cost-effective all-around LLM across all categories.
- [Best LLMs for AI Agents](/best/agents) ([md](/md/best/agents.md)) - AI models ranked for agentic work — tool use, browsing, computer use, and long-horizon task execution.
- [Best LLMs for Writing](/best/writing) ([md](/md/best/writing.md)) - AI models ranked for writing quality — instruction following, tone control, and human preference scores.
- [Best LLMs for Translation](/best/translation) ([md](/md/best/translation.md)) - AI models ranked for translation and multilingual work, from BenchLM's multilingual benchmark category.
- [Best Multimodal LLMs](/best/multimodal) ([md](/md/best/multimodal.md)) - AI models ranked for multimodal understanding — images, documents, charts, and grounded visual reasoning.
- [Best LLMs for Research](/best/research) ([md](/md/best/research.md)) - AI models ranked for research work — hard knowledge, agentic web research, and deep search benchmarks.
- [Best LLMs for Roleplay](/best/roleplay) ([md](/md/best/roleplay.md)) - AI models ranked for roleplay and persona work — instruction adherence, persona consistency, and creative response benchmarks.
- [Best LLMs for Data Analysis](/best/data-analysis) ([md](/md/best/data-analysis.md)) - AI models ranked for data analysis — quantitative reasoning, discrete reasoning over text, and analysis-code generation.

## Core Leaderboards

- [Overall leaderboard](/)
- [Coding leaderboard](/coding)
- [Reasoning leaderboard](/reasoning)
- [Math leaderboard](/math)
- [Knowledge leaderboard](/knowledge)
- [Agentic leaderboard](/agentic)
- [Multimodal and grounded leaderboard](/multimodal-grounded)
- [Instruction-following leaderboard](/instruction-following)
- [Multilingual leaderboard](/multilingual)


Canonical page: https://benchlm.ai/best
