AI Benchmarks Directory
BenchLM tracks 321 AI benchmarks across 10 categories — coding, agentic, reasoning, math, knowledge, multimodal, instruction following, and multilingual — each with a ranked leaderboard updated July 2026. SWE-bench Verified, GPQA Diamond, and Arena Elo are the most-watched evaluations for comparing frontier LLMs.
Explore 321 benchmarks used to evaluate AI language models across 10 categories.
Agentic(64 benchmarks)
View leaderboardDesign Arena Agentic Web Dev
2026Design Arena Agentic Web Dev Elo
A display-only Elo rating from blinded comparisons of multi-file web applications built by coding agents.
Design Arena Agentic Web Dev 2026 · updated July 20, 2026
AA Briefcase
2026Artificial Analysis Briefcase
An independently evaluated professional-work benchmark reported as Elo.
AA Briefcase 2026 · updated July 20, 2026
AA AutomationBench
2026Artificial Analysis AutomationBench
An independently evaluated automation benchmark from Artificial Analysis.
AA AutomationBench 2026 · updated July 20, 2026
AA EnterpriseOps-Gym
2026Artificial Analysis EnterpriseOps-Gym
An independently evaluated enterprise-operations benchmark from Artificial Analysis.
AA EnterpriseOps-Gym 2026 · updated July 20, 2026
AA Harvey LAB
2026Artificial Analysis Harvey LAB-AA
An independently evaluated legal-agent benchmark from Artificial Analysis.
AA Harvey LAB 2026 · updated July 20, 2026
AA ITBench
2026Artificial Analysis ITBench-AA
An independently evaluated IT-operations benchmark from Artificial Analysis.
AA ITBench 2026 · updated July 20, 2026
AA Tau3 Banking
2026Artificial Analysis Tau3-Banking
An independently evaluated Tau3 banking benchmark from Artificial Analysis.
AA Tau3 Banking 2026 · updated July 20, 2026
Terminal-Bench 2.0
2026Terminal-Bench 2.0
A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.
Terminal-Bench 2 · updated July 20, 2026
BrowseComp
2025BrowseComp
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
BrowseComp 2026 · updated July 20, 2026
HLE w/ tools
2026Humanity's Last Exam with tools
Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
HLE w/ tools 2026 · updated July 20, 2026
GDPval-AA
2026GDPval-AA
An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.
GDPval-AA 2026 · updated July 20, 2026
GDPval-AA
2026GDPval-AA normalized
A display-only Artificial Analysis normalized score for economically valuable tasks.
GDPval-AA 2026 · updated July 20, 2026
AA Agentic Index
2026Artificial Analysis Agentic Index
A display-only Artificial Analysis agentic index.
AA Agentic Index 2026 · updated July 20, 2026
APEX-Agents-AA
2026APEX-Agents-AA
Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.
APEX-Agents-AA 2026 · updated July 20, 2026
Gert Labs
2026Gert Labs Composite Game Benchmark
A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.
Gert Labs 2026 · updated July 20, 2026
OSWorld-Verified
2025OSWorld-Verified
OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.
OSWorld Verified · updated July 20, 2026
OSWorld 2.0
2026OSWorld 2.0
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
OSWorld 2.0 2026 · updated July 20, 2026
CyberGym
2026CyberGym
A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
CyberGym 2026 · updated July 20, 2026
Cybench
2025Cybench
A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.
Cybench 2025 · updated July 20, 2026
ExploitGym
2026ExploitGym
A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
ExploitGym 2026 · updated July 20, 2026
JobBench
2026JobBench
An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.
JobBench 2026 · updated July 20, 2026
BrowseComp-VL
2026BrowseComp-VL
A vision-language browsing benchmark for multimodal web research and tool-use workflows.
BrowseComp-VL 2026 · updated July 20, 2026
OSWorld
2026OSWorld
A computer-use benchmark for GUI task completion across the broader OSWorld task suite.
OSWorld 2026 · updated July 20, 2026
AndroidWorld
2026AndroidWorld
A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.
AndroidWorld 2026 · updated July 20, 2026
WebVoyager
2026WebVoyager
A browser-agent benchmark for completing multi-step workflows on live websites.
WebVoyager 2026 · updated July 20, 2026
MCP Atlas
2026MCP Atlas
A benchmark for tool-calling over Model Context Protocol integrations and external tools.
MCP Atlas 2026 · updated July 20, 2026
Kimi Claw 24/7
2026Kimi Claw 24/7 Bench
A Moonshot AI internal long-horizon agent benchmark for persistent professional coworking tasks.
Kimi Claw 24/7 2026 · updated July 20, 2026
MCP Mark Verified
2026MCPMark-Verified
A human-verified edition of MCPMark for MCP tool use across Notion, GitHub, Filesystem, Postgres, and Playwright server environments.
MCP Mark Verified 2026 · updated July 20, 2026
Toolathlon
2026Toolathlon
A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.
Toolathlon 2026 · updated July 20, 2026
Toolathlon-Verified
2026Toolathlon-Verified
A verified tool-use benchmark variant for completing multi-step workflows with external tools.
Toolathlon-Verified 2026 · updated July 20, 2026
AutomationBench
2026AutomationBench
An agent benchmark for completing automation workflows in reproducible task environments.
AutomationBench 2026 · updated July 20, 2026
APEX-Agents
2026APEX-Agents
A professional-services agent benchmark covering long-horizon knowledge-work tasks.
APEX-Agents 2026 · updated July 20, 2026
SpreadsheetBench 2
2026SpreadsheetBench 2
A spreadsheet-focused benchmark for agentic analysis and editing workflows.
SpreadsheetBench 2 2026 · updated July 20, 2026
DECK-Bench
2026DECK-Bench (Internal)
Moonshot AI's internal benchmark for presentation and deck-production workflows.
DECK-Bench 2026 · updated July 20, 2026
ZClawBench
2026ZClawBench
A Z.AI benchmark for OpenClaw-style agent workflows spanning information search, office work, data analysis, development and operations, automation, and security.
ZClawBench 2026 · updated July 20, 2026
τ²-bench results
2025τ²-Bench Tool-Agent-User Evaluation
This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.
τ²-Bench 2026 · updated July 20, 2026
DeepSearchQA
2026DeepSearchQA
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
DeepSearchQA 2026 · updated July 20, 2026
τ²-bench Airline
2025τ²-Bench Airline Domain
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
τ²-bench Airline 2025 · updated July 20, 2026
PinchBench
2026PinchBench
An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.
PinchBench 2026 · updated July 20, 2026
OpenHands Index
2025OpenHands Index
A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
OpenHands Index 2025 · updated July 20, 2026
SWE-Atlas Refactoring
2026SWE-Atlas Refactoring
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
SWE-Atlas Refactoring 2026 · updated July 20, 2026
InferenceBench
2026InferenceBench
A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed.
InferenceBench 2026 · updated July 20, 2026
EdgeBench
2026EdgeBench
A ByteDance Seed benchmark of 134 real-world, day-scale tasks that measures how autonomous agents learn from environment feedback over 12+ hour interaction horizons, spanning scientific and ML, systems and software engineering, optimization, knowledge, formal, and game domains.
EdgeBench 2026 · updated July 20, 2026
BFCL v4
2026Berkeley Function Calling Leaderboard v4
A function-calling benchmark for tool selection, schema adherence, and argument correctness.
BFCL v4 2026 · updated July 20, 2026
MLE-Bench Lite
2026MLE-Bench Lite
A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.
MLE-Bench Lite 2026 · updated July 20, 2026
MM-ClawBench
2026MM-ClawBench
An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.
MM-ClawBench 2026 · updated July 20, 2026
Claw-Eval
2026Claw-Eval
A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.
Claw-Eval 2026 · updated July 20, 2026
ResearchClawBench
2026ResearchClawBench
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
ResearchClawBench 2026 · updated July 20, 2026
QwenClawBench
2026QwenClawBench
Qwen's internal OpenClaw-style benchmark for measuring broad real-world agent performance across practical productivity and research tasks.
QwenClawBench 2026 · updated July 20, 2026
QwenWebBench
2026QwenWebBench
A Qwen benchmark for artifact and webpage generation quality reported as an Elo-style rating.
QwenWebBench 2026 · updated July 20, 2026
τ³-bench results
2026τ³-Bench Tool-Agent-User Evaluation
τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
τ³-bench results 2026 · updated July 20, 2026
VITA-Bench
2025VITA-Bench
An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.
VITA-Bench 2025 · updated July 20, 2026
DeepPlanning
2026DeepPlanning
A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.
DeepPlanning 2026 · updated July 20, 2026
MCP-Tasks
2026MCP-Tasks
A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.
MCP-Tasks 2026 · updated July 20, 2026
WideResearch
2026WideResearch
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
WideResearch 2026 · updated July 20, 2026
GAIA
2024General AI Assistants
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.
GAIA 2024 · updated July 20, 2026
TAU-bench
2024Tool-Agent-User Benchmark
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
TAU-bench 2024 · updated July 20, 2026
WebArena
2024WebArena Web Agent Benchmark
WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.
WebArena 2024 · updated July 20, 2026
WebArena-Verified
2025WebArena-Verified Browser Agent Benchmark
WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.
WebArena-Verified 2025 · updated July 20, 2026
MEWC
2026Multi-Environment Web Challenge
A benchmark that evaluates AI agents on multi-environment web challenges, testing navigation and task completion across diverse live web environments.
MEWC 2026 · updated July 20, 2026
Finance Agent v2
2026Finance Agent v2
Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.
Finance Agent v2 2026 · updated July 20, 2026
Market-Bench
2025Market-Bench
A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.
Market-Bench 2025 · updated July 20, 2026
GDPval rubrics
2026GDPval rubrics
A display-only provider-table GDPval rubric score for economically valuable work tasks.
GDPval rubrics 2026 · updated July 20, 2026
BankerToolBench
2026BankerToolBench
A display-only provider benchmark for finance-oriented tool-use and agent workflows.
BankerToolBench 2026 · updated July 20, 2026
Coding(47 benchmarks)
View leaderboardAA LiveCodeBench
2026Artificial Analysis LiveCodeBench
An independently evaluated LiveCodeBench result from Artificial Analysis.
AA LiveCodeBench 2026 · updated July 20, 2026
AA Terminal-Bench 2.1
2026Artificial Analysis Terminal-Bench v2.1
An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.
AA Terminal-Bench 2.1 2026 · updated July 20, 2026
HumanEval
2021Evaluating Large Language Models Trained on Code
A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.
HumanEval · updated July 20, 2026
BigCodeBench
2026BigCodeBench
A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.
BigCodeBench 2026 · updated July 20, 2026
Codeforces
2026Codeforces Rating
Competitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.
Codeforces 2026 · updated July 20, 2026
Terminal-Bench 2.0
2026Terminal-Bench 2.0
A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.
Terminal-Bench 2 · updated July 20, 2026
SWE-bench Verified
2024Software Engineering Benchmark Verified
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
SWE-bench Verified 2024 · updated July 20, 2026
SWE-Rebench
2026SWE-Rebench
A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.
Rolling 2026 window · updated July 20, 2026
LiveCodeBench
2024LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.
Rolling 2026 set · updated July 20, 2026
LiveCodeBench v6
2026LiveCodeBench v6
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
LiveCodeBench v6 2026 · updated July 20, 2026
LiveCodeBench v5
2025LiveCodeBench v5
LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside the rolling weighted lane.
LiveCodeBench v5 2025 · updated July 20, 2026
LiveCodeBench Pass@1-COT
2026LiveCodeBench Pass@1 with Chain-of-Thought
This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.
LiveCodeBench Pass@1-COT 2026 · updated July 20, 2026
LiveCodeBench Pro
2025LiveCodeBench Pro
A harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with quarter-specific public leaderboards and difficulty-aware reporting.
LiveCodeBench Pro 2025 · updated July 20, 2026
FLTEval
2026FLTEval
A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.
FLTEval 2026 · updated July 20, 2026
SWE-bench Pro
2025SWE-bench Pro
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
SWE-bench Pro 2025 · updated July 20, 2026
Senior SWE-Bench
2026Senior SWE-Bench
A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.
Senior SWE-Bench v2026.06 · updated July 20, 2026
FrontierCode 1.1 Main
2026FrontierCode 1.1 Main
Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.
FrontierCode 1.1 Main · updated July 20, 2026
FrontierCode 1.1 Extended
2026FrontierCode 1.1 Extended
Cognition's 150-task Extended subset of the FrontierCode 1.1 software-engineering benchmark.
FrontierCode 1.1 Extended · updated July 20, 2026
IDE-Bench
2026IDE-Bench
An 80-task software-engineering benchmark across eight repositories that tests whether autonomous IDE agents can explore, edit, run, and verify code changes end to end.
IDE-Bench 2026 · updated July 20, 2026
App-Bench
2025App-Bench
A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.
App-Bench 2025 · updated July 20, 2026
SWE Multilingual
2026SWE Multilingual
A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.
SWE Multilingual 2026 · updated July 20, 2026
SWE Multimodal
2025SWE-bench Multimodal
A multimodal variant of SWE-bench that adds visual context such as screenshots and design mockups to software engineering issue descriptions.
SWE Multimodal 2025 · updated July 20, 2026
CursorBench
2026CursorBench
Cursor's current first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions.
CursorBench 2026 · updated July 20, 2026
Multi-SWE Bench
2026Multi-SWE Bench
A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.
Multi-SWE Bench 2026 · updated July 20, 2026
VIBE-Pro
2026VIBE-Pro
A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.
VIBE-Pro 2026 · updated July 20, 2026
Vibe Code Bench
2026Vibe Code Bench v1.1
Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.
Vibe Code Bench 2026 · updated July 20, 2026
ProgramBench
2026ProgramBench: Can Language Models Rebuild Programs From Scratch?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.
ProgramBench 2026 · updated July 20, 2026
PostTrain Bench
2026PostTrain Bench
A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.
PostTrain Bench 2026 · updated July 20, 2026
FrontierSWE
2026FrontierSWE
An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.
FrontierSWE 2026 · updated July 20, 2026
Kimi Code Bench v2
2026Kimi Code Bench v2
A Moonshot AI internal coding-agent benchmark for realistic software-engineering tasks across mainstream programming languages and production technology stacks.
Kimi Code Bench v2 2026 · updated July 20, 2026
MLS-Bench Lite
2026MLS-Bench Lite
A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods.
MLS-Bench Lite 2026 · updated July 20, 2026
NL2Repo
2026NL2Repo
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
NL2Repo 2026 · updated July 20, 2026
React Native Evals
2026React Native Evals
An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.
React Native Evals 2026 · updated July 20, 2026
ReactBench
2026ReactBench v1
A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.
ReactBench 2026 · updated July 20, 2026
KernelBench
2026KernelBench Hard H100
An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.
KernelBench 2026 · updated July 20, 2026
Next.js Evals
2026AI Agent Evaluations for Next.js
A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.
Next.js Evals 2026 · updated July 20, 2026
SWE-bench Verified*
2026SWE-bench Verified (mini-swe-agent-v2)
A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.
SWE-bench Verified* 2026 · updated July 20, 2026
Spider 2.0-Lite
2024Spider 2.0-Lite
A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison.
Spider 2.0-Lite 2024 · updated July 20, 2026
SciCode
2024Scientific Code Benchmark
SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.
SciCode 2024 · updated July 20, 2026
AA Coding Index
2026Artificial Analysis Coding Index
A display-only Artificial Analysis coding index.
AA Coding Index 2026 · updated July 20, 2026
AA Coding Agents
2026Artificial Analysis Coding Agent Index
A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.
AA Coding Agents 2026 · updated July 20, 2026
AA-SciCode
2026Artificial Analysis SciCode
A display-only Artificial Analysis SciCode score.
AA-SciCode 2026 · updated July 20, 2026
Terminal-Bench Hard
2026Terminal-Bench Hard
A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.
Terminal-Bench Hard 2026 · updated July 20, 2026
VIBE V2
2026VIBE V2
A display-only MiniMax provider benchmark for end-to-end coding-agent and product-building tasks.
VIBE V2 2026 · updated July 20, 2026
SVG-Bench
2026SVG-Bench
A display-only provider benchmark for generating or manipulating SVG outputs from natural-language requirements.
SVG-Bench 2026 · updated July 20, 2026
KernelBench Hard
2026KernelBench Hard
A display-only benchmark for difficult GPU kernel implementation and optimization tasks.
KernelBench Hard 2026 · updated July 20, 2026
EdgeBench
2026EdgeBench
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
EdgeBench 2026 · updated July 20, 2026
Reasoning(25 benchmarks)
View leaderboardMuSR
2023Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.
MuSR 2023 · updated July 20, 2026
BBH
2022BIG-Bench Hard
A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.
BBH 2022 · updated July 20, 2026
DROP
2026Discrete Reasoning Over Paragraphs
A reading-comprehension benchmark requiring discrete reasoning over paragraphs, reported in DeepSeek-V4 base-model evaluations.
DROP 2026 · updated July 20, 2026
HellaSwag
2026HellaSwag
A commonsense natural-language inference benchmark reported in DeepSeek-V4 base-model evaluations.
HellaSwag 2026 · updated July 20, 2026
WinoGrande
2026WinoGrande
A commonsense coreference benchmark reported in DeepSeek-V4 base-model evaluations.
WinoGrande 2026 · updated July 20, 2026
CLUEWSC
2026CLUEWSC
A Chinese Winograd Schema Challenge benchmark reported in DeepSeek-V4 base-model evaluations.
CLUEWSC 2026 · updated July 20, 2026
LisanBench
2026LisanBench
A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.
LisanBench 2026 · updated July 20, 2026
Pencil Puzzle Bench
2026Pencil Puzzle Bench
A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.
Pencil Puzzle Bench 2026 · updated July 20, 2026
LongBench v2
2025LongBench v2
A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.
LongBench v2 2025 · updated July 20, 2026
MRCRv2
2025MRCRv2
A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.
MRCRv2 2025 · updated July 20, 2026
MRCR v2 64K-128K
2026OpenAI MRCR v2 8-needle 64K-128K
MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.
MRCR v2 64K-128K 2026 · updated July 20, 2026
MRCR v2 128K-256K
2026OpenAI MRCR v2 8-needle 128K-256K
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
MRCR v2 128K-256K 2026 · updated July 20, 2026
Graphwalks BFS 128K
2026Graphwalks BFS 0K-128K
Long-context graph traversal benchmark using breadth-first search tasks.
Graphwalks BFS 128K 2026 · updated July 20, 2026
Graphwalks Parents 128K
2026Graphwalks parents 0-128K
Long-context benchmark for recovering parent relationships inside graph tasks.
Graphwalks Parents 128K 2026 · updated July 20, 2026
MRCR 1M
2026MRCR 1M
A million-token MRCR long-context retrieval benchmark reported in DeepSeek-V4 model evaluations.
MRCR 1M 2026 · updated July 20, 2026
CorpusQA 1M
2026CorpusQA 1M
A million-token CorpusQA long-context question-answering benchmark reported in DeepSeek-V4 model evaluations.
CorpusQA 1M 2026 · updated July 20, 2026
ARC-AGI-2
2025Abstraction and Reasoning Corpus for AGI v2
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
ARC-AGI 2 · updated July 20, 2026
ARC-AGI-3
2026Abstraction and Reasoning Corpus for AGI v3
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
ARC-AGI 3 · updated July 20, 2026
GeneBench-Pro
2026GeneBench-Pro
A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.
GeneBench-Pro · updated July 20, 2026
AI-Needle
2026AI-Needle
A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.
AI-Needle 2026 · updated July 20, 2026
GPQA Diamond
2023GPQA Diamond
The hardest subset of GPQA featuring the most challenging graduate-level science questions. Sometimes reported separately from the standard GPQA benchmark.
GPQA Diamond 2023 · updated July 20, 2026
AA-LCR
2026Artificial Analysis Long Context Reasoning
A display-only Artificial Analysis long-context reasoning evaluation.
AA-LCR 2026 · updated July 20, 2026
CritPt
2026Critical Physics Tasks
A display-only Artificial Analysis metric for research-level physics reasoning.
CritPt 2026 · updated July 20, 2026
BullshitBench v2
2025BullshitBench v2
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.
BullshitBench v2 2025 · updated July 20, 2026
WildBench
2024WildBench
An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.
WildBench 2024 · updated July 20, 2026
Multimodal & Grounded(57 benchmarks)
View leaderboardMMMU
2024Massive Multi-discipline Multimodal Understanding
A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.
MMMU 2024 · updated July 20, 2026
MMMU-Pro
2024Massive Multi-discipline Multimodal Understanding Pro
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
MMMU-Pro 2024 · updated July 20, 2026
AA-MMMU-Pro
2026Artificial Analysis MMMU-Pro
A display-only Artificial Analysis MMMU-Pro score.
AA-MMMU-Pro 2026 · updated July 20, 2026
OCRBench V2
2025OCRBench V2
A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.
OCRBench V2 2025 · updated July 20, 2026
olmOCR
2025olmOCR-Bench
An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.
olmOCR 2025 · updated July 20, 2026
VoxPopuli WER
2026VoxPopuli-Cleaned-AA Word Error Rate
A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.
VoxPopuli WER 2026 · updated July 20, 2026
Design Arena Website
2026Design Arena Website Elo
A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.
Design Arena Website 2026 · updated July 20, 2026
OfficeQA Pro
2026OfficeQA Pro
A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.
OfficeQA Pro 2026 · updated July 20, 2026
MathVision w/ Python
2026MathVision with Python
A tool-augmented MathVision variant that permits Python during visual mathematics reasoning.
MathVision w/ Python 2026 · updated July 20, 2026
BabyVision w/ Python
2026BabyVision with Python
A Python-assisted BabyVision evaluation for fine-grained visual perception and grounded reasoning.
BabyVision w/ Python 2026 · updated July 20, 2026
ZeroBench w/ Python
2026ZeroBench_main with Python
A Python-assisted ZeroBench_main evaluation reported as pass@5.
ZeroBench w/ Python 2026 · updated July 20, 2026
WorldVQA ForceAnswer
2026WorldVQA ForceAnswer
A forced-answer WorldVQA variant for atomic visual world knowledge.
WorldVQA ForceAnswer 2026 · updated July 20, 2026
OmniDocBench
2026OmniDocBench
A document-understanding benchmark for parsing and reasoning over complex document layouts.
OmniDocBench 2026 · updated July 20, 2026
PerceptionBench
2026PerceptionBench (Internal)
Moonshot AI's internal benchmark for atomic visual perception capabilities.
PerceptionBench 2026 · updated July 20, 2026
MMMU-Pro w/ Python
2026MMMU-Pro with Python
Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning.
MMMU-Pro w/ Python 2026 · updated July 20, 2026
OmniDocBench 1.5
2026OmniDocBench 1.5
A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.
OmniDocBench 1.5 2026 · updated July 20, 2026
Liquid Extract JSON Validity
2026Liquid image-to-JSON extraction JSON validity
A display-only Liquid AI extraction metric measuring the share of image-to-JSON outputs that parse as strict JSON.
Liquid Extract JSON Validity 2026 · updated July 20, 2026
Liquid Extract F1
2026Liquid image-to-JSON extraction schema consistency F1
A display-only Liquid AI extraction metric measuring field-name agreement between requested schema fields and extracted JSON fields.
Liquid Extract F1 2026 · updated July 20, 2026
Liquid Extract VLM Judge
2026Liquid image-to-JSON extraction VLM judge score
A display-only Liquid AI extraction metric measuring judged agreement between extracted values and the source image.
Liquid Extract VLM Judge 2026 · updated July 20, 2026
RealWorldQA
2026RealWorldQA
A grounded visual QA benchmark focused on answering practical questions about real-world images and scenes.
RealWorldQA 2026 · updated July 20, 2026
Video-MME (with subtitle)
2026Video-MME with subtitle
A video understanding benchmark that allows subtitle access when answering multimodal questions about videos.
Video-MME (with subtitle) 2026 · updated July 20, 2026
Video-MME (w/o subtitle)
2026Video-MME without subtitle
A stricter Video-MME setting that removes subtitle help and tests video understanding from visual and audio context alone.
Video-MME (w/o subtitle) 2026 · updated July 20, 2026
Video-MME
2024Video-MME
A comprehensive benchmark for multimodal large language models on video understanding, covering temporal reasoning, perception, and question answering over videos.
Video-MME 2024 · updated July 20, 2026
MathVision
2026MathVision
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
MathVision 2026 · updated July 20, 2026
We-Math
2026We-Math
A multimodal math benchmark for visually grounded mathematical reasoning and answer generation.
We-Math 2026 · updated July 20, 2026
DynaMath
2026DynaMath
A multimodal benchmark for dynamic mathematical reasoning over visual and structured inputs.
DynaMath 2026 · updated July 20, 2026
MStar
2026MStar
A general visual question-answering benchmark used in provider tables for real-image reasoning quality.
MStar 2026 · updated July 20, 2026
ChatCVQA
2026ChatCVQA
A conversational visual QA benchmark that tests multi-turn grounded answering over images and documents.
ChatCVQA 2026 · updated July 20, 2026
MMLongBench-Doc
2026MMLongBench-Doc
A long-document multimodal benchmark for grounded reasoning over extended document contexts.
MMLongBench-Doc 2026 · updated July 20, 2026
CC-OCR
2026CC-OCR
An OCR-focused benchmark for reading and extracting text from visually complex documents and images.
CC-OCR 2026 · updated July 20, 2026
AI2D_TEST
2026AI2D test split
A diagram understanding benchmark focused on scientific and educational visual question answering.
AI2D_TEST 2026 · updated July 20, 2026
CountBench
2026CountBench
A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.
CountBench 2026 · updated July 20, 2026
RefCOCO (avg)
2026RefCOCO average
A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.
RefCOCO (avg) 2026 · updated July 20, 2026
ODINW13
2026ODINW13
A visual detection and grounding benchmark slice used to compare zero-shot object understanding across diverse domains.
ODINW13 2026 · updated July 20, 2026
ERQA
2026ERQA
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
ERQA 2026 · updated July 20, 2026
VideoMMMU
2026VideoMMMU
A video extension of MMMU-style multimodal reasoning over expert questions grounded in temporal media.
VideoMMMU 2026 · updated July 20, 2026
MLVU (M-Avg)
2026MLVU mean average
A multi-task video understanding benchmark averaged across MLVU categories.
MLVU (M-Avg) 2026 · updated July 20, 2026
MMVU
2026Multimodal Multi-disciplinary Video Understanding
A benchmark for evaluating multimodal models on video understanding tasks across multiple disciplines, emphasizing temporal reasoning and comprehension over video content.
MMVU 2026 · updated July 20, 2026
ScreenSpot Pro
2025ScreenSpot Pro
A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.
ScreenSpot Pro 2025 · updated July 20, 2026
TIR-Bench
2026TIR-Bench
A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.
TIR-Bench 2026 · updated July 20, 2026
GDPval-AA
2026GDPval-AA
An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.
GDPval-AA 2026 · updated July 20, 2026
MedXpertQA (MM)
2026MedXpertQA Multimodal
A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.
MedXpertQA (MM) 2026 · updated July 20, 2026
ZeroBench
2026ZeroBench
A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.
ZeroBench 2026 · updated July 20, 2026
Design2Code
2026Design2Code
A multimodal coding benchmark for turning visual designs into working frontend implementations.
Design2Code 2026 · updated July 20, 2026
Flame-VLM-Code
2026Flame-VLM-Code
A vision-language coding benchmark for generating correct code from visual and multimodal inputs.
Flame-VLM-Code 2026 · updated July 20, 2026
Vision2Web
2026Vision2Web
A benchmark for converting visual references into functional web implementations.
Vision2Web 2026 · updated July 20, 2026
ImageMining
2026ImageMining
A multimodal retrieval and extraction benchmark over image-heavy task settings.
ImageMining 2026 · updated July 20, 2026
MMSearch
2026MMSearch
A multimodal search benchmark for retrieval and grounded answering across mixed-media inputs.
MMSearch 2026 · updated July 20, 2026
MMSearch-Plus
2026MMSearch-Plus
A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.
MMSearch-Plus 2026 · updated July 20, 2026
SimpleVQA
2026SimpleVQA
A visual question answering benchmark focused on straightforward image-grounded understanding.
SimpleVQA 2026 · updated July 20, 2026
Facts-VLM
2026Facts-VLM
A grounded multimodal factuality benchmark for evidence-linked answer correctness.
Facts-VLM 2026 · updated July 20, 2026
V*
2026V*
A vision-centric benchmark for high-level multimodal reasoning and perception quality.
V* 2026 · updated July 20, 2026
CharXiv
2024CharXiv Reasoning
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
CharXiv 2024 · updated July 20, 2026
CharXiv w/o tools
2024CharXiv Reasoning without tools
Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.
CharXiv w/o tools 2024 · updated July 20, 2026
BabyVision
2026BabyVision
A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.
BabyVision 2026 · updated July 20, 2026
SWE-bench Multimodal
2025SWE-bench Multimodal
A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.
SWE-bench Multimodal 2025 · updated July 20, 2026
Blueprint-Bench 2
2026Blueprint-Bench 2
An agentic spatial reasoning benchmark reported as a normalized score.
Blueprint-Bench 2 2026 · updated July 20, 2026
Knowledge(34 benchmarks)
View leaderboardAA Openness Index
2026Artificial Analysis Openness Index
A display-only Artificial Analysis model-openness index.
AA Openness Index 2026 · updated July 20, 2026
AA MMLU-Pro
2026Artificial Analysis MMLU-Pro
An independently evaluated MMLU-Pro result from Artificial Analysis.
AA MMLU-Pro 2026 · updated July 20, 2026
MMLU
2020Massive Multitask Language Understanding
A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.
MMLU · updated July 20, 2026
GPQA
2023Graduate-Level Google-Proof Q&A
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
GPQA Diamond · updated July 20, 2026
GPQA-D
2026GPQA Diamond
A display-only GPQA Diamond reference from provider comparison charts.
GPQA-D 2026 · updated July 20, 2026
SuperGPQA
2025SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.
SuperGPQA 2025 · updated July 20, 2026
MMLU-Pro
2024Massive Multitask Language Understanding Professional
An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.
MMLU-Pro · updated July 20, 2026
AGIEval
2026AGIEval
A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.
AGIEval 2026 · updated July 20, 2026
HLE
2025Humanity's Last Exam
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Humanity's Last Exam · updated July 20, 2026
FrontierScience
2026FrontierScience
A benchmark for research-level scientific reasoning, designed to separate frontier models on difficult science tasks that mix domain knowledge with deep reasoning.
FrontierScience 2026 · updated July 20, 2026
Artificial Analysis Intelligence Index
2026Artificial Analysis Intelligence Index
A display-only intelligence index published by Artificial Analysis that aggregates provider-reported and benchmark-derived signals into a single model-level score.
Artificial Analysis Intelligence Index 2026 · updated July 20, 2026
AA-GPQA Diamond
2026Artificial Analysis GPQA Diamond
A display-only Artificial Analysis GPQA Diamond score.
AA-GPQA Diamond 2026 · updated July 20, 2026
AA-HLE
2026Artificial Analysis Humanity's Last Exam
A display-only Artificial Analysis Humanity's Last Exam score.
AA-HLE 2026 · updated July 20, 2026
AA-Omniscience Index
2026Artificial Analysis Omniscience Index
A display-only Artificial Analysis factual knowledge index.
AA-Omniscience Index 2026 · updated July 20, 2026
AA-Omniscience Accuracy
2026Artificial Analysis Omniscience Accuracy
A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.
AA-Omniscience Accuracy 2026 · updated July 20, 2026
AA-Omniscience Hallucination Rate
2026Artificial Analysis Omniscience Hallucination Rate
A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.
AA-Omniscience Hallucination Rate 2026 · updated July 20, 2026
SimpleQA
2024Measuring Short-Form Factuality in Large Language Models
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
SimpleQA 2024 · updated July 20, 2026
Chinese-SimpleQA
2026Chinese-SimpleQA
A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.
Chinese-SimpleQA 2026 · updated July 20, 2026
OpenBookQA
2018OpenBookQA
A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.
OpenBookQA 2018 · updated July 20, 2026
HealthBench Hard
2026HealthBench Hard
A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.
HealthBench Hard 2026 · updated July 20, 2026
HealthBench Professional
2026HealthBench Professional
An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.
HealthBench Professional 2026 · updated July 20, 2026
MedXpertQA (Text)
2026MedXpertQA Text
A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.
MedXpertQA (Text) 2026 · updated July 20, 2026
FrontierScience Research
2026FrontierScience Research
A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.
FrontierScience Research 2026 · updated July 20, 2026
TruthfulQA
2021TruthfulQA
A benchmark designed to measure whether language models produce truthful answers instead of repeating common misconceptions or misleading falsehoods.
TruthfulQA 2021 · updated July 20, 2026
HLE w/o tools
2026Humanity's Last Exam without tools
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
HLE w/o tools 2026 · updated July 20, 2026
MMLU-Pro (Arcee)
2026MMLU-Pro first-party comparison snapshot
A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.
MMLU-Pro (Arcee) 2026 · updated July 20, 2026
MMLU-Redux
2026MMLU-Redux
A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.
MMLU-Redux 2026 · updated July 20, 2026
MMMLU
2026MMMLU
A multilingual MMLU-style benchmark reported in provider evaluation tables.
MMMLU 2026 · updated July 20, 2026
C-Eval
2023C-Eval
A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.
C-Eval 2023 · updated July 20, 2026
CMMLU
2026Chinese Massive Multitask Language Understanding
A Chinese multitask academic benchmark reported in DeepSeek-V4 base-model evaluations.
CMMLU 2026 · updated July 20, 2026
MultiLoKo
2026MultiLoKo
A multilingual/localized knowledge benchmark reported in DeepSeek-V4 base-model evaluations.
MultiLoKo 2026 · updated July 20, 2026
FACTS Parametric
2026FACTS Parametric
A parametric factuality benchmark reported in DeepSeek-V4 base-model evaluations.
FACTS Parametric 2026 · updated July 20, 2026
TriviaQA
2026TriviaQA
A reading and trivia question-answering benchmark reported in DeepSeek-V4 base-model evaluations.
TriviaQA 2026 · updated July 20, 2026
FinanceArena
2025FinanceArena — FinanceQA Assumption-Based
An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.
FinanceArena 2025 · updated July 20, 2026
Multilingual(11 benchmarks)
View leaderboardAA Global-MMLU-Lite
2026Artificial Analysis Global-MMLU-Lite
An independently evaluated multilingual knowledge result from Artificial Analysis.
AA Global-MMLU-Lite 2026 · updated July 20, 2026
MGSM
2022Multilingual Grade School Math
A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.
MGSM 2022 · updated July 20, 2026
MMLU-ProX
2025MMLU-ProX
A multilingual extension of professional-level academic evaluation across many languages.
MMLU-ProX 2025 · updated July 20, 2026
NOVA-63
2026NOVA-63
A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.
NOVA-63 2026 · updated July 20, 2026
INCLUDE
2026INCLUDE
A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.
INCLUDE 2026 · updated July 20, 2026
PolyMath
2026PolyMath
A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.
PolyMath 2026 · updated July 20, 2026
VWT2k-lite
2026VWT2k-lite
A lighter multilingual benchmark slice published in provider tables for broad cross-lingual transfer and understanding.
VWT2k-lite 2026 · updated July 20, 2026
MAXIFE
2026MAXIFE
A multilingual instruction-following and understanding benchmark row published in Qwen's launch comparisons.
MAXIFE 2026 · updated July 20, 2026
SWE Multilingual
2025SWE-bench Multilingual
A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.
SWE Multilingual 2025 · updated July 20, 2026
NanoBEIR Multilingual
2026NanoBEIR Multilingual Extended
A display-only multilingual retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using NDCG@10 across 11 languages.
NanoBEIR Multilingual 2026 · updated July 20, 2026
MKQA-11
2026MKQA-11 multilingual retrieval
A display-only multilingual QA retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using Recall@20 across 11 languages.
MKQA-11 2026 · updated July 20, 2026
Instruction Following(4 benchmarks)
View leaderboardIFEval
2023Instruction-Following Eval
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
IFEval 2023 · updated July 20, 2026
IFBench
2025Instruction Following Benchmark
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
IFBench 2025 · updated July 20, 2026
AA-IFBench
2026Artificial Analysis IFBench
A display-only Artificial Analysis IFBench score.
AA-IFBench 2026 · updated July 20, 2026
SOB Value Acc
2026Structured Output Benchmark Value Accuracy
A structured-output benchmark from Interfaze measuring whether extracted JSON leaf values exactly match verified ground truth.
SOB Value Acc 2026 · updated July 20, 2026
Mathematics(27 benchmarks)
View leaderboardAA AIME 2025
2026Artificial Analysis AIME 2025
An independently evaluated AIME 2025 result from Artificial Analysis.
AA AIME 2025 2026 · updated July 20, 2026
AA MATH-500
2026Artificial Analysis MATH-500
An independently evaluated MATH-500 result from Artificial Analysis.
AA MATH-500 2026 · updated July 20, 2026
AIME 2023
2023American Invitational Mathematics Examination 2023
A 15-question, 3-hour examination where each answer is an integer from 000 to 999. Serves as the intermediate step between AMC 10/12 and the USA Mathematical Olympiad (USAMO).
AIME 2023 2023 · updated July 20, 2026
AIME 2024
2024American Invitational Mathematics Examination 2024
The 2024 edition of AIME, maintaining the same format of 15 challenging mathematics problems with integer answers from 000 to 999.
AIME 2024 2024 · updated July 20, 2026
AIME 2025
2025American Invitational Mathematics Examination 2025
The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.
AIME 2025 · updated July 20, 2026
GSM8K
2026Grade School Math 8K
A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
GSM8K 2026 · updated July 20, 2026
MATH
2026MATH
A competition-style mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
MATH 2026 · updated July 20, 2026
CMath
2026CMath
A Chinese mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
CMath 2026 · updated July 20, 2026
AIME25 (Arcee)
2026AIME25 first-party comparison snapshot
A display-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart.
AIME25 (Arcee) 2026 · updated July 20, 2026
HMMT Feb 2023
2023Harvard-MIT Mathematics Tournament February 2023
A prestigious high school mathematics competition hosted jointly by Harvard and MIT, featuring challenging problems across various mathematical disciplines.
HMMT Feb 2023 2023 · updated July 20, 2026
HMMT Feb 2024
2024Harvard-MIT Mathematics Tournament February 2024
The 2024 February edition of the Harvard-MIT Mathematics Tournament, continuing the tradition of challenging high school mathematics competition.
HMMT Feb 2024 2024 · updated July 20, 2026
HMMT Feb 2025
2025Harvard-MIT Mathematics Tournament February 2025
The most recent February edition of the Harvard-MIT Mathematics Tournament, featuring the latest challenging problems in competitive mathematics.
HMMT Feb 2025 2025 · updated July 20, 2026
BRUMO 2025
2025Bulgarian Mathematical Olympiad 2025
A challenging mathematical olympiad competition featuring problems that test advanced mathematical reasoning and problem-solving skills at the olympiad level.
BRUMO 2025 2025 · updated July 20, 2026
MATH-500
2021MATH-500 Problem Set
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
MATH-500 2021 · updated July 20, 2026
AIME26
2026AIME 2026
A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.
AIME26 2026 · updated July 20, 2026
IPhO 2025 (Theory)
2026International Physics Olympiad 2025 (Theory)
The three official theory problems from the 2025 International Physics Olympiad, scored with blinded human evaluation.
IPhO 2025 (Theory) 2026 · updated July 20, 2026
HMMT Feb 2025
2025Harvard-MIT Mathematics Tournament February 2025
A February 2025 HMMT slice used in exact-value provider tables for advanced contest-math reasoning.
HMMT Feb 2025 2025 · updated July 20, 2026
HMMT Nov 2025
2025Harvard-MIT Mathematics Tournament November 2025
A November 2025 HMMT slice for high-end mathematical reasoning comparisons.
HMMT Nov 2025 2025 · updated July 20, 2026
HMMT Feb 2026
2026Harvard-MIT Mathematics Tournament February 2026
A February 2026 HMMT slice used in newer frontier-model math comparisons.
HMMT Feb 2026 2026 · updated July 20, 2026
IMOAnswerBench
2026IMOAnswerBench
A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
IMOAnswerBench 2026 · updated July 20, 2026
Apex
2026Apex
A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Apex 2026 · updated July 20, 2026
Apex Shortlist
2026Apex Shortlist
A shortlist subset of the Apex mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Apex Shortlist 2026 · updated July 20, 2026
MMAnswerBench
2026MMAnswerBench
A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.
MMAnswerBench 2026 · updated July 20, 2026
FrontierMath (legacy)
2024FrontierMath legacy aggregate
Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.
FrontierMath (legacy) 2024 · updated July 20, 2026
FrontierMath v2 (Tiers 1-3)
2026FrontierMath v2 Tiers 1-3
Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.
FrontierMath v2 (Tiers 1-3) 2026 · updated July 20, 2026
FrontierMath v2 (Tier 4)
2026FrontierMath v2 Tier 4
Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.
FrontierMath v2 (Tier 4) 2026 · updated July 20, 2026
USAMO 2026
2026United States of America Mathematical Olympiad 2026
The premier US mathematical olympiad competition, featuring proof-based problems that require deep mathematical insight and rigorous argumentation at the highest competition level.
USAMO 2026 2026 · updated July 20, 2026
korean(8 benchmarks)
View leaderboardKMMLU
2024Korean Massive Multitask Language Understanding
Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.
KMMLU 2024 · updated July 20, 2026
KMMLU-Hard
2025KMMLU-Hard
A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.
KMMLU-Hard 2025 · updated July 20, 2026
KMMLU-Redux
KMMLU-Redux
Cleaned KMMLU from national technical qualification exams, with errors removed, decontaminated, and deduplicated.
KMMLU-Redux · updated July 20, 2026
KMMLU-Pro
KMMLU-Pro
Korean National Professional Licensure exams evaluating professional-grade knowledge.
KMMLU-Pro · updated July 20, 2026
CLIcK
Cultural and Linguistic Intelligence in Korean
Evaluates Korean culture and linguistics.
CLIcK · updated July 20, 2026
KoBALT
Korean Benchmark for Advanced Linguistic Tasks
Evaluates advanced Korean linguistic competence.
KoBALT · updated July 20, 2026
Korean CSAT
College Scholastic Ability Test (수능)
The Korean SAT exam.
Korean CSAT · updated July 20, 2026
HRM8K
HAE-RAE Math 8K
Korean mathematical reasoning (high-school to Olympiad level).
HRM8K · updated July 20, 2026
External benchmark mirrors(44 benchmarks)
View leaderboardKindBench
2026KindBench Psychological Safety Benchmark
A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.
KindBench v0.1.0 · updated July 20, 2026
LiveBench
2024LiveBench
A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.
LiveBench 2024 · updated July 20, 2026
Vals Index
2026Vals Index v1.2
Vals AI composite benchmark across finance and coding tasks, including Finance Agent v2, CorpFin v2, SWE-bench, Terminal-Bench 2.1, and Vibe Code Bench.
Vals Index 2026 · updated July 20, 2026
Vals Multimodal Index
2026Vals Multimodal Index v1.1
Vals AI multimodal composite across finance, coding, education, and mortgage-tax task families.
Vals Multimodal Index 2026 · updated July 20, 2026
CorpFin v2
2026Vals CorpFin v2
Vals AI private benchmark for understanding long-context credit agreements.
CorpFin v2 2026 · updated July 20, 2026
MedCode
2026Vals MedCode
Vals AI healthcare benchmark for whether models can support the medical billing process.
MedCode 2026 · updated July 20, 2026
MedScribe
2026Vals MedScribe
Vals AI healthcare benchmark for whether models can support doctors with administrative work.
MedScribe 2026 · updated July 20, 2026
MortgageTax
2026Vals MortgageTax
Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.
MortgageTax 2026 · updated July 20, 2026
ProofBench
2026Vals ProofBench
Vals AI automated theorem-proving benchmark.
ProofBench 2026 · updated July 20, 2026
LegalBench
2026Vals LegalBench
Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.
LegalBench 2026 · updated July 20, 2026
CaseLaw v2
2026Vals CaseLaw v2
Vals AI private question-answer benchmark over Canadian court cases.
CaseLaw v2 2026 · updated July 20, 2026
DeepSWE
2026DeepSWE
A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.
DeepSWE 2026 · updated July 20, 2026
SWE-Marathon
2026SWE-Marathon
A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.
SWE-Marathon 2026 · updated July 20, 2026
ExploitBench
2026ExploitBench v8-bench
A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.
ExploitBench 2026 · updated July 20, 2026
GBA-Eval
2026GBA-Eval
An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.
GBA-Eval 2026 · updated July 20, 2026
CAIS Text Leaderboard
2025CAIS AI Dashboard Text Capabilities Index
A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.
CAIS Text Leaderboard 2025 · updated July 20, 2026
WeirdML
2026WeirdML v2
A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.
WeirdML 2026 · updated July 20, 2026
ALE-Bench
2026Agents Last Exam
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
ALE-Bench 2026 · updated July 20, 2026
RuneScape-Bench
2026RuneBench / runescape-bench
An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.
RuneScape-Bench 2026 · updated July 20, 2026
Toloka Arena
2026Toloka Arena
An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.
Toloka Arena 2026 · updated July 20, 2026
Vals SWE-bench mirror
2026Vals-hosted SWE-bench mirror
Vals AI hosted SWE-bench view for solving production software engineering tasks.
Vals SWE-bench mirror 2026 · updated July 20, 2026
Vals Terminal-Bench 2.0 mirror
2026Vals-hosted Terminal-Bench 2.0 mirror
Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.
Vals Terminal-Bench 2.0 mirror 2026 · updated July 20, 2026
Vals LiveCodeBench mirror
2026Vals-hosted LiveCodeBench mirror
Vals AI implementation of LiveCodeBench with easy, medium, and hard task splits.
Vals LiveCodeBench mirror 2026 · updated July 20, 2026
Vals GPQA Diamond mirror
2026Vals-hosted GPQA Diamond mirror
Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.
Vals GPQA Diamond mirror 2026 · updated July 20, 2026
Vals MMLU-Pro mirror
2026Vals-hosted MMLU-Pro mirror
Vals AI hosted MMLU-Pro view with subject-level task splits.
Vals MMLU-Pro mirror 2026 · updated July 20, 2026
EMB
2026Vals EMB
Evaluating agents on Excel-based financial modeling tasks
EMB 2026 · updated July 20, 2026
CyberBench
2026Vals CyberBench
Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and stop crashing after the fix?
CyberBench 2026 · updated July 20, 2026
TaxEval v2
2026Vals TaxEval v2
A Vals-created set of questions and responses to tax questions
TaxEval v2 2026 · updated July 20, 2026
Harvey's Legal Agent Benchmark
2026Vals Harvey's Legal Agent Benchmark
Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools
Harvey's Legal Agent Benchmark 2026 · updated July 20, 2026
Terminal-Bench 2.1
2026Vals Terminal-Bench 2.1
State-of-the-art set of difficult terminal-based tasks
Terminal-Bench 2.1 2026 · updated July 20, 2026
Code Migration
2026Vals Code Migration
Can language models reimplement real-world programs in another language?
Code Migration 2026 · updated July 20, 2026
Legal Research Bench
2026Vals Legal Research Bench
Evaluating agents on legal research tasks across diverse areas of US law
Legal Research Bench 2026 · updated July 20, 2026
MedQA
2026Vals MedQA
Evaluating language model bias in medical questions.
MedQA 2026 · updated July 20, 2026
AIME
2026Vals AIME
Challenging national math exam given to top high-school students
AIME 2026 · updated July 20, 2026
MATH 500
2026Vals MATH 500
Academic math benchmark on probability, algebra, and trigonometry
MATH 500 2026 · updated July 20, 2026
MGSM
2026Vals MGSM
A multilingual benchmark for mathematical questions.
MGSM 2026 · updated July 20, 2026
MMMU
2026Vals MMMU
Multimodal Multi-task Benchmark
MMMU 2026 · updated July 20, 2026
SAGE
2026Vals SAGE
Student Assessment with Generative Evaluation
SAGE 2026 · updated July 20, 2026
IOI
2026Vals IOI
Based on the International Olympiad in Informatics
IOI 2026 · updated July 20, 2026
ProgramBench
2026Vals ProgramBench
Can language models rebuild programs from scratch?
ProgramBench 2026 · updated July 20, 2026
SkillsBench
2026Vals SkillsBench
How important are skills for agents?
SkillsBench 2026 · updated July 20, 2026
Agent Poker Bench
2026Vals Agent Poker Bench
Which model can make the most money playing poker?
Agent Poker Bench 2026 · updated July 20, 2026
Public Benefits Bench v1.1
2026Vals Public Benefits Bench v1.1
Can AI help people navigate SNAP benefits?
Public Benefits Bench v1.1 2026 · updated July 20, 2026
Public Benefits Bench v1
2026Vals Public Benefits Bench v1
Can AI help people navigate SNAP benefits?
Public Benefits Bench v1 2026 · updated July 20, 2026