Skip to main content

AI Benchmarks Directory

BenchLM tracks 321 AI benchmarks across 10 categories — coding, agentic, reasoning, math, knowledge, multimodal, instruction following, and multilingual — each with a ranked leaderboard updated July 2026. SWE-bench Verified, GPQA Diamond, and Arena Elo are the most-watched evaluations for comparing frontier LLMs.

Explore 321 benchmarks used to evaluate AI language models across 10 categories.

Agentic(64 benchmarks)

View leaderboard

Design Arena Agentic Web Dev

2026

Design Arena Agentic Web Dev Elo

A display-only Elo rating from blinded comparisons of multi-file web applications built by coding agents.

CurrentDisplay only
Multi-file web application developmentElo from blinded human preferencesAgentic frontend development
Display only

Design Arena Agentic Web Dev 2026 · updated July 20, 2026

AA Briefcase

2026

Artificial Analysis Briefcase

An independently evaluated professional-work benchmark reported as Elo.

CurrentDisplay only
Professional knowledge-work tasksEloProfessional work
Display only

AA Briefcase 2026 · updated July 20, 2026

AA AutomationBench

2026

Artificial Analysis AutomationBench

An independently evaluated automation benchmark from Artificial Analysis.

CurrentDisplay only
Business-process automation tasksTask success rateAgentic automation
Display only

AA AutomationBench 2026 · updated July 20, 2026

AA EnterpriseOps-Gym

2026

Artificial Analysis EnterpriseOps-Gym

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

CurrentDisplay only
Enterprise operations workflowsTask success rateEnterprise agent operations
Display only

AA EnterpriseOps-Gym 2026 · updated July 20, 2026

AA Harvey LAB

2026

Artificial Analysis Harvey LAB-AA

An independently evaluated legal-agent benchmark from Artificial Analysis.

CurrentDisplay only
Legal agent tasksTask success rateProfessional legal work
Display only

AA Harvey LAB 2026 · updated July 20, 2026

AA ITBench

2026

Artificial Analysis ITBench-AA

An independently evaluated IT-operations benchmark from Artificial Analysis.

CurrentDisplay only
IT incident-response tasksTask success rateEnterprise IT operations
Display only

AA ITBench 2026 · updated July 20, 2026

AA Tau3 Banking

2026

Artificial Analysis Tau3-Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

CurrentDisplay only
Banking tool-use workflowsTask success rateAgentic banking workflows
Display only

AA Tau3 Banking 2026 · updated July 20, 2026

Terminal-Bench 2.0

2026

Terminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.

Current
Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering
Weighted 38%

Terminal-Bench 2 · updated July 20, 2026

BrowseComp

2025

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Current
Research questions requiring browsingWeb search and evidence synthesisHard web research
Weighted 28%

BrowseComp 2026 · updated July 20, 2026

HLE w/ tools

2026

Humanity's Last Exam with tools

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

CurrentDisplay only
Expert questions with tool usePass@1Frontier tool-augmented reasoning
Display only

HLE w/ tools 2026 · updated July 20, 2026

GDPval-AA

2026

GDPval-AA

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

CurrentDisplay only
Agentic real-world work tasksEloProfessional agentic workflows
Display only

GDPval-AA 2026 · updated July 20, 2026

GDPval-AA

2026

GDPval-AA normalized

A display-only Artificial Analysis normalized score for economically valuable tasks.

CurrentDisplay only
Economically valuable tasksNormalized scoreProfessional agentic workflows
Display only

GDPval-AA 2026 · updated July 20, 2026

AA Agentic Index

2026

Artificial Analysis Agentic Index

A display-only Artificial Analysis agentic index.

CurrentDisplay only
Cross-benchmark agentic indexAggregated model scoreDisplay-only external reference
Display only

AA Agentic Index 2026 · updated July 20, 2026

APEX-Agents-AA

2026

APEX-Agents-AA

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

CurrentDisplay only
452 professional-services agent tasksPass@1Long-horizon workplace agent tasks
Display only

APEX-Agents-AA 2026 · updated July 20, 2026

Gert Labs

2026

Gert Labs Composite Game Benchmark

A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.

CurrentDisplay only
Novel game environmentsComposite game leaderboardAgentic coding and decision-making
Display only

Gert Labs 2026 · updated July 20, 2026

OSWorld-Verified

2025

OSWorld-Verified

OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

Current
369 real-world computer tasks (361 when eight Google Drive tasks are excluded)Execution-based interactive task successMulti-step desktop and cross-application workflows
Weighted 34%

OSWorld Verified · updated July 20, 2026

OSWorld 2.0

2026

OSWorld 2.0

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

CurrentDisplay only
108 long-horizon computer-use workflowsInteractive computer-use evaluationLong-horizon professional workflows
Display only

OSWorld 2.0 2026 · updated July 20, 2026

CyberGym

2026

CyberGym

A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.

CurrentDisplay only
1,507 vulnerability analysis instancesVulnerability reproduction and PoC generationReal-world cybersecurity
Display only

CyberGym 2026 · updated July 20, 2026

Cybench

2025

Cybench

A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.

CurrentDisplay only
40 professional CTF tasksCybersecurity agent task completionProfessional cybersecurity
Display only

Cybench 2025 · updated July 20, 2026

ExploitGym

2026

ExploitGym

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

CurrentDisplay only
898 exploitation tasksWorking exploit generationAdvanced cybersecurity exploitation
Display only

ExploitGym 2026 · updated July 20, 2026

JobBench

2026

JobBench

An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.

CurrentDisplay only
130 tasks across 35 occupationsAgentic workplace deliverablesProfessional multi-source workflows
Display only

JobBench 2026 · updated July 20, 2026

BrowseComp-VL

2026

BrowseComp-VL

A vision-language browsing benchmark for multimodal web research and tool-use workflows.

CurrentDisplay only
Multimodal browsing tasksVision-language web research evaluationMultimodal browser-agent
Display only

BrowseComp-VL 2026 · updated July 20, 2026

OSWorld

2026

OSWorld

A computer-use benchmark for GUI task completion across the broader OSWorld task suite.

CurrentDisplay only
Computer-use tasksInteractive GUI evaluationBroad computer-use suite
Display only

OSWorld 2026 · updated July 20, 2026

AndroidWorld

2026

AndroidWorld

A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.

CurrentDisplay only
Android app workflowsInteractive mobile-agent evaluationComplex mobile task completion
Display only

AndroidWorld 2026 · updated July 20, 2026

WebVoyager

2026

WebVoyager

A browser-agent benchmark for completing multi-step workflows on live websites.

CurrentDisplay only
Live website workflowsInteractive browser-agent evaluationMulti-step web navigation
Display only

WebVoyager 2026 · updated July 20, 2026

MCP Atlas

2026

MCP Atlas

A benchmark for tool-calling over Model Context Protocol integrations and external tools.

CurrentDisplay only
Tool-integrated agent tasksInteractive tool-calling evaluationAdvanced tool use
Display only

MCP Atlas 2026 · updated July 20, 2026

Kimi Claw 24/7

2026

Kimi Claw 24/7 Bench

A Moonshot AI internal long-horizon agent benchmark for persistent professional coworking tasks.

CurrentDisplay only
17 professional scenarios, 610 evaluation pointsAverage pass rate across repeated OpenClaw runsLong-horizon agentic work
Display only

Kimi Claw 24/7 2026 · updated July 20, 2026

MCP Mark Verified

2026

MCPMark-Verified

A human-verified edition of MCPMark for MCP tool use across Notion, GitHub, Filesystem, Postgres, and Playwright server environments.

CurrentDisplay only
MCP tool-use tasks across five server environmentsInteractive MCP task completionAdvanced tool use
Display only

MCP Mark Verified 2026 · updated July 20, 2026

Toolathlon

2026

Toolathlon

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

CurrentDisplay only
Multi-tool workflowsInteractive tool-calling evaluationAdvanced tool use
Display only

Toolathlon 2026 · updated July 20, 2026

Toolathlon-Verified

2026

Toolathlon-Verified

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

CurrentDisplay only
Verified multi-tool workflowsInteractive tool-use scoreAdvanced tool use
Display only

Toolathlon-Verified 2026 · updated July 20, 2026

AutomationBench

2026

AutomationBench

An agent benchmark for completing automation workflows in reproducible task environments.

CurrentDisplay only
600 public automation tasksAgent task-completion scoreLong-horizon automation
Display only

AutomationBench 2026 · updated July 20, 2026

APEX-Agents

2026

APEX-Agents

A professional-services agent benchmark covering long-horizon knowledge-work tasks.

CurrentDisplay only
Professional-services agent tasksAgent task-completion scoreLong-horizon professional work
Display only

APEX-Agents 2026 · updated July 20, 2026

SpreadsheetBench 2

2026

SpreadsheetBench 2

A spreadsheet-focused benchmark for agentic analysis and editing workflows.

CurrentDisplay only
Spreadsheet analysis and editing tasksAgent task-completion scoreProfessional spreadsheet work
Display only

SpreadsheetBench 2 2026 · updated July 20, 2026

DECK-Bench

2026

DECK-Bench (Internal)

Moonshot AI's internal benchmark for presentation and deck-production workflows.

CurrentDisplay only
Internal presentation workflowsInternal evaluation scoreProfessional presentation creation
Display only

DECK-Bench 2026 · updated July 20, 2026

ZClawBench

2026

ZClawBench

A Z.AI benchmark for OpenClaw-style agent workflows spanning information search, office work, data analysis, development and operations, automation, and security.

CurrentDisplay only
OpenClaw agent workflowsEnd-to-end agent benchmarkBroad productivity and operations workflows
Display only

ZClawBench 2026 · updated July 20, 2026

τ²-bench results

2025

τ²-Bench Tool-Agent-User Evaluation

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

CurrentDisplay only
Airline, retail, and telecom customer-service task setsPublished domain success or pass^k resultsDual-control customer-service workflows
Display only

τ²-Bench 2026 · updated July 20, 2026

DeepSearchQA

2026

DeepSearchQA

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

CurrentDisplay only
Agentic browsing and list-answer questionsSearch / open / find browser-agent evaluationAgentic web research
Display only

DeepSearchQA 2026 · updated July 20, 2026

τ²-bench Airline

2025

τ²-Bench Airline Domain

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

CurrentDisplay only
Airline customer-service tasksDomain success under a published trial policyPolicy-constrained airline support workflows
Display only

τ²-bench Airline 2025 · updated July 20, 2026

PinchBench

2026

PinchBench

An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

CurrentDisplay only
23 OpenClaw agent tasksAverage success rate from official runsLong-horizon agent workflows
Display only

PinchBench 2026 · updated July 20, 2026

OpenHands Index

2025

OpenHands Index

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

CurrentDisplay only
SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIAMacro-average across five coding-agent categoriesReal-world software engineering agent tasks
Display only

OpenHands Index 2025 · updated July 20, 2026

SWE-Atlas Refactoring

2026

SWE-Atlas Refactoring

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

CurrentDisplay only
SWE-Atlas refactoring tasksRefactoring score with confidence intervalsReal-world software-engineering agent tasks
Display only

SWE-Atlas Refactoring 2026 · updated July 20, 2026

InferenceBench

2026

InferenceBench

A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed.

CurrentDisplay only
4 inference-serving optimization scenariosTwo-hour autonomous CLI agent runOpen-ended ML systems engineering
Display only

InferenceBench 2026 · updated July 20, 2026

EdgeBench

2026

EdgeBench

A ByteDance Seed benchmark of 134 real-world, day-scale tasks that measures how autonomous agents learn from environment feedback over 12+ hour interaction horizons, spanning scientific and ML, systems and software engineering, optimization, knowledge, formal, and game domains.

CurrentDisplay only
134 tasks (51 public) across 6 domainsLong-horizon interactive agent evaluationDay-scale expert tasks
Display only

EdgeBench 2026 · updated July 20, 2026

BFCL v4

2026

Berkeley Function Calling Leaderboard v4

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

CurrentDisplay only
Function-calling tasksTool invocation and schema evaluationAdvanced tool use
Display only

BFCL v4 2026 · updated July 20, 2026

MLE-Bench Lite

2026

MLE-Bench Lite

A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.

CurrentDisplay only
Low-resource ML competitionsAutonomous iterative ML optimizationAgentic machine learning
Display only

MLE-Bench Lite 2026 · updated July 20, 2026

MM-ClawBench

2026

MM-ClawBench

An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.

CurrentDisplay only
OpenClaw-style real-world tasksAgent workflow evaluationBroad real-world agentic execution
Display only

MM-ClawBench 2026 · updated July 20, 2026

Claw-Eval

2026

Claw-Eval

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

CurrentDisplay only
300 tasks, 2,159 rubricsEnd-to-end autonomous-agent evaluation with Pass^3 scoringReal-world general, multi-turn, and native multimodal agent execution
Display only

Claw-Eval 2026 · updated July 20, 2026

ResearchClawBench

2026

ResearchClawBench

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

CurrentDisplay only
40 tasks across 10 scientific domainsEnd-to-end autonomous research evaluation with RADS scoringScientific research re-discovery
Display only

ResearchClawBench 2026 · updated July 20, 2026

QwenClawBench

2026

QwenClawBench

Qwen's internal OpenClaw-style benchmark for measuring broad real-world agent performance across practical productivity and research tasks.

CurrentDisplay only
Real-world agent workflowsEnd-to-end agent evaluationBroad real-world agentic execution
Display only

QwenClawBench 2026 · updated July 20, 2026

QwenWebBench

2026

QwenWebBench

A Qwen benchmark for artifact and webpage generation quality reported as an Elo-style rating.

CurrentDisplay only
Web artifacts and interactive deliverablesElo-style artifact benchmarkArtifact generation
Display only

QwenWebBench 2026 · updated July 20, 2026

τ³-bench results

2026

τ³-Bench Tool-Agent-User Evaluation

τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

CurrentDisplay only
Corrected customer-service tasks plus knowledge and voice evaluation modesPublished domain or average success resultsLong-horizon, multimodal, and knowledge-aware tool use
Display only

τ³-bench results 2026 · updated July 20, 2026

VITA-Bench

2025

VITA-Bench

An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.

CurrentDisplay only
Interactive consumer-service agent tasksEnd-to-end interactive agent evaluationLong-horizon real-world workflows
Display only

VITA-Bench 2025 · updated July 20, 2026

DeepPlanning

2026

DeepPlanning

A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

CurrentDisplay only
Travel planning and constrained shoppingLong-horizon planning benchmarkConstrained agent planning
Display only

DeepPlanning 2026 · updated July 20, 2026

MCP-Tasks

2026

MCP-Tasks

A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.

CurrentDisplay only
MCP-integrated tool tasksInteractive tool-use evaluationAdvanced MCP workflows
Display only

MCP-Tasks 2026 · updated July 20, 2026

WideResearch

2026

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

CurrentDisplay only
Open-ended research tasksMulti-source research evaluationBroad research-agent workflows
Display only

WideResearch 2026 · updated July 20, 2026

GAIA

2024

General AI Assistants

GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

RefreshingDisplay only
466
Display only

GAIA 2024 · updated July 20, 2026

TAU-bench

2024

Tool-Agent-User Benchmark

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

RefreshingDisplay only
Airline and retail task sets in the archived 2024 releaseDomain-specific pass^1 through pass^4 task successPolicy-constrained, multi-turn customer service
Display only

TAU-bench 2024 · updated July 20, 2026

WebArena

2024

WebArena Web Agent Benchmark

WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.

RefreshingDisplay only
812 long-horizon browser tasksEnd-state task successStateful multi-site browser work
Display only

WebArena 2024 · updated July 20, 2026

WebArena-Verified

2025

WebArena-Verified Browser Agent Benchmark

WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.

CurrentDisplay only
812 verified tasks; separate 258-task Hard subsetDeterministic end-state task successAudited stateful browser work
Display only

WebArena-Verified 2025 · updated July 20, 2026

MEWC

2026

Multi-Environment Web Challenge

A benchmark that evaluates AI agents on multi-environment web challenges, testing navigation and task completion across diverse live web environments.

CurrentDisplay only
Web-agent tasksBrowser task completionOpen-web agent workflows
Display only

MEWC 2026 · updated July 20, 2026

Finance Agent v2

2026

Finance Agent v2

Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.

CurrentDisplay only
Financial analyst task categoriesMean score across repeated runsProfessional expert-task agent workflow
Display only

Finance Agent v2 2026 · updated July 20, 2026

Market-Bench

2025

Market-Bench

A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.

CurrentDisplay only
3 quantitative-trading strategiesBacktester implementation scored by mean absolute errorMarket simulation and quantitative coding
Display only

Market-Bench 2025 · updated July 20, 2026

GDPval rubrics

2026

GDPval rubrics

A display-only provider-table GDPval rubric score for economically valuable work tasks.

CurrentDisplay only
Economically valuable work tasksRubric scoreProfessional agentic workflows
Display only

GDPval rubrics 2026 · updated July 20, 2026

BankerToolBench

2026

BankerToolBench

A display-only provider benchmark for finance-oriented tool-use and agent workflows.

CurrentDisplay only
Finance and banking tool-use tasksTask success rateProfessional finance-agent workflows
Display only

BankerToolBench 2026 · updated July 20, 2026

Coding(47 benchmarks)

View leaderboard

AA LiveCodeBench

2026

Artificial Analysis LiveCodeBench

An independently evaluated LiveCodeBench result from Artificial Analysis.

CurrentDisplay only
Contamination-resistant coding tasksPass rateCompetitive programming
Display only

AA LiveCodeBench 2026 · updated July 20, 2026

AA Terminal-Bench 2.1

2026

Artificial Analysis Terminal-Bench v2.1

An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.

CurrentDisplay only
Terminal-based agent tasksTask success rateAgentic software engineering
Display only

AA Terminal-Bench 2.1 2026 · updated July 20, 2026

HumanEval

2021

Evaluating Large Language Models Trained on Code

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

StaleSaturatedDisplay only
164 problemsPython function generationIntroductory to intermediate programming
Display only

HumanEval · updated July 20, 2026

BigCodeBench

2026

BigCodeBench

A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Code generation tasksPass@1Software engineering
Display only

BigCodeBench 2026 · updated July 20, 2026

Codeforces

2026

Codeforces Rating

Competitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.

CurrentDisplay only
Competitive programming contestsRatingElite competitive programming
Display only

Codeforces 2026 · updated July 20, 2026

Terminal-Bench 2.0

2026

Terminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.

CurrentDisplay only
Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering
Display only

Terminal-Bench 2 · updated July 20, 2026

SWE-bench Verified

2024

Software Engineering Benchmark Verified

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

Refreshing
500 verified issuesCode patch generationProfessional software engineering
Weighted 16%

SWE-bench Verified 2024 · updated July 20, 2026

SWE-Rebench

2026

SWE-Rebench

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

Current
Fresh GitHub issues (rolling window)Code patch generationProfessional software engineering
Weighted 20%

Rolling 2026 window · updated July 20, 2026

LiveCodeBench

2024

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.

Current
Continuously updated contest problemsCompetitive-programming evaluationCompetitive programming level
Weighted 38%

Rolling 2026 set · updated July 20, 2026

LiveCodeBench v6

2026

LiveCodeBench v6

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

CurrentDisplay only
Fresh programming problemsProvider-published v6 competitive programming resultsCompetitive programming level
Display only

LiveCodeBench v6 2026 · updated July 20, 2026

LiveCodeBench v5

2025

LiveCodeBench v5

LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside the rolling weighted lane.

CurrentDisplay only
July 2024 to May 2025 release windowProvider-published v5 competitive programming resultCompetitive programming level
Display only

LiveCodeBench v5 2025 · updated July 20, 2026

LiveCodeBench Pass@1-COT

2026

LiveCodeBench Pass@1 with Chain-of-Thought

This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.

CurrentDisplay only
DeepSeek-V4 report evaluation windowPass@1-COT competitive programming resultsCompetitive programming level
Display only

LiveCodeBench Pass@1-COT 2026 · updated July 20, 2026

LiveCodeBench Pro

2025

LiveCodeBench Pro

A harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with quarter-specific public leaderboards and difficulty-aware reporting.

CurrentDisplay only
Quarter-specific contest programming setsCompetitive programmingHigh-end contest programming
Display only

LiveCodeBench Pro 2025 · updated July 20, 2026

FLTEval

2026

FLTEval

A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

CurrentDisplay only
FLT project pull requestsLean 4 repository task completionFormal verification / proof engineering
Display only

FLTEval 2026 · updated July 20, 2026

SWE-bench Pro

2025

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

Current
1,865 repository problemsRepository task completionLong-horizon professional engineering
Weighted 10%

SWE-bench Pro 2025 · updated July 20, 2026

Senior SWE-Bench

2026

Senior SWE-Bench

A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

CurrentDisplay only
Senior-level repository tasksAgentic software-engineering evaluationProfessional senior engineering
Display only

Senior SWE-Bench v2026.06 · updated July 20, 2026

FrontierCode 1.1 Main

2026

FrontierCode 1.1 Main

Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

CurrentDisplay only
100 private Main tasks (150 in Extended)Repository task completion with maintainer rubricsFrontier coding-agent quality
Display only

FrontierCode 1.1 Main · updated July 20, 2026

FrontierCode 1.1 Extended

2026

FrontierCode 1.1 Extended

Cognition's 150-task Extended subset of the FrontierCode 1.1 software-engineering benchmark.

CurrentDisplay only
150 private software-engineering tasksRepository task completion with maintainer rubricsFrontier coding-agent quality
Display only

FrontierCode 1.1 Extended · updated July 20, 2026

IDE-Bench

2026

IDE-Bench

An 80-task software-engineering benchmark across eight repositories that tests whether autonomous IDE agents can explore, edit, run, and verify code changes end to end.

CurrentDisplay only
80 tasks across 8 repositoriesAutonomous IDE-agent task completion (pass@1)End-to-end software engineering
Display only

IDE-Bench 2026 · updated July 20, 2026

App-Bench

2025

App-Bench

A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.

CurrentDisplay only
6 full-stack app-building tasksBest-of-three one-shot feature completionProduction-style full-stack application generation
Display only

App-Bench 2025 · updated July 20, 2026

SWE Multilingual

2026

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

CurrentDisplay only
Multilingual software-engineering tasksRepository task completionProfessional software engineering
Display only

SWE Multilingual 2026 · updated July 20, 2026

SWE Multimodal

2025

SWE-bench Multimodal

A multimodal variant of SWE-bench that adds visual context such as screenshots and design mockups to software engineering issue descriptions.

CurrentDisplay only
Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding
Display only

SWE Multimodal 2025 · updated July 20, 2026

CursorBench

2026

CursorBench

Cursor's current first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions.

CurrentDisplay only
Harder long-horizon agentic coding tasksCursor agent-loop evaluationProfessional agentic software engineering
Display only

CursorBench 2026 · updated July 20, 2026

Multi-SWE Bench

2026

Multi-SWE Bench

A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.

CurrentDisplay only
Multi-language repo tasksRepository task completionProfessional software engineering
Display only

Multi-SWE Bench 2026 · updated July 20, 2026

VIBE-Pro

2026

VIBE-Pro

A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.

CurrentDisplay only
Full project delivery tasksRepository-level implementation benchmarkEnd-to-end software delivery
Display only

VIBE-Pro 2026 · updated July 20, 2026

Vibe Code Bench

2026

Vibe Code Bench v1.1

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

CurrentDisplay only
End-to-end web application buildsFull-stack app implementation benchmarkEnd-to-end software delivery
Display only

Vibe Code Bench 2026 · updated July 20, 2026

ProgramBench

2026

ProgramBench: Can Language Models Rebuild Programs From Scratch?

A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.

CurrentDisplay only
200 program reconstruction tasksCleanroom executable reimplementationFull-repository software architecture
Display only

ProgramBench 2026 · updated July 20, 2026

PostTrain Bench

2026

PostTrain Bench

A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.

CurrentDisplay only
Post-training software-engineering tasksHarbor agent evaluationFrontier software engineering
Display only

PostTrain Bench 2026 · updated July 20, 2026

FrontierSWE

2026

FrontierSWE

An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

CurrentDisplay only
17 ultra-long-horizon engineering and research tasksMean@5, best@5, average rank, and dominanceUltra-long-horizon frontier software engineering
Display only

FrontierSWE 2026 · updated July 20, 2026

Kimi Code Bench v2

2026

Kimi Code Bench v2

A Moonshot AI internal coding-agent benchmark for realistic software-engineering tasks across mainstream programming languages and production technology stacks.

CurrentDisplay only
Realistic coding-agent tasksCoding-agent pass rateProduction software engineering
Display only

Kimi Code Bench v2 2026 · updated July 20, 2026

MLS-Bench Lite

2026

MLS-Bench Lite

A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods.

CurrentDisplay only
30 machine-learning research tasksAgentic ML task evaluationML research and systems engineering
Display only

MLS-Bench Lite 2026 · updated July 20, 2026

NL2Repo

2026

NL2Repo

A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

CurrentDisplay only
Natural language to repository tasksRepository understanding benchmarkSystem-level software comprehension
Display only

NL2Repo 2026 · updated July 20, 2026

React Native Evals

2026

React Native Evals

An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.

CurrentDisplay only
React Native app implementation tasksFramework-specific app development evaluationProduction mobile app engineering
Display only

React Native Evals 2026 · updated July 20, 2026

ReactBench

2026

ReactBench v1

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

CurrentDisplay only
51 production React tasksPass@1 weighted rubric scoreProduction frontend engineering
Display only

ReactBench 2026 · updated July 20, 2026

KernelBench

2026

KernelBench Hard H100

An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.

CurrentDisplay only
6 GPU-kernel optimization problemsMean peak fraction of hardware roofline over valid cellsAgentic GPU systems engineering
Display only

KernelBench 2026 · updated July 20, 2026

Next.js Evals

2026

AI Agent Evaluations for Next.js

A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.

CurrentDisplay only
24 Next.js code generation and migration tasksAgent task completion with withheld Vitest assertionsFramework-specific web application engineering
Display only

Next.js Evals 2026 · updated July 20, 2026

SWE-bench Verified*

2026

SWE-bench Verified (mini-swe-agent-v2)

A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.

CurrentDisplay only
Repository task completionAgent scaffold benchmarkProfessional software engineering
Display only

SWE-bench Verified* 2026 · updated July 20, 2026

Spider 2.0-Lite

2024

Spider 2.0-Lite

A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison.

RefreshingDisplay only
Text-to-SQL queriesExecution accuracyEnterprise text-to-SQL
Display only

Spider 2.0-Lite 2024 · updated July 20, 2026

SciCode

2024

Scientific Code Benchmark

SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.

Refreshing
80
Weighted 16%

SciCode 2024 · updated July 20, 2026

AA Coding Index

2026

Artificial Analysis Coding Index

A display-only Artificial Analysis coding index.

CurrentDisplay only
Cross-benchmark coding indexAggregated model scoreDisplay-only external reference
Display only

AA Coding Index 2026 · updated July 20, 2026

AA Coding Agents

2026

Artificial Analysis Coding Agent Index

A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.

CurrentDisplay only
Composite over DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnAAverage pass@1 indexReal-world coding-agent workflows
Display only

AA Coding Agents 2026 · updated July 20, 2026

AA-SciCode

2026

Artificial Analysis SciCode

A display-only Artificial Analysis SciCode score.

CurrentDisplay only
Scientific coding subproblemsTask success rateScientific programming
Display only

AA-SciCode 2026 · updated July 20, 2026

Terminal-Bench Hard

2026

Terminal-Bench Hard

A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

CurrentDisplay only
Agentic coding and terminal tasksTask success rateProfessional software engineering
Display only

Terminal-Bench Hard 2026 · updated July 20, 2026

VIBE V2

2026

VIBE V2

A display-only MiniMax provider benchmark for end-to-end coding-agent and product-building tasks.

CurrentDisplay only
End-to-end coding-agent tasksTask success rateFrontier coding-agent workflows
Display only

VIBE V2 2026 · updated July 20, 2026

SVG-Bench

2026

SVG-Bench

A display-only provider benchmark for generating or manipulating SVG outputs from natural-language requirements.

CurrentDisplay only
SVG generation and editing tasksTask success rateVisual coding and structured graphics generation
Display only

SVG-Bench 2026 · updated July 20, 2026

KernelBench Hard

2026

KernelBench Hard

A display-only benchmark for difficult GPU kernel implementation and optimization tasks.

CurrentDisplay only
Hard GPU kernel coding tasksTask success rateSpecialized systems programming
Display only

KernelBench Hard 2026 · updated July 20, 2026

EdgeBench

2026

EdgeBench

A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

CurrentDisplay only
Systems and software-engineering tasksTime-budgeted agent learning curvesLong-horizon engineering
Display only

EdgeBench 2026 · updated July 20, 2026

Reasoning(25 benchmarks)

View leaderboard

MuSR

2023

Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

StaleDisplay only
Multi-step reasoningNarrative-based reasoningComplex reasoning tasks
Display only

MuSR 2023 · updated July 20, 2026

BBH

2022

BIG-Bench Hard

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

StaleSaturatedDisplay only
23 tasksMixed reasoning tasksAdvanced reasoning
Display only

BBH 2022 · updated July 20, 2026

DROP

2026

Discrete Reasoning Over Paragraphs

A reading-comprehension benchmark requiring discrete reasoning over paragraphs, reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Paragraph reasoning questionsF1Reading and numerical reasoning
Display only

DROP 2026 · updated July 20, 2026

HellaSwag

2026

HellaSwag

A commonsense natural-language inference benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Commonsense completion questionsExact matchCommonsense reasoning
Display only

HellaSwag 2026 · updated July 20, 2026

WinoGrande

2026

WinoGrande

A commonsense coreference benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Coreference resolution questionsExact matchCommonsense reasoning
Display only

WinoGrande 2026 · updated July 20, 2026

CLUEWSC

2026

CLUEWSC

A Chinese Winograd Schema Challenge benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Chinese coreference questionsExact matchChinese commonsense reasoning
Display only

CLUEWSC 2026 · updated July 20, 2026

LisanBench

2026

LisanBench

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

CurrentDisplay only
50 starting words × 3 trialsDifficulty-weighted word-chain reasoningOpen-ended lexical planning
Display only

LisanBench 2026 · updated July 20, 2026

Pencil Puzzle Bench

2026

Pencil Puzzle Bench

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

CurrentDisplay only
300 evaluation puzzlesDirect and agentic puzzle solve rateMulti-step verifiable reasoning
Display only

Pencil Puzzle Bench 2026 · updated July 20, 2026

LongBench v2

2025

LongBench v2

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

Current
Long-context tasksExtended-context retrieval and reasoningHard long-context
Weighted 38%

LongBench v2 2025 · updated July 20, 2026

MRCRv2

2025

MRCRv2

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

Current
Long-context retrievalMulti-round long-context evaluationHard long-context
Weighted 31%

MRCRv2 2025 · updated July 20, 2026

MRCR v2 64K-128K

2026

OpenAI MRCR v2 8-needle 64K-128K

MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.

CurrentDisplay only
8-needle retrieval tasksLong-context retrievalLong-context reasoning
Display only

MRCR v2 64K-128K 2026 · updated July 20, 2026

MRCR v2 128K-256K

2026

OpenAI MRCR v2 8-needle 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

CurrentDisplay only
8-needle retrieval tasksVery-long-context retrievalVery long-context reasoning
Display only

MRCR v2 128K-256K 2026 · updated July 20, 2026

Graphwalks BFS 128K

2026

Graphwalks BFS 0K-128K

Long-context graph traversal benchmark using breadth-first search tasks.

CurrentDisplay only
Graph traversal tasksLong-context graph reasoningAlgorithmic long-context reasoning
Display only

Graphwalks BFS 128K 2026 · updated July 20, 2026

Graphwalks Parents 128K

2026

Graphwalks parents 0-128K

Long-context benchmark for recovering parent relationships inside graph tasks.

CurrentDisplay only
Graph parent-retrieval tasksLong-context graph reasoningAlgorithmic long-context reasoning
Display only

Graphwalks Parents 128K 2026 · updated July 20, 2026

MRCR 1M

2026

MRCR 1M

A million-token MRCR long-context retrieval benchmark reported in DeepSeek-V4 model evaluations.

CurrentDisplay only
Million-token retrievalLong-context retrieval MMRMillion-token long context
Display only

MRCR 1M 2026 · updated July 20, 2026

CorpusQA 1M

2026

CorpusQA 1M

A million-token CorpusQA long-context question-answering benchmark reported in DeepSeek-V4 model evaluations.

CurrentDisplay only
Million-token corpus question answeringLong-context QA accuracyMillion-token long context
Display only

CorpusQA 1M 2026 · updated July 20, 2026

ARC-AGI-2

2025

Abstraction and Reasoning Corpus for AGI v2

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

Current
Visual pattern completion and abstract reasoningGrid transformation puzzles with novel rulesExpert-level — hardest public reasoning benchmark
Weighted 31%

ARC-AGI 2 · updated July 20, 2026

ARC-AGI-3

2026

Abstraction and Reasoning Corpus for AGI v3

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

CurrentDisplay only
Interactive game-like tasks with hidden rulesAgentic task completion under a capped evaluation budgetFrontier agentic reasoning
Display only

ARC-AGI 3 · updated July 20, 2026

GeneBench-Pro

2026

GeneBench-Pro

A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

CurrentDisplay only
129 genomics statistical-analysis workflowsEval-level pass rate across dependent analysis decisionsLong-horizon scientific reasoning
Display only

GeneBench-Pro · updated July 20, 2026

AI-Needle

2026

AI-Needle

A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.

CurrentDisplay only
Long-context retrievalNeedle-in-a-haystack recallLong-context memory
Display only

AI-Needle 2026 · updated July 20, 2026

GPQA Diamond

2023

GPQA Diamond

The hardest subset of GPQA featuring the most challenging graduate-level science questions. Sometimes reported separately from the standard GPQA benchmark.

StaleDisplay only
Expert-level science questionsMultiple choice questionsGraduate-level scientific reasoning
Display only

GPQA Diamond 2023 · updated July 20, 2026

AA-LCR

2026

Artificial Analysis Long Context Reasoning

A display-only Artificial Analysis long-context reasoning evaluation.

CurrentDisplay only
Long-context reasoning tasksAccuracyLong-context reasoning
Display only

AA-LCR 2026 · updated July 20, 2026

CritPt

2026

Critical Physics Tasks

A display-only Artificial Analysis metric for research-level physics reasoning.

CurrentDisplay only
Research-level physics questionsAccuracyResearch-level physics reasoning
Display only

CritPt 2026 · updated July 20, 2026

BullshitBench v2

2025

BullshitBench v2

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

CurrentDisplay only
Nonsensical and flawed prompts across multiple domainsPrompt challenge and refusal evaluationRobustness and critical reasoning
Display only

BullshitBench v2 2025 · updated July 20, 2026

WildBench

2024

WildBench

An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.

RefreshingDisplay only
1,024 real-world tasksReal-world task evaluationDiverse real-world scenarios
Display only

WildBench 2024 · updated July 20, 2026

Multimodal & Grounded(57 benchmarks)

View leaderboard

MMMU

2024

Massive Multi-discipline Multimodal Understanding

A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.

RefreshingDisplay only
Multimodal academic reasoningImage + text question answeringFrontier multimodal
Display only

MMMU 2024 · updated July 20, 2026

MMMU-Pro

2024

Massive Multi-discipline Multimodal Understanding Pro

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Refreshing
Multimodal academic reasoningImage + text question answeringFrontier multimodal
Weighted 45%

MMMU-Pro 2024 · updated July 20, 2026

AA-MMMU-Pro

2026

Artificial Analysis MMMU-Pro

A display-only Artificial Analysis MMMU-Pro score.

CurrentDisplay only
Multimodal academic reasoningImage + text question answeringFrontier multimodal
Display only

AA-MMMU-Pro 2026 · updated July 20, 2026

OCRBench V2

2025

OCRBench V2

A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.

CurrentDisplay only
Image OCR tasksAccuracyNative visual text understanding
Display only

OCRBench V2 2025 · updated July 20, 2026

olmOCR

2025

olmOCR-Bench

An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.

CurrentDisplay only
Layout-rich PDF understandingMean accuracyComplex document processing
Display only

olmOCR 2025 · updated July 20, 2026

VoxPopuli WER

2026

VoxPopuli-Cleaned-AA Word Error Rate

A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.

CurrentDisplay only
Speech-to-text transcriptionWord error rateAudio speech recognition
Display only

VoxPopuli WER 2026 · updated July 20, 2026

Design Arena Website

2026

Design Arena Website Elo

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

CurrentDisplay only
Website generation comparisonsEloDesign and website generation
Display only

Design Arena Website 2026 · updated July 20, 2026

OfficeQA Pro

2026

OfficeQA Pro

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

Current
Document and spreadsheet tasksGrounded QA over office artifactsEnterprise grounded reasoning
Weighted 30%

OfficeQA Pro 2026 · updated July 20, 2026

MathVision w/ Python

2026

MathVision with Python

A tool-augmented MathVision variant that permits Python during visual mathematics reasoning.

CurrentDisplay only
Visual mathematics problems with PythonImage and mathematics reasoning with toolsAdvanced multimodal mathematics
Display only

MathVision w/ Python 2026 · updated July 20, 2026

BabyVision w/ Python

2026

BabyVision with Python

A Python-assisted BabyVision evaluation for fine-grained visual perception and grounded reasoning.

CurrentDisplay only
Visual perception tasks with PythonTool-augmented multimodal scoreFine-grained visual perception
Display only

BabyVision w/ Python 2026 · updated July 20, 2026

ZeroBench w/ Python

2026

ZeroBench_main with Python

A Python-assisted ZeroBench_main evaluation reported as pass@5.

CurrentDisplay only
Visual reasoning questions with PythonPass@5Tool-augmented visual reasoning
Display only

ZeroBench w/ Python 2026 · updated July 20, 2026

WorldVQA ForceAnswer

2026

WorldVQA ForceAnswer

A forced-answer WorldVQA variant for atomic visual world knowledge.

CurrentDisplay only
Atomic visual world-knowledge questionsForced-answer visual QAFine-grained visual knowledge
Display only

WorldVQA ForceAnswer 2026 · updated July 20, 2026

OmniDocBench

2026

OmniDocBench

A document-understanding benchmark for parsing and reasoning over complex document layouts.

CurrentDisplay only
Complex document-understanding tasksDocument-understanding scoreGrounded document reasoning
Display only

OmniDocBench 2026 · updated July 20, 2026

PerceptionBench

2026

PerceptionBench (Internal)

Moonshot AI's internal benchmark for atomic visual perception capabilities.

CurrentDisplay only
Internal atomic visual-perception tasksInternal evaluation scoreFine-grained visual perception
Display only

PerceptionBench 2026 · updated July 20, 2026

MMMU-Pro w/ Python

2026

MMMU-Pro with Python

Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning.

CurrentDisplay only
Multimodal academic reasoningImage + text question answering with PythonFrontier multimodal
Display only

MMMU-Pro w/ Python 2026 · updated July 20, 2026

OmniDocBench 1.5

2026

OmniDocBench 1.5

A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.

CurrentDisplay only
Document understanding tasksDocument understanding benchmarkGrounded document reasoning
Display only

OmniDocBench 1.5 2026 · updated July 20, 2026

Liquid Extract JSON Validity

2026

Liquid image-to-JSON extraction JSON validity

A display-only Liquid AI extraction metric measuring the share of image-to-JSON outputs that parse as strict JSON.

CurrentDisplay only
Image-to-JSON extractionStrict JSON parseability rateStructured visual extraction
Display only

Liquid Extract JSON Validity 2026 · updated July 20, 2026

Liquid Extract F1

2026

Liquid image-to-JSON extraction schema consistency F1

A display-only Liquid AI extraction metric measuring field-name agreement between requested schema fields and extracted JSON fields.

CurrentDisplay only
Image-to-JSON extractionSchema field F1Structured visual extraction
Display only

Liquid Extract F1 2026 · updated July 20, 2026

Liquid Extract VLM Judge

2026

Liquid image-to-JSON extraction VLM judge score

A display-only Liquid AI extraction metric measuring judged agreement between extracted values and the source image.

CurrentDisplay only
Image-to-JSON extractionVLM-judged extraction accuracyStructured visual extraction
Display only

Liquid Extract VLM Judge 2026 · updated July 20, 2026

RealWorldQA

2026

RealWorldQA

A grounded visual QA benchmark focused on answering practical questions about real-world images and scenes.

CurrentDisplay only
Real-world visual question answeringImage-grounded QAGeneral visual reasoning
Display only

RealWorldQA 2026 · updated July 20, 2026

Video-MME (with subtitle)

2026

Video-MME with subtitle

A video understanding benchmark that allows subtitle access when answering multimodal questions about videos.

CurrentDisplay only
Video understandingVideo QA with subtitle contextMultimodal video reasoning
Display only

Video-MME (with subtitle) 2026 · updated July 20, 2026

Video-MME (w/o subtitle)

2026

Video-MME without subtitle

A stricter Video-MME setting that removes subtitle help and tests video understanding from visual and audio context alone.

CurrentDisplay only
Video understandingVideo QA without subtitle contextMultimodal video reasoning
Display only

Video-MME (w/o subtitle) 2026 · updated July 20, 2026

Video-MME

2024

Video-MME

A comprehensive benchmark for multimodal large language models on video understanding, covering temporal reasoning, perception, and question answering over videos.

RefreshingDisplay only
Video understandingVideo QA and analysisBroad multimodal video reasoning
Display only

Video-MME 2024 · updated July 20, 2026

MathVision

2026

MathVision

A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.

CurrentDisplay only
Visually grounded math problemsImage + math reasoningAdvanced multimodal mathematics
Display only

MathVision 2026 · updated July 20, 2026

We-Math

2026

We-Math

A multimodal math benchmark for visually grounded mathematical reasoning and answer generation.

CurrentDisplay only
Visually grounded math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics
Display only

We-Math 2026 · updated July 20, 2026

DynaMath

2026

DynaMath

A multimodal benchmark for dynamic mathematical reasoning over visual and structured inputs.

CurrentDisplay only
Dynamic visual math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics
Display only

DynaMath 2026 · updated July 20, 2026

MStar

2026

MStar

A general visual question-answering benchmark used in provider tables for real-image reasoning quality.

CurrentDisplay only
Real-image visual QAImage-grounded QAGeneral visual reasoning
Display only

MStar 2026 · updated July 20, 2026

ChatCVQA

2026

ChatCVQA

A conversational visual QA benchmark that tests multi-turn grounded answering over images and documents.

CurrentDisplay only
Conversational visual QAMulti-turn image-grounded QAConversational multimodal reasoning
Display only

ChatCVQA 2026 · updated July 20, 2026

MMLongBench-Doc

2026

MMLongBench-Doc

A long-document multimodal benchmark for grounded reasoning over extended document contexts.

CurrentDisplay only
Long document understandingDocument-grounded reasoningLong-context document reasoning
Display only

MMLongBench-Doc 2026 · updated July 20, 2026

CC-OCR

2026

CC-OCR

An OCR-focused benchmark for reading and extracting text from visually complex documents and images.

CurrentDisplay only
Optical character recognitionText extraction from images and documentsDocument reading
Display only

CC-OCR 2026 · updated July 20, 2026

AI2D_TEST

2026

AI2D test split

A diagram understanding benchmark focused on scientific and educational visual question answering.

CurrentDisplay only
Diagram understandingDiagram-grounded QAStructured visual reasoning
Display only

AI2D_TEST 2026 · updated July 20, 2026

CountBench

2026

CountBench

A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.

CurrentDisplay only
Visual counting tasksImage-grounded countingFine-grained visual perception
Display only

CountBench 2026 · updated July 20, 2026

RefCOCO (avg)

2026

RefCOCO average

A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.

CurrentDisplay only
Referring-expression groundingGrounded visual localizationFine-grained visual grounding
Display only

RefCOCO (avg) 2026 · updated July 20, 2026

ODINW13

2026

ODINW13

A visual detection and grounding benchmark slice used to compare zero-shot object understanding across diverse domains.

CurrentDisplay only
Out-of-distribution object understandingDetection and groundingRobust visual grounding
Display only

ODINW13 2026 · updated July 20, 2026

ERQA

2026

ERQA

A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

CurrentDisplay only
Evidence-based visual QAGrounded image reasoningGrounded multimodal reasoning
Display only

ERQA 2026 · updated July 20, 2026

VideoMMMU

2026

VideoMMMU

A video extension of MMMU-style multimodal reasoning over expert questions grounded in temporal media.

CurrentDisplay only
Video-grounded expert reasoningVideo + text reasoningFrontier multimodal video reasoning
Display only

VideoMMMU 2026 · updated July 20, 2026

MLVU (M-Avg)

2026

MLVU mean average

A multi-task video understanding benchmark averaged across MLVU categories.

CurrentDisplay only
General video understandingVideo QA and understandingBroad multimodal video reasoning
Display only

MLVU (M-Avg) 2026 · updated July 20, 2026

MMVU

2026

Multimodal Multi-disciplinary Video Understanding

A benchmark for evaluating multimodal models on video understanding tasks across multiple disciplines, emphasizing temporal reasoning and comprehension over video content.

CurrentDisplay only
Video understandingVideo reasoning benchmarkMulti-disciplinary multimodal video reasoning
Display only

MMVU 2026 · updated July 20, 2026

ScreenSpot Pro

2025

ScreenSpot Pro

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

CurrentDisplay only
1,581 grounding instructionsStatic interface element localizationProfessional GUI grounding
Display only

ScreenSpot Pro 2025 · updated July 20, 2026

TIR-Bench

2026

TIR-Bench

A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.

CurrentDisplay only
Visual agent and interface reasoningScreenshot-grounded task reasoningComputer-use visual reasoning
Display only

TIR-Bench 2026 · updated July 20, 2026

GDPval-AA

2026

GDPval-AA

An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.

CurrentDisplay only
Professional office deliveryELO-style office benchmarkProfessional knowledge work
Display only

GDPval-AA 2026 · updated July 20, 2026

MedXpertQA (MM)

2026

MedXpertQA Multimodal

A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

CurrentDisplay only
2,000 multimodal medical questionsMedical visual MCQClinical multimodal reasoning
Display only

MedXpertQA (MM) 2026 · updated July 20, 2026

ZeroBench

2026

ZeroBench

A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.

CurrentDisplay only
100 visual reasoning questionsMulti-step visual reasoningTool-augmented visual reasoning
Display only

ZeroBench 2026 · updated July 20, 2026

Design2Code

2026

Design2Code

A multimodal coding benchmark for turning visual designs into working frontend implementations.

CurrentDisplay only
Design-to-code tasksVisual input to frontend implementationMultimodal coding
Display only

Design2Code 2026 · updated July 20, 2026

Flame-VLM-Code

2026

Flame-VLM-Code

A vision-language coding benchmark for generating correct code from visual and multimodal inputs.

CurrentDisplay only
Multimodal coding tasksVision-language code generationMultimodal coding
Display only

Flame-VLM-Code 2026 · updated July 20, 2026

Vision2Web

2026

Vision2Web

A benchmark for converting visual references into functional web implementations.

CurrentDisplay only
Screenshot-to-web tasksVisual reference to web implementationMultimodal web generation
Display only

Vision2Web 2026 · updated July 20, 2026

ImageMining

2026

ImageMining

A multimodal retrieval and extraction benchmark over image-heavy task settings.

CurrentDisplay only
Visual retrieval tasksImage-grounded retrieval and extractionMultimodal retrieval
Display only

ImageMining 2026 · updated July 20, 2026

MMSearch

2026

MMSearch

A multimodal search benchmark for retrieval and grounded answering across mixed-media inputs.

CurrentDisplay only
Multimodal search tasksMixed-media retrieval and grounded answeringMultimodal search
Display only

MMSearch 2026 · updated July 20, 2026

MMSearch-Plus

2026

MMSearch-Plus

A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.

CurrentDisplay only
Hard multimodal search tasksAdvanced mixed-media retrieval benchmarkAdvanced multimodal search
Display only

MMSearch-Plus 2026 · updated July 20, 2026

SimpleVQA

2026

SimpleVQA

A visual question answering benchmark focused on straightforward image-grounded understanding.

CurrentDisplay only
Visual QA tasksImage-grounded question answeringGeneral visual understanding
Display only

SimpleVQA 2026 · updated July 20, 2026

Facts-VLM

2026

Facts-VLM

A grounded multimodal factuality benchmark for evidence-linked answer correctness.

CurrentDisplay only
Grounded factuality tasksEvidence-linked multimodal factualityGrounded multimodal factuality
Display only

Facts-VLM 2026 · updated July 20, 2026

V*

2026

V*

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

CurrentDisplay only
Frontier multimodal reasoning tasksVision-centric reasoning benchmarkFrontier multimodal
Display only

V* 2026 · updated July 20, 2026

CharXiv

2024

CharXiv Reasoning

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Refreshing
Scientific chart reasoningChart understanding and reasoningScientific visualization reasoning
Weighted 25%

CharXiv 2024 · updated July 20, 2026

CharXiv w/o tools

2024

CharXiv Reasoning without tools

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

RefreshingDisplay only
Scientific chart reasoning (tool-free)Chart understanding without toolsScientific visualization reasoning
Display only

CharXiv w/o tools 2024 · updated July 20, 2026

BabyVision

2026

BabyVision

A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

CurrentDisplay only
Visual perception tasksMultimodal visual reasoningFine-grained visual perception
Display only

BabyVision 2026 · updated July 20, 2026

SWE-bench Multimodal

2025

SWE-bench Multimodal

A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.

CurrentDisplay only
Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding
Display only

SWE-bench Multimodal 2025 · updated July 20, 2026

Blueprint-Bench 2

2026

Blueprint-Bench 2

An agentic spatial reasoning benchmark reported as a normalized score.

CurrentDisplay only
Spatial reasoning from blueprintsNormalized scoreAgentic spatial reasoning
Display only

Blueprint-Bench 2 2026 · updated July 20, 2026

Knowledge(34 benchmarks)

View leaderboard

AA Openness Index

2026

Artificial Analysis Openness Index

A display-only Artificial Analysis model-openness index.

CurrentDisplay only
Model openness assessmentIndex scoreDisplay-only external reference
Display only

AA Openness Index 2026 · updated July 20, 2026

AA MMLU-Pro

2026

Artificial Analysis MMLU-Pro

An independently evaluated MMLU-Pro result from Artificial Analysis.

CurrentDisplay only
Professional multi-subject questionsAccuracyProfessional knowledge and reasoning
Display only

AA MMLU-Pro 2026 · updated July 20, 2026

MMLU

2020

Massive Multitask Language Understanding

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

StaleSaturatedDisplay only
57 subjectsMultiple choice questionsElementary to professional level
Display only

MMLU · updated July 20, 2026

GPQA

2023

Graduate-Level Google-Proof Q&A

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

Refreshing
448 questionsMultiple choice questionsGraduate level
Weighted 7%

GPQA Diamond · updated July 20, 2026

GPQA-D

2026

GPQA Diamond

A display-only GPQA Diamond reference from provider comparison charts.

CurrentDisplay only
Graduate-level science questionsMultiple choice questionsGraduate level
Display only

GPQA-D 2026 · updated July 20, 2026

SuperGPQA

2025

SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines

An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.

Current
285 disciplinesMultiple choice questionsGraduate level
Weighted 7%

SuperGPQA 2025 · updated July 20, 2026

MMLU-Pro

2024

Massive Multitask Language Understanding Professional

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Refreshing
Multiple subjects10-way multiple choiceProfessional level
Weighted 30%

MMLU-Pro · updated July 20, 2026

AGIEval

2026

AGIEval

A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
General academic and professional exam questionsExact matchGeneral knowledge
Display only

AGIEval 2026 · updated July 20, 2026

HLE

2025

Humanity's Last Exam

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Current
Expert-level questionsOpen-ended and multiple choiceFrontier expert level
Weighted 45%

Humanity's Last Exam · updated July 20, 2026

FrontierScience

2026

FrontierScience

A benchmark for research-level scientific reasoning, designed to separate frontier models on difficult science tasks that mix domain knowledge with deep reasoning.

CurrentDisplay only
Research-level science tasksScientific reasoning benchmarkResearch frontier
Display only

FrontierScience 2026 · updated July 20, 2026

Artificial Analysis Intelligence Index

2026

Artificial Analysis Intelligence Index

A display-only intelligence index published by Artificial Analysis that aggregates provider-reported and benchmark-derived signals into a single model-level score.

CurrentDisplay only
Cross-benchmark intelligence indexAggregated model scoreDisplay-only external reference
Display only

Artificial Analysis Intelligence Index 2026 · updated July 20, 2026

AA-GPQA Diamond

2026

Artificial Analysis GPQA Diamond

A display-only Artificial Analysis GPQA Diamond score.

CurrentDisplay only
Graduate-level science questionsAccuracyGraduate-level science reasoning
Display only

AA-GPQA Diamond 2026 · updated July 20, 2026

AA-HLE

2026

Artificial Analysis Humanity's Last Exam

A display-only Artificial Analysis Humanity's Last Exam score.

CurrentDisplay only
Expert-level questionsAccuracyFrontier expert reasoning
Display only

AA-HLE 2026 · updated July 20, 2026

AA-Omniscience Index

2026

Artificial Analysis Omniscience Index

A display-only Artificial Analysis factual knowledge index.

CurrentDisplay only
Knowledge questionsIndex scoreBroad factual knowledge
Display only

AA-Omniscience Index 2026 · updated July 20, 2026

AA-Omniscience Accuracy

2026

Artificial Analysis Omniscience Accuracy

A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.

CurrentDisplay only
Knowledge questionsAccuracyBroad knowledge
Display only

AA-Omniscience Accuracy 2026 · updated July 20, 2026

AA-Omniscience Hallucination Rate

2026

Artificial Analysis Omniscience Hallucination Rate

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

CurrentDisplay only
Knowledge questionsHallucination rateFactuality
Display only

AA-Omniscience Hallucination Rate 2026 · updated July 20, 2026

SimpleQA

2024

Measuring Short-Form Factuality in Large Language Models

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Refreshing
Factual questionsShort-form Q&AFactual accuracy focused
Weighted 11%

SimpleQA 2024 · updated July 20, 2026

Chinese-SimpleQA

2026

Chinese-SimpleQA

A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.

CurrentDisplay only
Chinese factual questionsShort-form factual QAFactual accuracy focused
Display only

Chinese-SimpleQA 2026 · updated July 20, 2026

OpenBookQA

2018

OpenBookQA

A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.

StaleDisplay only
Elementary science questions4-way multiple choiceElementary science reasoning
Display only

OpenBookQA 2018 · updated July 20, 2026

HealthBench Hard

2026

HealthBench Hard

A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

CurrentDisplay only
1,000 health promptsOpen-ended health evaluationAdvanced health reasoning
Display only

HealthBench Hard 2026 · updated July 20, 2026

HealthBench Professional

2026

HealthBench Professional

An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.

CurrentDisplay only
Clinician chat tasksRubric-graded open-ended responsesProfessional clinical workflows
Display only

HealthBench Professional 2026 · updated July 20, 2026

MedXpertQA (Text)

2026

MedXpertQA Text

A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.

CurrentDisplay only
2,450 medical multiple-choice questionsMedical MCQProfessional medical knowledge
Display only

MedXpertQA (Text) 2026 · updated July 20, 2026

FrontierScience Research

2026

FrontierScience Research

A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.

CurrentDisplay only
Scientific research problemsResearch evaluationFrontier scientific research
Display only

FrontierScience Research 2026 · updated July 20, 2026

TruthfulQA

2021

TruthfulQA

A benchmark designed to measure whether language models produce truthful answers instead of repeating common misconceptions or misleading falsehoods.

StaleDisplay only
Truthfulness and misconception resistanceQuestion answeringHallucination and factuality stress test
Display only

TruthfulQA 2021 · updated July 20, 2026

HLE w/o tools

2026

Humanity's Last Exam without tools

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

CurrentDisplay only
Expert-level questionsTool-free expert QAFrontier expert level
Display only

HLE w/o tools 2026 · updated July 20, 2026

MMLU-Pro (Arcee)

2026

MMLU-Pro first-party comparison snapshot

A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.

CurrentDisplay only
Professional academic QA10-way multiple choiceProfessional level
Display only

MMLU-Pro (Arcee) 2026 · updated July 20, 2026

MMLU-Redux

2026

MMLU-Redux

A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.

CurrentDisplay only
Broad academic QAMultiple choice questionsAdvanced general knowledge
Display only

MMLU-Redux 2026 · updated July 20, 2026

MMMLU

2026

MMMLU

A multilingual MMLU-style benchmark reported in provider evaluation tables.

CurrentDisplay only
Multilingual academic QAExact matchBroad multilingual knowledge
Display only

MMMLU 2026 · updated July 20, 2026

C-Eval

2023

C-Eval

A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.

StaleDisplay only
Chinese academic and professional examsMultiple choice questionsHigh school to professional level
Display only

C-Eval 2023 · updated July 20, 2026

CMMLU

2026

Chinese Massive Multitask Language Understanding

A Chinese multitask academic benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Chinese academic QAExact matchBroad Chinese knowledge
Display only

CMMLU 2026 · updated July 20, 2026

MultiLoKo

2026

MultiLoKo

A multilingual/localized knowledge benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Localized multilingual knowledge questionsExact matchMultilingual knowledge
Display only

MultiLoKo 2026 · updated July 20, 2026

FACTS Parametric

2026

FACTS Parametric

A parametric factuality benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Parametric factual recallExact matchFactual accuracy focused
Display only

FACTS Parametric 2026 · updated July 20, 2026

TriviaQA

2026

TriviaQA

A reading and trivia question-answering benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Trivia and reading-comprehension QAExact matchGeneral factual QA
Display only

TriviaQA 2026 · updated July 20, 2026

FinanceArena

2025

FinanceArena — FinanceQA Assumption-Based

An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.

CurrentDisplay only
Professional financial-analysis questionsOpen-ended financial QA with exact-match gradingProfessional finance reasoning
Display only

FinanceArena 2025 · updated July 20, 2026

Multilingual(11 benchmarks)

View leaderboard

AA Global-MMLU-Lite

2026

Artificial Analysis Global-MMLU-Lite

An independently evaluated multilingual knowledge result from Artificial Analysis.

CurrentDisplay only
Multilingual knowledge questionsAccuracyMultilingual professional knowledge
Display only

AA Global-MMLU-Lite 2026 · updated July 20, 2026

MGSM

2022

Multilingual Grade School Math

A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.

StaleDisplay only
250 problems × 11 languagesMath word problemsGrade school math, multilingual
Display only

MGSM 2022 · updated July 20, 2026

MMLU-ProX

2025

MMLU-ProX

A multilingual extension of professional-level academic evaluation across many languages.

Current
Multilingual professional QAMultilingual multiple choiceProfessional multilingual
Weighted 100%

MMLU-ProX 2025 · updated July 20, 2026

NOVA-63

2026

NOVA-63

A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.

CurrentDisplay only
Broad multilingual evaluationCross-lingual benchmarkBroad multilingual capability
Display only

NOVA-63 2026 · updated July 20, 2026

INCLUDE

2026

INCLUDE

A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.

CurrentDisplay only
Cross-lingual understandingMultilingual benchmarkBroad multilingual capability
Display only

INCLUDE 2026 · updated July 20, 2026

PolyMath

2026

PolyMath

A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

CurrentDisplay only
Multilingual math problemsCross-lingual mathematical reasoningAdvanced multilingual reasoning
Display only

PolyMath 2026 · updated July 20, 2026

VWT2k-lite

2026

VWT2k-lite

A lighter multilingual benchmark slice published in provider tables for broad cross-lingual transfer and understanding.

CurrentDisplay only
Multilingual transfer tasksCross-lingual benchmarkBroad multilingual capability
Display only

VWT2k-lite 2026 · updated July 20, 2026

MAXIFE

2026

MAXIFE

A multilingual instruction-following and understanding benchmark row published in Qwen's launch comparisons.

CurrentDisplay only
Multilingual instruction followingCross-lingual benchmarkAdvanced multilingual instruction following
Display only

MAXIFE 2026 · updated July 20, 2026

SWE Multilingual

2025

SWE-bench Multilingual

A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.

CurrentDisplay only
300 problems across 9 languagesMulti-language code patch generationProfessional multilingual software engineering
Display only

SWE Multilingual 2025 · updated July 20, 2026

NanoBEIR Multilingual

2026

NanoBEIR Multilingual Extended

A display-only multilingual retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using NDCG@10 across 11 languages.

CurrentDisplay only
Multilingual document retrievalNDCG@10 averageMultilingual retrieval
Display only

NanoBEIR Multilingual 2026 · updated July 20, 2026

MKQA-11

2026

MKQA-11 multilingual retrieval

A display-only multilingual QA retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using Recall@20 across 11 languages.

CurrentDisplay only
Cross-lingual open-domain QA retrievalRecall@20 averageMultilingual retrieval
Display only

MKQA-11 2026 · updated July 20, 2026

Instruction Following(4 benchmarks)

View leaderboard

Mathematics(27 benchmarks)

View leaderboard

AA AIME 2025

2026

Artificial Analysis AIME 2025

An independently evaluated AIME 2025 result from Artificial Analysis.

CurrentDisplay only
30 AIME 2025 problemsAccuracyOlympiad mathematics
Display only

AA AIME 2025 2026 · updated July 20, 2026

AA MATH-500

2026

Artificial Analysis MATH-500

An independently evaluated MATH-500 result from Artificial Analysis.

CurrentDisplay only
500 competition mathematics problemsAccuracyHigh school to undergraduate mathematics
Display only

AA MATH-500 2026 · updated July 20, 2026

AIME 2023

2023

American Invitational Mathematics Examination 2023

A 15-question, 3-hour examination where each answer is an integer from 000 to 999. Serves as the intermediate step between AMC 10/12 and the USA Mathematical Olympiad (USAMO).

StaleDisplay only
15 problemsInteger answers 000-999High school olympiad level
Display only

AIME 2023 2023 · updated July 20, 2026

AIME 2024

2024

American Invitational Mathematics Examination 2024

The 2024 edition of AIME, maintaining the same format of 15 challenging mathematics problems with integer answers from 000 to 999.

RefreshingDisplay only
15 problemsInteger answers 000-999High school olympiad level
Display only

AIME 2024 2024 · updated July 20, 2026

AIME 2025

2025

American Invitational Mathematics Examination 2025

The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.

CurrentDisplay only
15 problemsInteger answers 000-999High school olympiad level
Display only

AIME 2025 · updated July 20, 2026

GSM8K

2026

Grade School Math 8K

A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Grade-school math word problemsExact matchGrade-school math
Display only

GSM8K 2026 · updated July 20, 2026

MATH

2026

MATH

A competition-style mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Competition math problemsExact matchAdvanced math reasoning
Display only

MATH 2026 · updated July 20, 2026

CMath

2026

CMath

A Chinese mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

CurrentDisplay only
Chinese math problemsExact matchMath reasoning
Display only

CMath 2026 · updated July 20, 2026

AIME25 (Arcee)

2026

AIME25 first-party comparison snapshot

A display-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart.

CurrentDisplay only
15 problemsInteger answers 000-999High school olympiad level
Display only

AIME25 (Arcee) 2026 · updated July 20, 2026

HMMT Feb 2023

2023

Harvard-MIT Mathematics Tournament February 2023

A prestigious high school mathematics competition hosted jointly by Harvard and MIT, featuring challenging problems across various mathematical disciplines.

StaleDisplay only
Tournament problemsCompetition mathematicsHigh school olympiad level
Display only

HMMT Feb 2023 2023 · updated July 20, 2026

HMMT Feb 2024

2024

Harvard-MIT Mathematics Tournament February 2024

The 2024 February edition of the Harvard-MIT Mathematics Tournament, continuing the tradition of challenging high school mathematics competition.

RefreshingDisplay only
Tournament problemsCompetition mathematicsHigh school olympiad level
Display only

HMMT Feb 2024 2024 · updated July 20, 2026

HMMT Feb 2025

2025

Harvard-MIT Mathematics Tournament February 2025

The most recent February edition of the Harvard-MIT Mathematics Tournament, featuring the latest challenging problems in competitive mathematics.

CurrentDisplay only
Tournament problemsCompetition mathematicsHigh school olympiad level
Display only

HMMT Feb 2025 2025 · updated July 20, 2026

BRUMO 2025

2025

Bulgarian Mathematical Olympiad 2025

A challenging mathematical olympiad competition featuring problems that test advanced mathematical reasoning and problem-solving skills at the olympiad level.

CurrentDisplay only
Olympiad problemsMathematical olympiadMathematical olympiad level
Display only

BRUMO 2025 2025 · updated July 20, 2026

MATH-500

2021

MATH-500 Problem Set

A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

StaleDisplay only
500 problemsFree-form mathematical answersHigh school to undergraduate
Display only

MATH-500 2021 · updated July 20, 2026

AIME26

2026

AIME 2026

A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.

Current
Competition math problemsShort-answer mathematicsOlympiad-style mathematics
Weighted 25%

AIME26 2026 · updated July 20, 2026

IPhO 2025 (Theory)

2026

International Physics Olympiad 2025 (Theory)

The three official theory problems from the 2025 International Physics Olympiad, scored with blinded human evaluation.

CurrentDisplay only
3 olympiad theory problemsPhysics olympiad theoryInternational olympiad physics
Display only

IPhO 2025 (Theory) 2026 · updated July 20, 2026

HMMT Feb 2025

2025

Harvard-MIT Mathematics Tournament February 2025

A February 2025 HMMT slice used in exact-value provider tables for advanced contest-math reasoning.

CurrentDisplay only
Competition math problemsContest mathematicsOlympiad-style mathematics
Display only

HMMT Feb 2025 2025 · updated July 20, 2026

HMMT Nov 2025

2025

Harvard-MIT Mathematics Tournament November 2025

A November 2025 HMMT slice for high-end mathematical reasoning comparisons.

CurrentDisplay only
Competition math problemsContest mathematicsOlympiad-style mathematics
Display only

HMMT Nov 2025 2025 · updated July 20, 2026

HMMT Feb 2026

2026

Harvard-MIT Mathematics Tournament February 2026

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Current
Competition math problemsContest mathematicsOlympiad-style mathematics
Weighted 25%

HMMT Feb 2026 2026 · updated July 20, 2026

IMOAnswerBench

2026

IMOAnswerBench

A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

CurrentDisplay only
Advanced mathematical answer generationPass@1 math benchmarkOlympiad-level mathematics
Display only

IMOAnswerBench 2026 · updated July 20, 2026

Apex

2026

Apex

A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

CurrentDisplay only
Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning
Display only

Apex 2026 · updated July 20, 2026

Apex Shortlist

2026

Apex Shortlist

A shortlist subset of the Apex mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

CurrentDisplay only
Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning
Display only

Apex Shortlist 2026 · updated July 20, 2026

MMAnswerBench

2026

MMAnswerBench

A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.

CurrentDisplay only
Multimodal math questionsVisual and structured mathematical QAAdvanced mathematical reasoning
Display only

MMAnswerBench 2026 · updated July 20, 2026

FrontierMath (legacy)

2024

FrontierMath legacy aggregate

Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.

RefreshingDisplay only
Historical aggregateOpen-ended mathematical reasoning with tool accessResearch-level mathematics
Display only

FrontierMath (legacy) 2024 · updated July 20, 2026

FrontierMath v2 (Tiers 1-3)

2026

FrontierMath v2 Tiers 1-3

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

Current
295 private advanced mathematics problemsPython-enabled iterative mathematical problem solvingFrom olympiad-plus to early research mathematics
Weighted 30%

FrontierMath v2 (Tiers 1-3) 2026 · updated July 20, 2026

FrontierMath v2 (Tier 4)

2026

FrontierMath v2 Tier 4

Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.

Current
43 private extreme-difficulty mathematics problemsPython-enabled iterative mathematical problem solvingResearch-level mathematics requiring hours or days of expert work
Weighted 10%

FrontierMath v2 (Tier 4) 2026 · updated July 20, 2026

USAMO 2026

2026

United States of America Mathematical Olympiad 2026

The premier US mathematical olympiad competition, featuring proof-based problems that require deep mathematical insight and rigorous argumentation at the highest competition level.

Current
6 proof-based problemsMathematical proof constructionInternational olympiad level
Weighted 10%

USAMO 2026 2026 · updated July 20, 2026

korean(8 benchmarks)

View leaderboard

KMMLU

2024

Korean Massive Multitask Language Understanding

Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.

RefreshingDisplay only
35,030 questionsMultiple choice questionsElementary to professional level in Korean
Display only

KMMLU 2024 · updated July 20, 2026

KMMLU-Hard

2025

KMMLU-Hard

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

CurrentDisplay only
~5,000 questionsMultiple choice questionsAdvanced Korean reasoning
Display only

KMMLU-Hard 2025 · updated July 20, 2026

KMMLU-Redux

KMMLU-Redux

Cleaned KMMLU from national technical qualification exams, with errors removed, decontaminated, and deduplicated.

RefreshingDisplay only
~3,500 questionsTechnical multiple choiceIndustrial/technical
Display only

KMMLU-Redux · updated July 20, 2026

KMMLU-Pro

KMMLU-Pro

Korean National Professional Licensure exams evaluating professional-grade knowledge.

RefreshingDisplay only
~2,500 questionsProfessional licensure examsProfessional
Display only

KMMLU-Pro · updated July 20, 2026

CLIcK

Cultural and Linguistic Intelligence in Korean

Evaluates Korean culture and linguistics.

RefreshingDisplay only
1,995 questionsCultural/linguistic QAKorean cultural nuances
Display only

CLIcK · updated July 20, 2026

KoBALT

Korean Benchmark for Advanced Linguistic Tasks

Evaluates advanced Korean linguistic competence.

RefreshingDisplay only
Linguistics questionsAdvanced linguisticsAdvanced linguistic phenomena
Display only

KoBALT · updated July 20, 2026

Korean CSAT

College Scholastic Ability Test (수능)

The Korean SAT exam.

RefreshingDisplay only
Multi-subject examStandardized testHigh school to college level
Display only

Korean CSAT · updated July 20, 2026

HRM8K

HAE-RAE Math 8K

Korean mathematical reasoning (high-school to Olympiad level).

RefreshingDisplay only
8,011 instancesMath word problemsOlympiad level
Display only

HRM8K · updated July 20, 2026

External benchmark mirrors(44 benchmarks)

View leaderboard

KindBench

2026

KindBench Psychological Safety Benchmark

A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.

CurrentDisplay only
16 multi-turn scenarios, 72 criteriaJudge-scored behavioral audit with human reviewAdversarial psychological-safety evaluation
Display only

KindBench v0.1.0 · updated July 20, 2026

LiveBench

2024

LiveBench

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

RefreshingDisplay only
23 objective tasks across 7 categoriesMean of category averagesBroad frontier-model evaluation
Display only

LiveBench 2024 · updated July 20, 2026

Vals Index

2026

Vals Index v1.2

Vals AI composite benchmark across finance and coding tasks, including Finance Agent v2, CorpFin v2, SWE-bench, Terminal-Bench 2.1, and Vibe Code Bench.

CurrentDisplay only
Finance and coding componentsComposite scorePrivate economic-work benchmark composite
Display only

Vals Index 2026 · updated July 20, 2026

Vals Multimodal Index

2026

Vals Multimodal Index v1.1

Vals AI multimodal composite across finance, coding, education, and mortgage-tax task families.

CurrentDisplay only
Finance, coding, education, and mortgage-tax componentsComposite scorePrivate multimodal economic-work benchmark composite
Display only

Vals Multimodal Index 2026 · updated July 20, 2026

CorpFin v2

2026

Vals CorpFin v2

Vals AI private benchmark for understanding long-context credit agreements.

CurrentDisplay only
Credit-agreement understanding tasksAccuracy scoreProfessional finance document reasoning
Display only

CorpFin v2 2026 · updated July 20, 2026

MedCode

2026

Vals MedCode

Vals AI healthcare benchmark for whether models can support the medical billing process.

CurrentDisplay only
Medical billing support tasksAccuracy scoreProfessional healthcare administration
Display only

MedCode 2026 · updated July 20, 2026

MedScribe

2026

Vals MedScribe

Vals AI healthcare benchmark for whether models can support doctors with administrative work.

CurrentDisplay only
Medical administrative support tasksAccuracy scoreProfessional healthcare administration
Display only

MedScribe 2026 · updated July 20, 2026

MortgageTax

2026

Vals MortgageTax

Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.

CurrentDisplay only
Mortgage and tax extraction tasksAccuracy scoreProfessional mortgage-tax document reasoning
Display only

MortgageTax 2026 · updated July 20, 2026

ProofBench

2026

Vals ProofBench

Vals AI automated theorem-proving benchmark.

CurrentDisplay only
Automated theorem provingAccuracy scoreFormal proof reasoning
Display only

ProofBench 2026 · updated July 20, 2026

LegalBench

2026

Vals LegalBench

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

CurrentDisplay only
Legal reasoning task viewsAccuracy scoreProfessional legal reasoning
Display only

LegalBench 2026 · updated July 20, 2026

CaseLaw v2

2026

Vals CaseLaw v2

Vals AI private question-answer benchmark over Canadian court cases.

CurrentDisplay only
Canadian case-law question answeringAccuracy scoreProfessional legal retrieval and reasoning
Display only

CaseLaw v2 2026 · updated July 20, 2026

DeepSWE

2026

DeepSWE

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

CurrentDisplay only
113 software engineering tasks across 91 repositories and 5 languagesPass@1 with confidence interval, cost, time, and token metadataLong-horizon software engineering
Display only

DeepSWE 2026 · updated July 20, 2026

SWE-Marathon

2026

SWE-Marathon

A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.

CurrentDisplay only
20 multi-hour software engineering tasksTask resolution and trajectory reviewUltra-long-horizon software engineering
Display only

SWE-Marathon 2026 · updated July 20, 2026

ExploitBench

2026

ExploitBench v8-bench

A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

CurrentDisplay only
V8 exploit synthesis runsCapability coverage percentage over 16 flagsBrowser exploitation and cybersecurity
Display only

ExploitBench 2026 · updated July 20, 2026

GBA-Eval

2026

GBA-Eval

An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.

CurrentDisplay only
27 emulator test casesOverall emulator scoreLong-horizon systems programming
Display only

GBA-Eval 2026 · updated July 20, 2026

CAIS Text Leaderboard

2025

CAIS AI Dashboard Text Capabilities Index

A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

CurrentDisplay only
HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuestsAverage component scoreComposite frontier text capability
Display only

CAIS Text Leaderboard 2025 · updated July 20, 2026

WeirdML

2026

WeirdML v2

A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.

CurrentDisplay only
17 novel ML engineering tasksAverage accuracy across tasksNovel dataset modeling and iterative debugging
Display only

WeirdML 2026 · updated July 20, 2026

ALE-Bench

2026

Agents Last Exam

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

CurrentDisplay only
152 ALE-V1 professional workflow tasks across 13 top-level domainsPass rate, partial-credit score, cost, token, and duration metadataReal-world agentic workflows
Display only

ALE-Bench 2026 · updated July 20, 2026

RuneScape-Bench

2026

RuneBench / runescape-bench

An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.

CurrentDisplay only
16 RuneScape skill-training tasksAverage log XP-rate scoreAgentic gameplay automation
Display only

RuneScape-Bench 2026 · updated July 20, 2026

Toloka Arena

2026

Toloka Arena

An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.

CurrentDisplay only
Private simulated enterprise workflowspass^5 arena scoreAgentic workflow reliability
Display only

Toloka Arena 2026 · updated July 20, 2026

Vals SWE-bench mirror

2026

Vals-hosted SWE-bench mirror

Vals AI hosted SWE-bench view for solving production software engineering tasks.

CurrentDisplay only
Software engineering issue-resolution tasksAccuracy scoreProduction software engineering
Display only

Vals SWE-bench mirror 2026 · updated July 20, 2026

Vals Terminal-Bench 2.0 mirror

2026

Vals-hosted Terminal-Bench 2.0 mirror

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

CurrentDisplay only
Terminal task difficulty splitsAccuracy scoreTerminal-based agent execution
Display only

Vals Terminal-Bench 2.0 mirror 2026 · updated July 20, 2026

Vals LiveCodeBench mirror

2026

Vals-hosted LiveCodeBench mirror

Vals AI implementation of LiveCodeBench with easy, medium, and hard task splits.

CurrentDisplay only
Coding problem difficulty splitsAccuracy scoreContamination-resistant coding problems
Display only

Vals LiveCodeBench mirror 2026 · updated July 20, 2026

Vals GPQA Diamond mirror

2026

Vals-hosted GPQA Diamond mirror

Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.

CurrentDisplay only
GPQA Diamond task splitsAccuracy scoreGraduate science reasoning
Display only

Vals GPQA Diamond mirror 2026 · updated July 20, 2026

Vals MMLU-Pro mirror

2026

Vals-hosted MMLU-Pro mirror

Vals AI hosted MMLU-Pro view with subject-level task splits.

CurrentDisplay only
MMLU-Pro subject splitsAccuracy scoreProfessional academic reasoning
Display only

Vals MMLU-Pro mirror 2026 · updated July 20, 2026

EMB

2026

Vals EMB

Evaluating agents on Excel-based financial modeling tasks

CurrentDisplay only
Excel-based financial modeling tasksAccuracy scoreProfessional finance modeling
Display only

EMB 2026 · updated July 20, 2026

CyberBench

2026

Vals CyberBench

Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and stop crashing after the fix?

CurrentDisplay only
OSS-Fuzz PoC and patch-verification tasksAccuracy scoreAutonomous cybersecurity exploit reproduction
Display only

CyberBench 2026 · updated July 20, 2026

TaxEval v2

2026

Vals TaxEval v2

A Vals-created set of questions and responses to tax questions

CurrentDisplay only
Tax question answering and response evaluationAccuracy scoreProfessional tax reasoning
Display only

TaxEval v2 2026 · updated July 20, 2026

Harvey's Legal Agent Benchmark

2026

Vals Harvey's Legal Agent Benchmark

Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools

CurrentDisplay only
Legal agent work across documents, spreadsheets, presentations, and filesAccuracy scoreProfessional legal workflow automation
Display only

Harvey's Legal Agent Benchmark 2026 · updated July 20, 2026

Terminal-Bench 2.1

2026

Vals Terminal-Bench 2.1

State-of-the-art set of difficult terminal-based tasks

CurrentDisplay only
Terminal-based task executionAccuracy scoreFrontier terminal-agent execution
Display only

Terminal-Bench 2.1 2026 · updated July 20, 2026

Code Migration

2026

Vals Code Migration

Can language models reimplement real-world programs in another language?

CurrentDisplay only
Real-world program reimplementation in another languageAccuracy scoreProduction code migration
Display only

Code Migration 2026 · updated July 20, 2026

Legal Research Bench

2026

Vals Legal Research Bench

Evaluating agents on legal research tasks across diverse areas of US law

CurrentDisplay only
US-law legal research tasksAccuracy scoreProfessional legal research
Display only

Legal Research Bench 2026 · updated July 20, 2026

MedQA

2026

Vals MedQA

Evaluating language model bias in medical questions.

CurrentDisplay only
Medical question answeringAccuracy scoreMedical knowledge and bias evaluation
Display only

MedQA 2026 · updated July 20, 2026

AIME

2026

Vals AIME

Challenging national math exam given to top high-school students

CurrentDisplay only
AIME math problemsAccuracy scoreCompetition math
Display only

AIME 2026 · updated July 20, 2026

MATH 500

2026

Vals MATH 500

Academic math benchmark on probability, algebra, and trigonometry

CurrentDisplay only
MATH 500 academic math problemsAccuracy scoreAdvanced academic math
Display only

MATH 500 2026 · updated July 20, 2026

MGSM

2026

Vals MGSM

A multilingual benchmark for mathematical questions.

CurrentDisplay only
Multilingual grade-school math questionsAccuracy scoreMultilingual mathematical reasoning
Display only

MGSM 2026 · updated July 20, 2026

MMMU

2026

Vals MMMU

Multimodal Multi-task Benchmark

CurrentDisplay only
Multimodal academic task suiteAccuracy scoreMultimodal college-level reasoning
Display only

MMMU 2026 · updated July 20, 2026

SAGE

2026

Vals SAGE

Student Assessment with Generative Evaluation

CurrentDisplay only
Student assessment with generative evaluationAccuracy scoreEducation assessment reasoning
Display only

SAGE 2026 · updated July 20, 2026

IOI

2026

Vals IOI

Based on the International Olympiad in Informatics

CurrentDisplay only
International Olympiad in Informatics-style programming tasksAccuracy scoreOlympiad programming
Display only

IOI 2026 · updated July 20, 2026

ProgramBench

2026

Vals ProgramBench

Can language models rebuild programs from scratch?

CurrentDisplay only
Program reconstruction tasksAccuracy scoreCleanroom software engineering
Display only

ProgramBench 2026 · updated July 20, 2026

SkillsBench

2026

Vals SkillsBench

How important are skills for agents?

CurrentDisplay only
Agent skill-importance tasksAccuracy scoreAgent skill evaluation
Display only

SkillsBench 2026 · updated July 20, 2026

Agent Poker Bench

2026

Vals Agent Poker Bench

Which model can make the most money playing poker?

CurrentDisplay only
Poker-playing agent trialsAccuracy scoreStrategic game-agent decision making
Display only

Agent Poker Bench 2026 · updated July 20, 2026

Public Benefits Bench v1.1

2026

Vals Public Benefits Bench v1.1

Can AI help people navigate SNAP benefits?

CurrentDisplay only
SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation
Display only

Public Benefits Bench v1.1 2026 · updated July 20, 2026

Public Benefits Bench v1

2026

Vals Public Benefits Bench v1

Can AI help people navigate SNAP benefits?

CurrentDisplay only
SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation
Display only

Public Benefits Bench v1 2026 · updated July 20, 2026