AI Benchmarks Directory
BenchLM tracks 416 AI benchmarks across coding, agentic, reasoning, math, knowledge, multimodal, instruction following, and multilingual tasks. Each record links the evaluation to its ranked leaderboard and September 2026 data.
Agentic(87 benchmarks)
DRACO
Data Research and Analysis with Complex Operations
Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.
Agentic data research and analysis tasksNormalized rubric scoreProfessional data analysis
CurrentDisplay only
DRACO 2026 · updated September 2, 2026
BrowseComp (10-agent, prerelease)
Multi-Agent BrowseComp — 10-agent team prerelease configuration
BrowseComp accuracy from ten collaborating Opus 5 agents on a pre-release model and unreleased effort configuration.
BrowseComp web-research tasks10-agent team accuracyLong-horizon web research
CurrentDisplay only
BrowseComp (10-agent, prerelease) 2026 · updated September 2, 2026
MCP-Atlas claim coverage
MCP-Atlas mean claim coverage
Average coverage of required claims in answers produced during real-world MCP tool-use workflows.
Production-like multi-server MCP workflowsMean claim coverageReal-world tool use
CurrentDisplay only
MCP-Atlas claim coverage 2026 · updated September 2, 2026
LAB all-pass (Anthropic harness)
Legal Agent Benchmark all-pass rate — Anthropic harness
Strict task success requiring every expert-written legal-work rubric criterion to pass.
1,235 legal-agent tasksAll-criteria pass rateProfessional legal work
CurrentDisplay only
LAB all-pass (Anthropic harness) 2026 · updated September 2, 2026
LAB criterion-pass (Anthropic harness)
Legal Agent Benchmark mean criterion-pass rate — Anthropic harness
Mean fraction of expert-written rubric criteria passed across legal-agent tasks.
1,235 legal-agent tasksMean criterion-pass rateProfessional legal work
CurrentDisplay only
LAB criterion-pass (Anthropic harness) 2026 · updated September 2, 2026
LAB all-pass (Harvey held-out)
Legal Agent Benchmark all-pass rate — Harvey held-out set
Harvey AI's strict held-out task success rate requiring every rubric criterion to pass.
Harvey-held-out legal-agent tasksAll-criteria pass rateProfessional legal work
CurrentDisplay only
LAB all-pass (Harvey held-out) 2026 · updated September 2, 2026
LAB criterion-pass (Harvey held-out)
Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set
Harvey AI's mean criterion-level score on its held-out legal-agent evaluation.
Harvey-held-out legal-agent tasksMean criterion-pass rateProfessional legal work
CurrentDisplay only
LAB criterion-pass (Harvey held-out) 2026 · updated September 2, 2026
Toolathlon Verified Pass@3
Toolathlon Verified Pass@3
Fraction of Toolathlon Verified tasks solved in at least one of three trials.
108 verified real-world tool-use tasksPass@3Long-horizon application tool use
CurrentDisplay only
Toolathlon Verified Pass@3 2026 · updated September 2, 2026
Toolathlon Verified Pass³
Toolathlon Verified Pass cubed
Fraction of Toolathlon Verified tasks solved in all three independent trials.
108 verified real-world tool-use tasksAll-three-trials pass rateLong-horizon application tool use
CurrentDisplay only
Toolathlon Verified Pass³ 2026 · updated September 2, 2026
Toolathlon Verified avg. turns
Toolathlon Verified average assistant turns
Average assistant turns per Toolathlon Verified trajectory.
108 verified real-world tool-use tasksAverage trajectory lengthLong-horizon application tool use
CurrentDisplay only
Toolathlon Verified avg. turns 2026 · updated September 2, 2026
Terminal-Bench 3.0
Terminal-Bench 3.0
A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.
74 professional computer-work tasks across 7 domainsTask completion rateFrontier autonomous knowledge work
CurrentDisplay only
Terminal-Bench 3.0 · updated September 2, 2026
Terminal-Bench 4.0
Terminal-Bench 4.0
The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing unstable tasks, and removing tasks that no longer separate frontier systems.
66 professional computer-work tasksTask completion rate across 5 trials per taskFrontier autonomous knowledge work
CurrentDisplay only
Terminal-Bench 4.0.0 · updated September 2, 2026
Terminal-Bench-Science 0.1
Terminal-Bench-Science 0.1
A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.
70 expert-curated scientific research workflowsResolution rate across 3 trials per taskFrontier scientific research workflows
CurrentDisplay only
Terminal-Bench-Science 0.1.0 · updated September 2, 2026
EQ-Bench 4
EQ-Bench 4
A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.
120 multi-turn persona scenariosPairwise EloEmotional and social intelligence
CurrentDisplay only
EQ-Bench 4 2026 · updated September 2, 2026
AutoCAD-Bench
AutoCAD-Bench
A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.
21 2D drawing tasks and 29 3D modeling tasksTask completion rate at a 75-point rubric thresholdProfessional CAD computer use
CurrentDisplay only
AutoCAD-Bench 2026 · updated September 2, 2026
Design Arena Agentic Web Dev
Design Arena Agentic Web Dev Elo
A display-only Elo rating from blinded comparisons of multi-file web applications built by coding agents.
Multi-file web application developmentElo from blinded human preferencesAgentic frontend development
CurrentDisplay only
Design Arena Agentic Web Dev 2026 · updated September 2, 2026
AA Briefcase
Artificial Analysis Briefcase
An independently evaluated professional-work benchmark reported as Elo.
Professional knowledge-work tasksEloProfessional work
CurrentDisplay only
AA Briefcase 2026 · updated September 2, 2026
AA AutomationBench
Artificial Analysis AutomationBench
An independently evaluated automation benchmark from Artificial Analysis.
Business-process automation tasksTask success rateAgentic automation
CurrentDisplay only
AA AutomationBench 2026 · updated September 2, 2026
AA EnterpriseOps-Gym
Artificial Analysis EnterpriseOps-Gym
An independently evaluated enterprise-operations benchmark from Artificial Analysis.
Enterprise operations workflowsTask success rateEnterprise agent operations
CurrentDisplay only
AA EnterpriseOps-Gym 2026 · updated September 2, 2026
AA Harvey LAB
Artificial Analysis Harvey LAB-AA
An independently evaluated legal-agent benchmark from Artificial Analysis.
Legal agent tasksTask success rateProfessional legal work
CurrentDisplay only
AA Harvey LAB 2026 · updated September 2, 2026
AA ITBench
Artificial Analysis ITBench-AA
An independently evaluated IT-operations benchmark from Artificial Analysis.
IT incident-response tasksTask success rateEnterprise IT operations
CurrentDisplay only
AA ITBench 2026 · updated September 2, 2026
AA Tau3 Banking
Artificial Analysis Tau3-Banking
An independently evaluated Tau3 banking benchmark from Artificial Analysis.
Banking tool-use workflowsTask success rateAgentic banking workflows
CurrentDisplay only
AA Tau3 Banking 2026 · updated September 2, 2026
Terminal-Bench 2.0
Terminal-Bench 2.0
A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.
Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering
CurrentWeighted 38%
Terminal-Bench 2 · updated September 2, 2026
Terminal-Bench 2.1
Terminal-Bench 2.1 (provider run)
A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.
Terminal-based software-agent tasksInteractive task success rateProfessional software engineering
CurrentDisplay only
Terminal-Bench 2.1 2026 · updated September 2, 2026
BrowseComp
BrowseComp
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
Research questions requiring browsingWeb search and evidence synthesisHard web research
CurrentWeighted 28%
BrowseComp 2026 · updated September 2, 2026
HLE w/ tools
Humanity's Last Exam with tools
Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.
Expert questions with tool usePass@1Frontier tool-augmented reasoning
CurrentDisplay only
HLE w/ tools 2026 · updated September 2, 2026
GDPval-AA
GDPval-AA
An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.
Agentic real-world work tasksEloProfessional agentic workflows
CurrentDisplay only
GDPval-AA 2026 · updated September 2, 2026
GDPval-AA
GDPval-AA normalized
A display-only Artificial Analysis normalized score for economically valuable tasks.
Economically valuable tasksNormalized scoreProfessional agentic workflows
CurrentDisplay only
GDPval-AA 2026 · updated September 2, 2026
AA Agentic Index
Artificial Analysis Agentic Index
A display-only Artificial Analysis agentic index.
Cross-benchmark agentic indexAggregated model scoreDisplay-only external reference
CurrentDisplay only
AA Agentic Index 2026 · updated September 2, 2026
APEX-Agents-AA
APEX-Agents-AA
Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.
452 professional-services agent tasksPass@1Long-horizon workplace agent tasks
CurrentDisplay only
APEX-Agents-AA 2026 · updated September 2, 2026
Gert Labs
Gert Labs Composite Game Benchmark
A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.
Novel game environmentsComposite game leaderboardAgentic coding and decision-making
CurrentDisplay only
Gert Labs 2026 · updated September 2, 2026
OSWorld-Verified
OSWorld-Verified
OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.
369 real-world computer tasks (361 when eight Google Drive tasks are excluded)Execution-based interactive task successMulti-step desktop and cross-application workflows
CurrentWeighted 34%
OSWorld Verified · updated September 2, 2026
OSWorld 2.0
OSWorld 2.0
A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.
108 long-horizon computer-use workflowsInteractive computer-use evaluationLong-horizon professional workflows
CurrentDisplay only
OSWorld 2.0 2026 · updated September 2, 2026
CyberGym
CyberGym
A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
1,507 vulnerability analysis instancesVulnerability reproduction and PoC generationReal-world cybersecurity
CurrentDisplay only
CyberGym 2026 · updated September 2, 2026
CWE-Bench
CWE-Bench
An external benchmark for evaluating whether coding agents can produce correct patches for real-world software vulnerabilities.
Real-world vulnerability patching tasksPass@1Automated vulnerability remediation
CurrentDisplay only
CWE-Bench 2026 · updated September 2, 2026
CTI-REALM
CTI-REALM
A cybersecurity benchmark that measures whether an agent can turn raw threat-intelligence reports into working detection rules.
Threat-intelligence-to-detection-rule workflowsSuccess rateProfessional cyber threat detection
CurrentDisplay only
CTI-REALM 2026 · updated September 2, 2026
Cybench
Cybench
A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.
40 professional CTF tasksCybersecurity agent task completionProfessional cybersecurity
CurrentDisplay only
Cybench 2025 · updated September 2, 2026
ExploitGym
ExploitGym
A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.
898 exploitation tasksWorking exploit generationAdvanced cybersecurity exploitation
CurrentDisplay only
ExploitGym 2026 · updated September 2, 2026
JobBench
JobBench
An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.
130 tasks across 35 occupationsAgentic workplace deliverablesProfessional multi-source workflows
CurrentDisplay only
JobBench 2026 · updated September 2, 2026
BrowseComp-VL
BrowseComp-VL
A vision-language browsing benchmark for multimodal web research and tool-use workflows.
Multimodal browsing tasksVision-language web research evaluationMultimodal browser-agent
CurrentDisplay only
BrowseComp-VL 2026 · updated September 2, 2026
OSWorld
OSWorld
A computer-use benchmark for GUI task completion across the broader OSWorld task suite.
Computer-use tasksInteractive GUI evaluationBroad computer-use suite
CurrentDisplay only
OSWorld 2026 · updated September 2, 2026
AndroidWorld
AndroidWorld
A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.
Android app workflowsInteractive mobile-agent evaluationComplex mobile task completion
CurrentDisplay only
AndroidWorld 2026 · updated September 2, 2026
WebVoyager
WebVoyager
A browser-agent benchmark for completing multi-step workflows on live websites.
Live website workflowsInteractive browser-agent evaluationMulti-step web navigation
CurrentDisplay only
WebVoyager 2026 · updated September 2, 2026
MCP Atlas
MCP Atlas
A benchmark for tool-calling over Model Context Protocol integrations and external tools.
Tool-integrated agent tasksInteractive tool-calling evaluationAdvanced tool use
CurrentDisplay only
MCP Atlas 2026 · updated September 2, 2026
Kimi Claw 24/7
Kimi Claw 24/7 Bench
A Moonshot AI internal long-horizon agent benchmark for persistent professional coworking tasks.
17 professional scenarios, 610 evaluation pointsAverage pass rate across repeated OpenClaw runsLong-horizon agentic work
CurrentDisplay only
Kimi Claw 24/7 2026 · updated September 2, 2026
MCP Mark Verified
MCPMark-Verified
A human-verified edition of MCPMark for MCP tool use across Notion, GitHub, Filesystem, Postgres, and Playwright server environments.
MCP tool-use tasks across five server environmentsInteractive MCP task completionAdvanced tool use
CurrentDisplay only
MCP Mark Verified 2026 · updated September 2, 2026
Toolathlon
Toolathlon
A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.
Multi-tool workflowsInteractive tool-calling evaluationAdvanced tool use
CurrentDisplay only
Toolathlon 2026 · updated September 2, 2026
Toolathlon-Verified
Toolathlon-Verified
A verified tool-use benchmark variant for completing multi-step workflows with external tools.
Verified multi-tool workflowsInteractive tool-use scoreAdvanced tool use
CurrentDisplay only
Toolathlon-Verified 2026 · updated September 2, 2026
AutomationBench
AutomationBench
An agent benchmark for completing automation workflows in reproducible task environments.
600 public automation tasksAgent task-completion scoreLong-horizon automation
CurrentDisplay only
AutomationBench 2026 · updated September 2, 2026
Agents' Last Exam
Agents' Last Exam
An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.
Agent tasksProvider-reported task scoreAdvanced agentic work
CurrentDisplay only
Agents' Last Exam 2026 · updated September 2, 2026
APEX-Agents
APEX-Agents
A professional-services agent benchmark covering long-horizon knowledge-work tasks.
Professional-services agent tasksAgent task-completion scoreLong-horizon professional work
CurrentDisplay only
APEX-Agents 2026 · updated September 2, 2026
SpreadsheetBench 2
SpreadsheetBench 2
A spreadsheet-focused benchmark for agentic analysis and editing workflows.
Spreadsheet analysis and editing tasksAgent task-completion scoreProfessional spreadsheet work
CurrentDisplay only
SpreadsheetBench 2 2026 · updated September 2, 2026
DECK-Bench
DECK-Bench (Internal)
Moonshot AI's internal benchmark for presentation and deck-production workflows.
Internal presentation workflowsInternal evaluation scoreProfessional presentation creation
CurrentDisplay only
DECK-Bench 2026 · updated September 2, 2026
ZClawBench
ZClawBench
A Z.AI benchmark for OpenClaw-style agent workflows spanning information search, office work, data analysis, development and operations, automation, and security.
OpenClaw agent workflowsEnd-to-end agent benchmarkBroad productivity and operations workflows
CurrentDisplay only
ZClawBench 2026 · updated September 2, 2026
τ²-bench results
τ²-Bench Tool-Agent-User Evaluation
This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.
Airline, retail, and telecom customer-service task setsPublished domain success or pass^k resultsDual-control customer-service workflows
CurrentDisplay only
τ²-Bench 2026 · updated September 2, 2026
DeepSearchQA
DeepSearchQA
An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.
Agentic browsing and list-answer questionsSearch / open / find browser-agent evaluationAgentic web research
CurrentDisplay only
DeepSearchQA 2026 · updated September 2, 2026
τ²-bench Airline
τ²-Bench Airline Domain
τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.
Airline customer-service tasksDomain success under a published trial policyPolicy-constrained airline support workflows
CurrentDisplay only
τ²-bench Airline 2025 · updated September 2, 2026
PinchBench
PinchBench
An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.
23 OpenClaw agent tasksAverage success rate from official runsLong-horizon agent workflows
CurrentDisplay only
PinchBench 2026 · updated September 2, 2026
OpenHands Index
OpenHands Index
A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.
SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIAMacro-average across five coding-agent categoriesReal-world software engineering agent tasks
CurrentDisplay only
OpenHands Index 2025 · updated September 2, 2026
SWE-Atlas Refactoring
SWE-Atlas Refactoring
A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.
SWE-Atlas refactoring tasksRefactoring score with confidence intervalsReal-world software-engineering agent tasks
CurrentDisplay only
SWE-Atlas Refactoring 2026 · updated September 2, 2026
SWE Refactor Bench
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.
20 whole-repository stack migrationsMigration audit, frozen behavioral checks, and agentic verification6- to 30-hour autonomous repository migrations
CurrentDisplay only
SWE Refactor Bench 2026 · updated September 2, 2026
AI4AI-Bench
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.
10 AI training-algorithm design tasksFour-hour code rewrite followed by sealed training and held-out evaluationEnd-to-end AI research and algorithm design
CurrentDisplay only
AI4AI-Bench v1.5 · updated September 2, 2026
InferenceBench
InferenceBench
A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed.
4 inference-serving optimization scenariosTwo-hour autonomous CLI agent runOpen-ended ML systems engineering
CurrentDisplay only
InferenceBench 2026 · updated September 2, 2026
EdgeBench
EdgeBench
A ByteDance Seed benchmark of 134 real-world, day-scale tasks that measures how autonomous agents learn from environment feedback over 12+ hour interaction horizons, spanning scientific and ML, systems and software engineering, optimization, knowledge, formal, and game domains.
134 tasks (51 public) across 6 domainsLong-horizon interactive agent evaluationDay-scale expert tasks
CurrentDisplay only
EdgeBench 2026 · updated September 2, 2026
BFCL v4
Berkeley Function Calling Leaderboard v4
A function-calling benchmark for tool selection, schema adherence, and argument correctness.
Function-calling tasksTool invocation and schema evaluationAdvanced tool use
CurrentDisplay only
BFCL v4 2026 · updated September 2, 2026
MLE-Bench Lite
MLE-Bench Lite
A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.
Low-resource ML competitionsAutonomous iterative ML optimizationAgentic machine learning
CurrentDisplay only
MLE-Bench Lite 2026 · updated September 2, 2026
MM-ClawBench
MM-ClawBench
An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.
OpenClaw-style real-world tasksAgent workflow evaluationBroad real-world agentic execution
CurrentDisplay only
MM-ClawBench 2026 · updated September 2, 2026
Claw-Eval
Claw-Eval
A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.
300 tasks, 2,159 rubricsEnd-to-end autonomous-agent evaluation with Pass^3 scoringReal-world general, multi-turn, and native multimodal agent execution
CurrentDisplay only
Claw-Eval 2026 · updated September 2, 2026
ResearchClawBench
ResearchClawBench
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
40 tasks across 10 scientific domainsEnd-to-end autonomous research evaluation with RADS scoringScientific research re-discovery
CurrentDisplay only
ResearchClawBench 2026 · updated September 2, 2026
QwenClawBench
QwenClawBench
Qwen's internal OpenClaw-style benchmark for measuring broad real-world agent performance across practical productivity and research tasks.
Real-world agent workflowsEnd-to-end agent evaluationBroad real-world agentic execution
CurrentDisplay only
QwenClawBench 2026 · updated September 2, 2026
QwenWebBench
QwenWebBench
A Qwen benchmark for artifact and webpage generation quality reported as an Elo-style rating.
Web artifacts and interactive deliverablesElo-style artifact benchmarkArtifact generation
CurrentDisplay only
QwenWebBench 2026 · updated September 2, 2026
τ³-bench results
τ³-Bench Tool-Agent-User Evaluation
τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
Corrected customer-service tasks plus knowledge and voice evaluation modesPublished domain or average success resultsLong-horizon, multimodal, and knowledge-aware tool use
CurrentDisplay only
τ³-bench results 2026 · updated September 2, 2026
VITA-Bench
VITA-Bench
An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.
Interactive consumer-service agent tasksEnd-to-end interactive agent evaluationLong-horizon real-world workflows
CurrentDisplay only
VITA-Bench 2025 · updated September 2, 2026
DeepPlanning
DeepPlanning
A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.
Travel planning and constrained shoppingLong-horizon planning benchmarkConstrained agent planning
CurrentDisplay only
DeepPlanning 2026 · updated September 2, 2026
MCP-Tasks
MCP-Tasks
A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.
MCP-integrated tool tasksInteractive tool-use evaluationAdvanced MCP workflows
CurrentDisplay only
MCP-Tasks 2026 · updated September 2, 2026
WideResearch
WideResearch
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
Open-ended research tasksMulti-source research evaluationBroad research-agent workflows
CurrentDisplay only
WideResearch 2026 · updated September 2, 2026
CoWorkBench
CoWorkBench
Qwen's internal benchmark for long-horizon professional work across computer science, finance, law, medicine, and other productivity domains.
Long-horizon professional workflowsProvider-run agent scoreCross-domain professional work
CurrentDisplay only
CoWorkBench 2026 · updated September 2, 2026
MobileWorld
MobileWorld
A mobile-use agent benchmark for completing interactive tasks in smartphone environments.
Interactive mobile-device workflowsMobile agent task scoreLong-horizon mobile computer use
CurrentDisplay only
MobileWorld 2026 · updated September 2, 2026
GAIA
General AI Assistants
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.
466
RefreshingDisplay only
GAIA 2024 · updated September 2, 2026
TAU-bench
Tool-Agent-User Benchmark
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
Airline and retail task sets in the archived 2024 releaseDomain-specific pass^1 through pass^4 task successPolicy-constrained, multi-turn customer service
RefreshingDisplay only
TAU-bench 2024 · updated September 2, 2026
WebArena
WebArena Web Agent Benchmark
WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.
812 long-horizon browser tasksEnd-state task successStateful multi-site browser work
RefreshingDisplay only
WebArena 2024 · updated September 2, 2026
WebArena-Verified
WebArena-Verified Browser Agent Benchmark
WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.
812 verified tasks; separate 258-task Hard subsetDeterministic end-state task successAudited stateful browser work
CurrentDisplay only
WebArena-Verified 2025 · updated September 2, 2026
MEWC
Multi-Environment Web Challenge
A benchmark that evaluates AI agents on multi-environment web challenges, testing navigation and task completion across diverse live web environments.
Web-agent tasksBrowser task completionOpen-web agent workflows
CurrentDisplay only
MEWC 2026 · updated September 2, 2026
Finance Agent v2
Finance Agent v2
Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.
Financial analyst task categoriesMean score across repeated runsProfessional expert-task agent workflow
CurrentDisplay only
Finance Agent v2 2026 · updated September 2, 2026
Market-Bench
Market-Bench
A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.
3 quantitative-trading strategiesBacktester implementation scored by mean absolute errorMarket simulation and quantitative coding
CurrentDisplay only
Market-Bench 2025 · updated September 2, 2026
GDPval rubrics
GDPval rubrics
A display-only provider-table GDPval rubric score for economically valuable work tasks.
Economically valuable work tasksRubric scoreProfessional agentic workflows
CurrentDisplay only
GDPval rubrics 2026 · updated September 2, 2026
BankerToolBench
BankerToolBench
A display-only provider benchmark for finance-oriented tool-use and agent workflows.
Finance and banking tool-use tasksTask success rateProfessional finance-agent workflows
CurrentDisplay only
BankerToolBench 2026 · updated September 2, 2026
Coding(60 benchmarks)
ProgramBench (episode 1)
ProgramBench hidden-test pass rate after episode 1
Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.
166 golden program-reconstruction tasksHidden-test pass rate after episode 1Long-context clean-room software engineering
CurrentDisplay only
ProgramBench (episode 1) 2026 · updated September 2, 2026
AA LiveCodeBench
Artificial Analysis LiveCodeBench
An independently evaluated LiveCodeBench result from Artificial Analysis.
Contamination-resistant coding tasksPass rateCompetitive programming
CurrentDisplay only
AA LiveCodeBench 2026 · updated September 2, 2026
AA Terminal-Bench 2.1
Artificial Analysis Terminal-Bench v2.1
An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.
Terminal-based agent tasksTask success rateAgentic software engineering
CurrentDisplay only
AA Terminal-Bench 2.1 2026 · updated September 2, 2026
Terminal-Bench 2.1
Terminal-Bench 2.1 (provider run)
A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.
Terminal-based software-agent tasksInteractive task success rateProfessional software engineering
CurrentDisplay only
Terminal-Bench 2.1 2026 · updated September 2, 2026
HumanEval
Evaluating Large Language Models Trained on Code
A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.
164 problemsPython function generationIntroductory to intermediate programming
StaleSaturatedDisplay only
HumanEval · updated September 2, 2026
BigCodeBench
BigCodeBench
A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.
Code generation tasksPass@1Software engineering
CurrentDisplay only
BigCodeBench 2026 · updated September 2, 2026
Codeforces
Codeforces Rating
Competitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.
Competitive programming contestsRatingElite competitive programming
CurrentDisplay only
Codeforces 2026 · updated September 2, 2026
Terminal-Bench 2.0
Terminal-Bench 2.0
A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.
Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering
CurrentDisplay only
Terminal-Bench 2 · updated September 2, 2026
SWE-bench Verified
Software Engineering Benchmark Verified
A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.
500 verified issuesCode patch generationProfessional software engineering
RefreshingWeighted 16%
SWE-bench Verified 2024 · updated September 2, 2026
SWE-Rebench
SWE-Rebench
A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.
Fresh GitHub issues (rolling window)Code patch generationProfessional software engineering
CurrentWeighted 20%
Rolling 2026 window · updated September 2, 2026
LiveCodeBench
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.
Continuously updated contest problemsCompetitive-programming evaluationCompetitive programming level
CurrentWeighted 38%
Rolling 2026 set · updated September 2, 2026
LiveCodeBench v6
LiveCodeBench v6
LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.
Fresh programming problemsProvider-published v6 competitive programming resultsCompetitive programming level
CurrentDisplay only
LiveCodeBench v6 2026 · updated September 2, 2026
LiveCodeBench v5
LiveCodeBench v5
LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside the rolling weighted lane.
July 2024 to May 2025 release windowProvider-published v5 competitive programming resultCompetitive programming level
CurrentDisplay only
LiveCodeBench v5 2025 · updated September 2, 2026
LiveCodeBench Pass@1-COT
LiveCodeBench Pass@1 with Chain-of-Thought
This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.
DeepSeek-V4 report evaluation windowPass@1-COT competitive programming resultsCompetitive programming level
CurrentDisplay only
LiveCodeBench Pass@1-COT 2026 · updated September 2, 2026
LiveCodeBench Pro
LiveCodeBench Pro
A harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with quarter-specific public leaderboards and difficulty-aware reporting.
Quarter-specific contest programming setsCompetitive programmingHigh-end contest programming
CurrentDisplay only
LiveCodeBench Pro 2025 · updated September 2, 2026
FLTEval
FLTEval
A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.
FLT project pull requestsLean 4 repository task completionFormal verification / proof engineering
CurrentDisplay only
FLTEval 2026 · updated September 2, 2026
SWE-bench Pro
SWE-bench Pro
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
1,865 repository problemsRepository task completionLong-horizon professional engineering
CurrentWeighted 10%
SWE-bench Pro 2025 · updated September 2, 2026
Senior SWE-Bench
Senior SWE-Bench
A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.
Senior-level repository tasksAgentic software-engineering evaluationProfessional senior engineering
CurrentDisplay only
Senior SWE-Bench v2026.06 · updated September 2, 2026
VulcanBench v3
VulcanBench v3
An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories.
23 post-cutoff repository tasks in the v3 reportPass@1 with low, medium, and high effortProfessional multi-file software engineering
CurrentDisplay only
VulcanBench v3 2026 · updated September 2, 2026
OpenHarmony Bench
OpenHarmony Bench v1.0
An app-level coding benchmark that asks DevEco Code configurations to implement observable behavior in buildable OpenHarmony ArkTS applications.
153 app-development and bug-fix tasksTask completion through DevEco CodeEnd-to-end OpenHarmony application development
CurrentDisplay only
OpenHarmony Bench 2026 · updated September 2, 2026
VulcanBench CII v1
VulcanBench Coding Intelligence Index v1
A post-cutoff software-engineering benchmark with hidden functional tests and regression guards, reported for vendor coding-agent harnesses.
38 validated post-cutoff repository tasksPass@1 with vendor coding-agent harnessesMid-band frontier software engineering
CurrentDisplay only
VulcanBench CII v1 2026 · updated September 2, 2026
FrontierCode 1.1 Main
FrontierCode 1.1 Main
Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.
100 private Main tasks (150 in Extended)Repository task completion with maintainer rubricsFrontier coding-agent quality
CurrentDisplay only
FrontierCode 1.1 Main · updated September 2, 2026
FrontierCode 1.1 Extended
FrontierCode 1.1 Extended
Cognition's 150-task Extended subset of the FrontierCode 1.1 software-engineering benchmark.
150 private software-engineering tasksRepository task completion with maintainer rubricsFrontier coding-agent quality
CurrentDisplay only
FrontierCode 1.1 Extended · updated September 2, 2026
IDE-Bench
IDE-Bench
An 80-task software-engineering benchmark across eight repositories that tests whether autonomous IDE agents can explore, edit, run, and verify code changes end to end.
80 tasks across 8 repositoriesAutonomous IDE-agent task completion (pass@1)End-to-end software engineering
CurrentDisplay only
IDE-Bench 2026 · updated September 2, 2026
App-Bench
App-Bench
A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.
6 full-stack app-building tasksBest-of-three one-shot feature completionProduction-style full-stack application generation
CurrentDisplay only
App-Bench 2025 · updated September 2, 2026
SWE Multilingual
SWE Multilingual
A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.
Multilingual software-engineering tasksRepository task completionProfessional software engineering
CurrentDisplay only
SWE Multilingual 2026 · updated September 2, 2026
SWE Multimodal
SWE-bench Multimodal
A multimodal variant of SWE-bench that adds visual context such as screenshots and design mockups to software engineering issue descriptions.
Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding
CurrentDisplay only
SWE Multimodal 2025 · updated September 2, 2026
CursorBench
CursorBench
Cursor's current first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions.
Harder long-horizon agentic coding tasksCursor agent-loop evaluationProfessional agentic software engineering
CurrentDisplay only
CursorBench 2026 · updated September 2, 2026
Multi-SWE Bench
Multi-SWE Bench
A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.
Multi-language repo tasksRepository task completionProfessional software engineering
CurrentDisplay only
Multi-SWE Bench 2026 · updated September 2, 2026
VIBE-Pro
VIBE-Pro
A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.
Full project delivery tasksRepository-level implementation benchmarkEnd-to-end software delivery
CurrentDisplay only
VIBE-Pro 2026 · updated September 2, 2026
Vibe Code Bench
Vibe Code Bench v1.1
Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.
End-to-end web application buildsFull-stack app implementation benchmarkEnd-to-end software delivery
CurrentDisplay only
Vibe Code Bench 2026 · updated September 2, 2026
ProgramBench
ProgramBench: Can Language Models Rebuild Programs From Scratch?
A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.
200 program reconstruction tasksCleanroom executable reimplementationFull-repository software architecture
CurrentDisplay only
ProgramBench 2026 · updated September 2, 2026
PostTrain Bench
PostTrain Bench
A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.
Post-training software-engineering tasksHarbor agent evaluationFrontier software engineering
CurrentDisplay only
PostTrain Bench 2026 · updated September 2, 2026
FrontierSWE
FrontierSWE
An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.
17 ultra-long-horizon engineering and research tasksMean@5, best@5, average rank, and dominanceUltra-long-horizon frontier software engineering
CurrentDisplay only
FrontierSWE 2026 · updated September 2, 2026
FrontierSWE v2
FrontierSWE v2
A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.
34 ultra-long-horizon engineering and research tasksFive-trial mean task score (Mean@5), 0-100Ultra-long-horizon frontier software engineering
CurrentDisplay only
FrontierSWE v2 2026 · updated September 2, 2026
SWE-Atlas Codebase QnA
SWE-Atlas Codebase QnA
A code-comprehension benchmark covering production repositories across Go, Python, C, and TypeScript.
124 codebase questions across 11 repositoriesMean pass@1Production codebase comprehension
CurrentDisplay only
SWE-Atlas Codebase QnA 2026 · updated September 2, 2026
Bug Hunt Bench
Bug Hunt Bench
A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.
105 planted bugs across two production TypeScript repositoriesStrict planted bugs fixedBlind production-repository bug finding and repair
CurrentDisplay only
Bug Hunt Bench 2026 · updated September 2, 2026
Kimi Code Bench v2
Kimi Code Bench v2
A Moonshot AI internal coding-agent benchmark for realistic software-engineering tasks across mainstream programming languages and production technology stacks.
Realistic coding-agent tasksCoding-agent pass rateProduction software engineering
CurrentDisplay only
Kimi Code Bench v2 2026 · updated September 2, 2026
MLS-Bench Lite
MLS-Bench Lite
A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods.
30 machine-learning research tasksAgentic ML task evaluationML research and systems engineering
CurrentDisplay only
MLS-Bench Lite 2026 · updated September 2, 2026
PaperBench
PaperBench
A research-reproduction benchmark that asks agents to recreate the contributions of AI papers from the paper alone.
AI research-paper reproductionLong-horizon agent evaluationFrontier autonomous research and engineering
CurrentDisplay only
PaperBench 2026 · updated September 2, 2026
QwenReactBench
QwenReactBench
Qwen's internal benchmark for building and rendering bilingual React projects across seven categories.
Bilingual React project constructionBradley-Terry/Elo ratingProduction frontend development
CurrentDisplay only
QwenReactBench 2026 · updated September 2, 2026
NL2Repo
NL2Repo
A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.
Natural language to repository tasksRepository understanding benchmarkSystem-level software comprehension
CurrentDisplay only
NL2Repo 2026 · updated September 2, 2026
DSBench-FullStack
DeepSeek DSBench FullStack
DeepSeek's internal full-stack coding-agent benchmark.
Internal full-stack coding-agent tasksProvider-reported scoreFull-stack software engineering
CurrentDisplay only
DSBench-FullStack 2026 · updated September 2, 2026
DSBench-Hard
DeepSeek DSBench Hard
DeepSeek's internal hard coding-agent benchmark.
Internal hard coding-agent tasksProvider-reported scoreAdvanced coding-agent challenges
CurrentDisplay only
DSBench-Hard 2026 · updated September 2, 2026
React Native Evals
React Native Evals
An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.
React Native app implementation tasksFramework-specific app development evaluationProduction mobile app engineering
CurrentDisplay only
React Native Evals 2026 · updated September 2, 2026
ReactBench
ReactBench v1
A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.
51 production React tasksPass@1 weighted rubric scoreProduction frontend engineering
CurrentDisplay only
ReactBench 2026 · updated September 2, 2026
KernelBench
KernelBench Hard H100
An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.
6 GPU-kernel optimization problemsMean peak fraction of hardware roofline over valid cellsAgentic GPU systems engineering
CurrentDisplay only
KernelBench 2026 · updated September 2, 2026
Next.js Evals
AI Agent Evaluations for Next.js
A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.
24 Next.js code generation and migration tasksAgent task completion with withheld Vitest assertionsFramework-specific web application engineering
CurrentDisplay only
Next.js Evals 2026 · updated September 2, 2026
SWE-bench Verified*
SWE-bench Verified (mini-swe-agent-v2)
A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.
Repository task completionAgent scaffold benchmarkProfessional software engineering
CurrentDisplay only
SWE-bench Verified* 2026 · updated September 2, 2026
Spider 2.0-Lite
Spider 2.0-Lite
A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison.
Text-to-SQL queriesExecution accuracyEnterprise text-to-SQL
RefreshingDisplay only
Spider 2.0-Lite 2024 · updated September 2, 2026
SciCode
Scientific Code Benchmark
SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.
80
RefreshingWeighted 16%
SciCode 2024 · updated September 2, 2026
AA Coding Index
Artificial Analysis Coding Index
A display-only Artificial Analysis coding index.
Cross-benchmark coding indexAggregated model scoreDisplay-only external reference
CurrentDisplay only
AA Coding Index 2026 · updated September 2, 2026
AA Coding Agents
Artificial Analysis Coding Agent Index
A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.
Composite over DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnAAverage pass@1 indexReal-world coding-agent workflows
CurrentDisplay only
AA Coding Agents 2026 · updated September 2, 2026
AA-SciCode
Artificial Analysis SciCode
A display-only Artificial Analysis SciCode score.
Scientific coding subproblemsTask success rateScientific programming
CurrentDisplay only
AA-SciCode 2026 · updated September 2, 2026
Terminal-Bench Hard
Terminal-Bench Hard
A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.
Agentic coding and terminal tasksTask success rateProfessional software engineering
CurrentDisplay only
Terminal-Bench Hard 2026 · updated September 2, 2026
VIBE V2
VIBE V2
A display-only MiniMax provider benchmark for end-to-end coding-agent and product-building tasks.
End-to-end coding-agent tasksTask success rateFrontier coding-agent workflows
CurrentDisplay only
VIBE V2 2026 · updated September 2, 2026
SVG-Bench
SVG-Bench
A display-only provider benchmark for generating or manipulating SVG outputs from natural-language requirements.
SVG generation and editing tasksTask success rateVisual coding and structured graphics generation
CurrentDisplay only
SVG-Bench 2026 · updated September 2, 2026
KernelBench Hard
KernelBench Hard
A display-only benchmark for difficult GPU kernel implementation and optimization tasks.
Hard GPU kernel coding tasksTask success rateSpecialized systems programming
CurrentDisplay only
KernelBench Hard 2026 · updated September 2, 2026
GameDevBench
GameDevBench
Evaluates coding agents on 333 multimodal game-development tasks in Godot, spanning 2D graphics, 3D graphics, user interfaces, and gameplay logic.
333 tasks from 88 tutorialsPass@1 on the full task set with 95% confidence intervalsMultimodal game development in Godot 4.4.1
CurrentDisplay only
GameDevBench 2026 · updated September 2, 2026
EdgeBench
EdgeBench
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
Systems and software-engineering tasksTime-budgeted agent learning curvesLong-horizon engineering
CurrentDisplay only
EdgeBench 2026 · updated September 2, 2026
Reasoning(29 benchmarks)
ARC-AGI-1
ARC-AGI-1 Semi-Private Evaluation
ARC Prize fluid-intelligence benchmark using novel visual grid transformations.
Semi-private ARC-AGI-1 evaluation setVerified accuracyAbstract visual reasoning
CurrentDisplay only
ARC-AGI-1 2026 · updated September 2, 2026
MuSR
Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.
Multi-step reasoningNarrative-based reasoningComplex reasoning tasks
StaleDisplay only
MuSR 2023 · updated September 2, 2026
BBH
BIG-Bench Hard
A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.
23 tasksMixed reasoning tasksAdvanced reasoning
StaleSaturatedDisplay only
BBH 2022 · updated September 2, 2026
DROP
Discrete Reasoning Over Paragraphs
A reading-comprehension benchmark requiring discrete reasoning over paragraphs, reported in DeepSeek-V4 base-model evaluations.
Paragraph reasoning questionsF1Reading and numerical reasoning
CurrentDisplay only
DROP 2026 · updated September 2, 2026
HellaSwag
HellaSwag
A commonsense natural-language inference benchmark reported in DeepSeek-V4 base-model evaluations.
Commonsense completion questionsExact matchCommonsense reasoning
CurrentDisplay only
HellaSwag 2026 · updated September 2, 2026
WinoGrande
WinoGrande
A commonsense coreference benchmark reported in DeepSeek-V4 base-model evaluations.
Coreference resolution questionsExact matchCommonsense reasoning
CurrentDisplay only
WinoGrande 2026 · updated September 2, 2026
CLUEWSC
CLUEWSC
A Chinese Winograd Schema Challenge benchmark reported in DeepSeek-V4 base-model evaluations.
Chinese coreference questionsExact matchChinese commonsense reasoning
CurrentDisplay only
CLUEWSC 2026 · updated September 2, 2026
LisanBench
LisanBench
A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.
50 starting words × 3 trialsDifficulty-weighted word-chain reasoningOpen-ended lexical planning
CurrentDisplay only
LisanBench 2026 · updated September 2, 2026
Conceptual Reasoning
Conceptual Reasoning Benchmark
Tests whether model judgments rank argumentative critiques in the same order as expert human ratings across philosophy, AI alignment, and other concept-heavy texts.
224 texts and 608 within-text critique pairsAverage pairwise-ranking loss against expert ratingsFuzzy, expert-rated argumentative reasoning
CurrentDisplay only
Early results snapshot · updated September 2, 2026
Pencil Puzzle Bench
Pencil Puzzle Bench
A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.
300 evaluation puzzlesDirect and agentic puzzle solve rateMulti-step verifiable reasoning
CurrentDisplay only
Pencil Puzzle Bench 2026 · updated September 2, 2026
LongBench v2
LongBench v2
A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.
Long-context tasksExtended-context retrieval and reasoningHard long-context
CurrentWeighted 38%
LongBench v2 2025 · updated September 2, 2026
MRCRv2
MRCRv2
A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.
Long-context retrievalMulti-round long-context evaluationHard long-context
CurrentWeighted 31%
MRCRv2 2025 · updated September 2, 2026
MRCR v2 64K-128K
OpenAI MRCR v2 8-needle 64K-128K
MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.
8-needle retrieval tasksLong-context retrievalLong-context reasoning
CurrentDisplay only
MRCR v2 64K-128K 2026 · updated September 2, 2026
MRCR v2 128K-256K
OpenAI MRCR v2 8-needle 128K-256K
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
8-needle retrieval tasksVery-long-context retrievalVery long-context reasoning
CurrentDisplay only
MRCR v2 128K-256K 2026 · updated September 2, 2026
MRCR v2 256K-512K
OpenAI MRCR v2 8-needle 256K-512K
MRCR v2 slice focused on retrieval across 256K-512K-token contexts.
100 eight-needle retrieval examplesMean sequence-matcher ratioVery long-context retrieval
CurrentDisplay only
MRCR v2 256K-512K 2026 · updated September 2, 2026
MRCR v2 512K-1M
OpenAI MRCR v2 8-needle 512K-1M
MRCR v2 slice focused on retrieval across 512K-1M-token contexts.
100 eight-needle retrieval examplesMean sequence-matcher ratioMillion-token retrieval
CurrentDisplay only
MRCR v2 512K-1M 2026 · updated September 2, 2026
Graphwalks BFS 128K
Graphwalks BFS 0K-128K
Long-context graph traversal benchmark using breadth-first search tasks.
Graph traversal tasksLong-context graph reasoningAlgorithmic long-context reasoning
CurrentDisplay only
Graphwalks BFS 128K 2026 · updated September 2, 2026
Graphwalks Parents 128K
Graphwalks parents 0-128K
Long-context benchmark for recovering parent relationships inside graph tasks.
Graph parent-retrieval tasksLong-context graph reasoningAlgorithmic long-context reasoning
CurrentDisplay only
Graphwalks Parents 128K 2026 · updated September 2, 2026
MRCR 1M
MRCR 1M
A million-token MRCR long-context retrieval benchmark reported in DeepSeek-V4 model evaluations.
Million-token retrievalLong-context retrieval MMRMillion-token long context
CurrentDisplay only
MRCR 1M 2026 · updated September 2, 2026
CorpusQA 1M
CorpusQA 1M
A million-token CorpusQA long-context question-answering benchmark reported in DeepSeek-V4 model evaluations.
Million-token corpus question answeringLong-context QA accuracyMillion-token long context
CurrentDisplay only
CorpusQA 1M 2026 · updated September 2, 2026
ARC-AGI-2
Abstraction and Reasoning Corpus for AGI v2
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
Visual pattern completion and abstract reasoningGrid transformation puzzles with novel rulesExpert-level — hardest public reasoning benchmark
CurrentWeighted 31%
ARC-AGI 2 · updated September 2, 2026
ARC-AGI-3
Abstraction and Reasoning Corpus for AGI v3
An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.
Interactive game-like tasks with hidden rulesAgentic task completion under a capped evaluation budgetFrontier agentic reasoning
CurrentDisplay only
ARC-AGI 3 · updated September 2, 2026
GeneBench-Pro
GeneBench-Pro
A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.
129 genomics statistical-analysis workflowsEval-level pass rate across dependent analysis decisionsLong-horizon scientific reasoning
CurrentDisplay only
GeneBench-Pro · updated September 2, 2026
AI-Needle
AI-Needle
A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.
Long-context retrievalNeedle-in-a-haystack recallLong-context memory
CurrentDisplay only
AI-Needle 2026 · updated September 2, 2026
GPQA Diamond
GPQA Diamond
The hardest subset of GPQA featuring the most challenging graduate-level science questions. Sometimes reported separately from the standard GPQA benchmark.
Expert-level science questionsMultiple choice questionsGraduate-level scientific reasoning
StaleDisplay only
GPQA Diamond 2023 · updated September 2, 2026
AA-LCR
Artificial Analysis Long Context Reasoning
A display-only Artificial Analysis long-context reasoning evaluation.
Long-context reasoning tasksAccuracyLong-context reasoning
CurrentDisplay only
AA-LCR 2026 · updated September 2, 2026
CritPt
Critical Physics Tasks
A display-only Artificial Analysis metric for research-level physics reasoning.
Research-level physics questionsAccuracyResearch-level physics reasoning
CurrentDisplay only
CritPt 2026 · updated September 2, 2026
BullshitBench v2
BullshitBench v2
A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.
Nonsensical and flawed prompts across multiple domainsPrompt challenge and refusal evaluationRobustness and critical reasoning
CurrentDisplay only
BullshitBench v2 2025 · updated September 2, 2026
WildBench
WildBench
An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.
1,024 real-world tasksReal-world task evaluationDiverse real-world scenarios
RefreshingDisplay only
WildBench 2024 · updated September 2, 2026
Multimodal & Grounded(65 benchmarks)
Chartography (no tools)
Chartography without tools
Professional chart understanding across 100 specialized chart types with expert-set answer tolerances.
100 specialized chart typesAccuracy without toolsProfessional chart reasoning
CurrentDisplay only
Chartography (no tools) 2026 · updated September 2, 2026
Chartography (tools)
Chartography with image and code tools
Professional chart understanding with a container, standard libraries, and image cropping.
100 specialized chart typesAccuracy with toolsProfessional chart reasoning
CurrentDisplay only
Chartography (tools) 2026 · updated September 2, 2026
BenchCAD Vision2Code (no tools)
BenchCAD Vision2Code voxel IoU without tools
Generates CadQuery code from multi-view renders and scores geometric similarity by voxel intersection-over-union.
1,000-file Vision2Code subsetVoxel IoUProgrammatic CAD generation
CurrentDisplay only
BenchCAD Vision2Code (no tools) 2026 · updated September 2, 2026
BenchCAD Vision2Code (tools)
BenchCAD Vision2Code voxel IoU with tools
Generates CadQuery code from multi-view renders with image inspection and code-execution tools.
1,000-file Vision2Code subsetVoxel IoU with toolsProgrammatic CAD generation
CurrentDisplay only
BenchCAD Vision2Code (tools) 2026 · updated September 2, 2026
GDP.pdf (no tools)
GDP.pdf mean criteria pass rate without tools
Professional document understanding over 100 real-world PDFs from ten domains.
100 professional document promptsMean criteria pass rateProfessional document reasoning
CurrentDisplay only
GDP.pdf (no tools) 2026 · updated September 2, 2026
GDP.pdf (tools)
GDP.pdf mean criteria pass rate with tools
Professional document understanding with a container, standard libraries, and image cropping.
100 professional document promptsMean criteria pass rate with toolsProfessional document reasoning
CurrentDisplay only
GDP.pdf (tools) 2026 · updated September 2, 2026
OfficeQA
OfficeQA
Grounded numerical reasoning over a corpus of historical U.S. Treasury Bulletin documents.
Historical Treasury Bulletin questionsAgentic grounded QA accuracyProfessional document reasoning
CurrentDisplay only
OfficeQA 2026 · updated September 2, 2026
MMMU
Massive Multi-discipline Multimodal Understanding
A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.
Multimodal academic reasoningImage + text question answeringFrontier multimodal
RefreshingDisplay only
MMMU 2024 · updated September 2, 2026
MMMU-Pro
Massive Multi-discipline Multimodal Understanding Pro
A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.
Multimodal academic reasoningImage + text question answeringFrontier multimodal
RefreshingWeighted 45%
MMMU-Pro 2024 · updated September 2, 2026
AA-MMMU-Pro
Artificial Analysis MMMU-Pro
A display-only Artificial Analysis MMMU-Pro score.
Multimodal academic reasoningImage + text question answeringFrontier multimodal
CurrentDisplay only
AA-MMMU-Pro 2026 · updated September 2, 2026
OCRBench V2
OCRBench V2
A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.
Image OCR tasksAccuracyNative visual text understanding
CurrentDisplay only
OCRBench V2 2025 · updated September 2, 2026
olmOCR
olmOCR-Bench
An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.
Layout-rich PDF understandingMean accuracyComplex document processing
CurrentDisplay only
olmOCR 2025 · updated September 2, 2026
VoxPopuli WER
VoxPopuli-Cleaned-AA Word Error Rate
A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.
Speech-to-text transcriptionWord error rateAudio speech recognition
CurrentDisplay only
VoxPopuli WER 2026 · updated September 2, 2026
Design Arena Website
Design Arena Website Elo
A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.
Website generation comparisonsEloDesign and website generation
CurrentDisplay only
Design Arena Website 2026 · updated September 2, 2026
OfficeQA Pro
OfficeQA Pro
A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.
Document and spreadsheet tasksGrounded QA over office artifactsEnterprise grounded reasoning
CurrentWeighted 30%
OfficeQA Pro 2026 · updated September 2, 2026
MathVision w/ Python
MathVision with Python
A tool-augmented MathVision variant that permits Python during visual mathematics reasoning.
Visual mathematics problems with PythonImage and mathematics reasoning with toolsAdvanced multimodal mathematics
CurrentDisplay only
MathVision w/ Python 2026 · updated September 2, 2026
BabyVision w/ Python
BabyVision with Python
A Python-assisted BabyVision evaluation for fine-grained visual perception and grounded reasoning.
Visual perception tasks with PythonTool-augmented multimodal scoreFine-grained visual perception
CurrentDisplay only
BabyVision w/ Python 2026 · updated September 2, 2026
ZeroBench w/ Python
ZeroBench_main with Python
A Python-assisted ZeroBench_main evaluation reported as pass@5.
Visual reasoning questions with PythonPass@5Tool-augmented visual reasoning
CurrentDisplay only
ZeroBench w/ Python 2026 · updated September 2, 2026
WorldVQA ForceAnswer
WorldVQA ForceAnswer
A forced-answer WorldVQA variant for atomic visual world knowledge.
Atomic visual world-knowledge questionsForced-answer visual QAFine-grained visual knowledge
CurrentDisplay only
WorldVQA ForceAnswer 2026 · updated September 2, 2026
OmniDocBench
OmniDocBench
A document-understanding benchmark for parsing and reasoning over complex document layouts.
Complex document-understanding tasksDocument-understanding scoreGrounded document reasoning
CurrentDisplay only
OmniDocBench 2026 · updated September 2, 2026
PerceptionBench
PerceptionBench (Internal)
Moonshot AI's internal benchmark for atomic visual perception capabilities.
Internal atomic visual-perception tasksInternal evaluation scoreFine-grained visual perception
CurrentDisplay only
PerceptionBench 2026 · updated September 2, 2026
MMMU-Pro w/ Python
MMMU-Pro with Python
Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning.
Multimodal academic reasoningImage + text question answering with PythonFrontier multimodal
CurrentDisplay only
MMMU-Pro w/ Python 2026 · updated September 2, 2026
OmniDocBench 1.5
OmniDocBench 1.5
A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.
Document understanding tasksDocument understanding benchmarkGrounded document reasoning
CurrentDisplay only
OmniDocBench 1.5 2026 · updated September 2, 2026
Liquid Extract JSON Validity
Liquid image-to-JSON extraction JSON validity
A display-only Liquid AI extraction metric measuring the share of image-to-JSON outputs that parse as strict JSON.
Image-to-JSON extractionStrict JSON parseability rateStructured visual extraction
CurrentDisplay only
Liquid Extract JSON Validity 2026 · updated September 2, 2026
Liquid Extract F1
Liquid image-to-JSON extraction schema consistency F1
A display-only Liquid AI extraction metric measuring field-name agreement between requested schema fields and extracted JSON fields.
Image-to-JSON extractionSchema field F1Structured visual extraction
CurrentDisplay only
Liquid Extract F1 2026 · updated September 2, 2026
Liquid Extract VLM Judge
Liquid image-to-JSON extraction VLM judge score
A display-only Liquid AI extraction metric measuring judged agreement between extracted values and the source image.
Image-to-JSON extractionVLM-judged extraction accuracyStructured visual extraction
CurrentDisplay only
Liquid Extract VLM Judge 2026 · updated September 2, 2026
RealWorldQA
RealWorldQA
A grounded visual QA benchmark focused on answering practical questions about real-world images and scenes.
Real-world visual question answeringImage-grounded QAGeneral visual reasoning
CurrentDisplay only
RealWorldQA 2026 · updated September 2, 2026
Video-MME (with subtitle)
Video-MME with subtitle
A video understanding benchmark that allows subtitle access when answering multimodal questions about videos.
Video understandingVideo QA with subtitle contextMultimodal video reasoning
CurrentDisplay only
Video-MME (with subtitle) 2026 · updated September 2, 2026
Video-MME (w/o subtitle)
Video-MME without subtitle
A stricter Video-MME setting that removes subtitle help and tests video understanding from visual and audio context alone.
Video understandingVideo QA without subtitle contextMultimodal video reasoning
CurrentDisplay only
Video-MME (w/o subtitle) 2026 · updated September 2, 2026
Video-MME
Video-MME
A comprehensive benchmark for multimodal large language models on video understanding, covering temporal reasoning, perception, and question answering over videos.
Video understandingVideo QA and analysisBroad multimodal video reasoning
RefreshingDisplay only
Video-MME 2024 · updated September 2, 2026
MathVision
MathVision
A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.
Visually grounded math problemsImage + math reasoningAdvanced multimodal mathematics
CurrentDisplay only
MathVision 2026 · updated September 2, 2026
We-Math
We-Math
A multimodal math benchmark for visually grounded mathematical reasoning and answer generation.
Visually grounded math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics
CurrentDisplay only
We-Math 2026 · updated September 2, 2026
DynaMath
DynaMath
A multimodal benchmark for dynamic mathematical reasoning over visual and structured inputs.
Dynamic visual math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics
CurrentDisplay only
DynaMath 2026 · updated September 2, 2026
MStar
MStar
A general visual question-answering benchmark used in provider tables for real-image reasoning quality.
Real-image visual QAImage-grounded QAGeneral visual reasoning
CurrentDisplay only
MStar 2026 · updated September 2, 2026
ChatCVQA
ChatCVQA
A conversational visual QA benchmark that tests multi-turn grounded answering over images and documents.
Conversational visual QAMulti-turn image-grounded QAConversational multimodal reasoning
CurrentDisplay only
ChatCVQA 2026 · updated September 2, 2026
MMLongBench-Doc
MMLongBench-Doc
A long-document multimodal benchmark for grounded reasoning over extended document contexts.
Long document understandingDocument-grounded reasoningLong-context document reasoning
CurrentDisplay only
MMLongBench-Doc 2026 · updated September 2, 2026
CC-OCR
CC-OCR
An OCR-focused benchmark for reading and extracting text from visually complex documents and images.
Optical character recognitionText extraction from images and documentsDocument reading
CurrentDisplay only
CC-OCR 2026 · updated September 2, 2026
AI2D_TEST
AI2D test split
A diagram understanding benchmark focused on scientific and educational visual question answering.
Diagram understandingDiagram-grounded QAStructured visual reasoning
CurrentDisplay only
AI2D_TEST 2026 · updated September 2, 2026
CountBench
CountBench
A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.
Visual counting tasksImage-grounded countingFine-grained visual perception
CurrentDisplay only
CountBench 2026 · updated September 2, 2026
RefCOCO (avg)
RefCOCO average
A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.
Referring-expression groundingGrounded visual localizationFine-grained visual grounding
CurrentDisplay only
RefCOCO (avg) 2026 · updated September 2, 2026
ODINW13
ODINW13
A visual detection and grounding benchmark slice used to compare zero-shot object understanding across diverse domains.
Out-of-distribution object understandingDetection and groundingRobust visual grounding
CurrentDisplay only
ODINW13 2026 · updated September 2, 2026
ERQA
ERQA
A grounded visual reasoning benchmark focused on evidence-based question answering over real images.
Evidence-based visual QAGrounded image reasoningGrounded multimodal reasoning
CurrentDisplay only
ERQA 2026 · updated September 2, 2026
VideoMMMU
VideoMMMU
A video extension of MMMU-style multimodal reasoning over expert questions grounded in temporal media.
Video-grounded expert reasoningVideo + text reasoningFrontier multimodal video reasoning
CurrentDisplay only
VideoMMMU 2026 · updated September 2, 2026
MLVU (M-Avg)
MLVU mean average
A multi-task video understanding benchmark averaged across MLVU categories.
General video understandingVideo QA and understandingBroad multimodal video reasoning
CurrentDisplay only
MLVU (M-Avg) 2026 · updated September 2, 2026
LVBench
LVBench
A long-video understanding benchmark for retrieving and reasoning over information distributed across extended video inputs.
Long-form video question answeringLong-video understanding scoreExtended temporal reasoning
CurrentDisplay only
LVBench 2026 · updated September 2, 2026
MMVU
Multimodal Multi-disciplinary Video Understanding
A benchmark for evaluating multimodal models on video understanding tasks across multiple disciplines, emphasizing temporal reasoning and comprehension over video content.
Video understandingVideo reasoning benchmarkMulti-disciplinary multimodal video reasoning
CurrentDisplay only
MMVU 2026 · updated September 2, 2026
ScreenSpot Pro
ScreenSpot Pro
A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.
1,581 grounding instructionsStatic interface element localizationProfessional GUI grounding
CurrentDisplay only
ScreenSpot Pro 2025 · updated September 2, 2026
TIR-Bench
TIR-Bench
A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.
Visual agent and interface reasoningScreenshot-grounded task reasoningComputer-use visual reasoning
CurrentDisplay only
TIR-Bench 2026 · updated September 2, 2026
GDPval-AA
GDPval-AA
An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.
Professional office deliveryELO-style office benchmarkProfessional knowledge work
CurrentDisplay only
GDPval-AA 2026 · updated September 2, 2026
MedXpertQA (MM)
MedXpertQA Multimodal
A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.
2,000 multimodal medical questionsMedical visual MCQClinical multimodal reasoning
CurrentDisplay only
MedXpertQA (MM) 2026 · updated September 2, 2026
ZeroBench
ZeroBench
A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.
100 visual reasoning questionsMulti-step visual reasoningTool-augmented visual reasoning
CurrentDisplay only
ZeroBench 2026 · updated September 2, 2026
Design2Code
Design2Code
A multimodal coding benchmark for turning visual designs into working frontend implementations.
Design-to-code tasksVisual input to frontend implementationMultimodal coding
CurrentDisplay only
Design2Code 2026 · updated September 2, 2026
Flame-VLM-Code
Flame-VLM-Code
A vision-language coding benchmark for generating correct code from visual and multimodal inputs.
Multimodal coding tasksVision-language code generationMultimodal coding
CurrentDisplay only
Flame-VLM-Code 2026 · updated September 2, 2026
Vision2Web
Vision2Web
A benchmark for converting visual references into functional web implementations.
Screenshot-to-web tasksVisual reference to web implementationMultimodal web generation
CurrentDisplay only
Vision2Web 2026 · updated September 2, 2026
ImageMining
ImageMining
A multimodal retrieval and extraction benchmark over image-heavy task settings.
Visual retrieval tasksImage-grounded retrieval and extractionMultimodal retrieval
CurrentDisplay only
ImageMining 2026 · updated September 2, 2026
MMSearch
MMSearch
A multimodal search benchmark for retrieval and grounded answering across mixed-media inputs.
Multimodal search tasksMixed-media retrieval and grounded answeringMultimodal search
CurrentDisplay only
MMSearch 2026 · updated September 2, 2026
MMSearch-Plus
MMSearch-Plus
A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.
Hard multimodal search tasksAdvanced mixed-media retrieval benchmarkAdvanced multimodal search
CurrentDisplay only
MMSearch-Plus 2026 · updated September 2, 2026
SimpleVQA
SimpleVQA
A visual question answering benchmark focused on straightforward image-grounded understanding.
Visual QA tasksImage-grounded question answeringGeneral visual understanding
CurrentDisplay only
SimpleVQA 2026 · updated September 2, 2026
Facts-VLM
Facts-VLM
A grounded multimodal factuality benchmark for evidence-linked answer correctness.
Grounded factuality tasksEvidence-linked multimodal factualityGrounded multimodal factuality
CurrentDisplay only
Facts-VLM 2026 · updated September 2, 2026
V*
V*
A vision-centric benchmark for high-level multimodal reasoning and perception quality.
Frontier multimodal reasoning tasksVision-centric reasoning benchmarkFrontier multimodal
CurrentDisplay only
V* 2026 · updated September 2, 2026
CharXiv
CharXiv Reasoning
A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.
Scientific chart reasoningChart understanding and reasoningScientific visualization reasoning
RefreshingWeighted 25%
CharXiv 2024 · updated September 2, 2026
CharXiv w/o tools
CharXiv Reasoning without tools
Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.
Scientific chart reasoning (tool-free)Chart understanding without toolsScientific visualization reasoning
RefreshingDisplay only
CharXiv w/o tools 2024 · updated September 2, 2026
BabyVision
BabyVision
A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.
Visual perception tasksMultimodal visual reasoningFine-grained visual perception
CurrentDisplay only
BabyVision 2026 · updated September 2, 2026
SWE-bench Multimodal
SWE-bench Multimodal
A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.
Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding
CurrentDisplay only
SWE-bench Multimodal 2025 · updated September 2, 2026
Blueprint-Bench 2
Blueprint-Bench 2
An agentic spatial reasoning benchmark reported as a normalized score.
Spatial reasoning from blueprintsNormalized scoreAgentic spatial reasoning
CurrentDisplay only
Blueprint-Bench 2 2026 · updated September 2, 2026
Knowledge(48 benchmarks)
HealthBench (raw)
HealthBench raw score
Raw score on realistic multi-turn healthcare conversations graded against expert-written rubrics.
5,000 multi-turn patient conversationsRaw rubric scoreRealistic healthcare conversations
CurrentDisplay only
HealthBench (raw) 2026 · updated September 2, 2026
HealthBench (length-adjusted)
HealthBench length-adjusted score
HealthBench score after applying a verbosity penalty to model responses.
5,000 multi-turn patient conversationsLength-adjusted rubric scoreRealistic healthcare conversations
CurrentDisplay only
HealthBench (length-adjusted) 2026 · updated September 2, 2026
HealthBench Professional (raw)
HealthBench Professional raw score
Raw score on physician-authored clinical consult, documentation, and research conversations.
525 physician-authored conversationsRaw rubric scoreProfessional clinical tasks
CurrentDisplay only
HealthBench Professional (raw) 2026 · updated September 2, 2026
BioMysteryBench (human-solvable)
BioMysteryBench Human Solvable
Computational biology challenges that independent human experts were able to solve.
Human-solvable computational biology investigationsTask scoreExpert computational biology
CurrentDisplay only
BioMysteryBench (human-solvable) 2026 · updated September 2, 2026
BioMysteryBench (human-difficult)
BioMysteryBench Human Difficult
Computational biology challenges with objective answers that remained unsolved by independent human experts.
Human-difficult computational biology investigationsTask scoreFrontier computational biology
CurrentDisplay only
BioMysteryBench (human-difficult) 2026 · updated September 2, 2026
SpatialBench Verified
LatchBio SpatialBench Verified
Analysis of spatial transcriptomics data across externally validated biological problems.
115 externally validated spatial transcriptomics problemsTask scoreProfessional bioinformatics
CurrentDisplay only
SpatialBench Verified 2026 · updated September 2, 2026
SingleCellBench
LatchBio SingleCellBench
Single-cell RNA sequencing analysis tasks spanning common bioinformatics workflows.
195 single-cell RNA sequencing problemsTask scoreProfessional bioinformatics
CurrentDisplay only
SingleCellBench 2026 · updated September 2, 2026
ProteinGym Hard
ProteinGym Hard
Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.
Hard protein mutation-effect ranking tasksRank correlationComputational protein science
CurrentDisplay only
ProteinGym Hard 2026 · updated September 2, 2026
Protein Design
Anthropic Protein Design evaluation
Generates novel protein sequences under family, topology, globularity, and structural-motif constraints.
Constrained protein-sequence design tasksComposite scoreComputational protein design
CurrentDisplay only
Protein Design 2026 · updated September 2, 2026
Organic chemistry V2
Anthropic Organic Chemistry V2 evaluation
Chemistry tasks covering spectroscopy, synthesis planning, reaction prediction, and chemical structure images.
Organic chemistry reasoning tasksTask scoreExpert organic chemistry
CurrentDisplay only
Organic chemistry V2 2026 · updated September 2, 2026
Protocols (troubleshooting)
Molecular Biology Protocols Troubleshooting
Detects and fixes errors in molecular-biology protocols using document, code, and web-search tools.
Molecular-biology protocol troubleshootingTask scoreExpert laboratory protocols
CurrentDisplay only
Protocols (troubleshooting) 2026 · updated September 2, 2026
Protocols (understanding)
Benchling Molecular Biology Protocols Understanding
Extends online molecular-biology protocols in additional directions using document and web-search tools.
Molecular-biology protocol extensionTask scoreExpert laboratory protocols
CurrentDisplay only
Protocols (understanding) 2026 · updated September 2, 2026
AA Openness Index
Artificial Analysis Openness Index
A display-only Artificial Analysis model-openness index.
Model openness assessmentIndex scoreDisplay-only external reference
CurrentDisplay only
AA Openness Index 2026 · updated September 2, 2026
AA MMLU-Pro
Artificial Analysis MMLU-Pro
An independently evaluated MMLU-Pro result from Artificial Analysis.
Professional multi-subject questionsAccuracyProfessional knowledge and reasoning
CurrentDisplay only
AA MMLU-Pro 2026 · updated September 2, 2026
MMLU
Massive Multitask Language Understanding
A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.
57 subjectsMultiple choice questionsElementary to professional level
StaleSaturatedDisplay only
MMLU · updated September 2, 2026
GPQA
Graduate-Level Google-Proof Q&A
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
448 questionsMultiple choice questionsGraduate level
RefreshingWeighted 7%
GPQA Diamond · updated September 2, 2026
GPQA-D
GPQA Diamond
A display-only GPQA Diamond reference from provider comparison charts.
Graduate-level science questionsMultiple choice questionsGraduate level
CurrentDisplay only
GPQA-D 2026 · updated September 2, 2026
SuperGPQA
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.
285 disciplinesMultiple choice questionsGraduate level
CurrentWeighted 7%
SuperGPQA 2025 · updated September 2, 2026
MMLU-Pro
Massive Multitask Language Understanding Professional
An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.
Multiple subjects10-way multiple choiceProfessional level
RefreshingWeighted 30%
MMLU-Pro · updated September 2, 2026
AGIEval
AGIEval
A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.
General academic and professional exam questionsExact matchGeneral knowledge
CurrentDisplay only
AGIEval 2026 · updated September 2, 2026
HLE
Humanity's Last Exam
An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.
Expert-level questionsOpen-ended and multiple choiceFrontier expert level
CurrentWeighted 45%
Humanity's Last Exam · updated September 2, 2026
HLE-Verified
HLE-Verified
A verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions before evaluating frontier expert reasoning.
1,811 verified or revised expert questionsFull-set accuracyFrontier multidisciplinary expert reasoning
CurrentDisplay only
HLE-Verified 2026 · updated September 2, 2026
LABBench2
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
A benchmark of realistic biology-research tasks involving literature, figures, tables, databases, and bioinformatics files.
Nearly 1,900 biology-research tasksAggregate accuracyReal-world biology research
CurrentDisplay only
LABBench2 2026 · updated September 2, 2026
FrontierScience
FrontierScience
A benchmark for research-level scientific reasoning, designed to separate frontier models on difficult science tasks that mix domain knowledge with deep reasoning.
Research-level science tasksScientific reasoning benchmarkResearch frontier
CurrentDisplay only
FrontierScience 2026 · updated September 2, 2026
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index
A display-only intelligence index published by Artificial Analysis that aggregates provider-reported and benchmark-derived signals into a single model-level score.
Cross-benchmark intelligence indexAggregated model scoreDisplay-only external reference
CurrentDisplay only
Artificial Analysis Intelligence Index 2026 · updated September 2, 2026
AA-GPQA Diamond
Artificial Analysis GPQA Diamond
A display-only Artificial Analysis GPQA Diamond score.
Graduate-level science questionsAccuracyGraduate-level science reasoning
CurrentDisplay only
AA-GPQA Diamond 2026 · updated September 2, 2026
AA-HLE
Artificial Analysis Humanity's Last Exam
A display-only Artificial Analysis Humanity's Last Exam score.
Expert-level questionsAccuracyFrontier expert reasoning
CurrentDisplay only
AA-HLE 2026 · updated September 2, 2026
AA-Omniscience Index
Artificial Analysis Omniscience Index
A display-only Artificial Analysis factual knowledge index.
Knowledge questionsIndex scoreBroad factual knowledge
CurrentDisplay only
AA-Omniscience Index 2026 · updated September 2, 2026
AA-Omniscience Accuracy
Artificial Analysis Omniscience Accuracy
A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.
Knowledge questionsAccuracyBroad knowledge
CurrentDisplay only
AA-Omniscience Accuracy 2026 · updated September 2, 2026
AA-Omniscience Hallucination Rate
Artificial Analysis Omniscience Hallucination Rate
A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.
Knowledge questionsHallucination rateFactuality
CurrentDisplay only
AA-Omniscience Hallucination Rate 2026 · updated September 2, 2026
SimpleQA
Measuring Short-Form Factuality in Large Language Models
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
Factual questionsShort-form Q&AFactual accuracy focused
RefreshingWeighted 11%
SimpleQA 2024 · updated September 2, 2026
Chinese-SimpleQA
Chinese-SimpleQA
A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.
Chinese factual questionsShort-form factual QAFactual accuracy focused
CurrentDisplay only
Chinese-SimpleQA 2026 · updated September 2, 2026
OpenBookQA
OpenBookQA
A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.
Elementary science questions4-way multiple choiceElementary science reasoning
StaleDisplay only
OpenBookQA 2018 · updated September 2, 2026
HealthBench Hard
HealthBench Hard
A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.
1,000 health promptsOpen-ended health evaluationAdvanced health reasoning
CurrentDisplay only
HealthBench Hard 2026 · updated September 2, 2026
HealthBench Professional
HealthBench Professional
An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.
Clinician chat tasksRubric-graded open-ended responsesProfessional clinical workflows
CurrentDisplay only
HealthBench Professional 2026 · updated September 2, 2026
MedXpertQA (Text)
MedXpertQA Text
A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.
2,450 medical multiple-choice questionsMedical MCQProfessional medical knowledge
CurrentDisplay only
MedXpertQA (Text) 2026 · updated September 2, 2026
FrontierScience Research
FrontierScience Research
A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.
Scientific research problemsResearch evaluationFrontier scientific research
CurrentDisplay only
FrontierScience Research 2026 · updated September 2, 2026
TruthfulQA
TruthfulQA
A benchmark designed to measure whether language models produce truthful answers instead of repeating common misconceptions or misleading falsehoods.
Truthfulness and misconception resistanceQuestion answeringHallucination and factuality stress test
StaleDisplay only
TruthfulQA 2021 · updated September 2, 2026
HLE w/o tools
Humanity's Last Exam without tools
Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.
Expert-level questionsTool-free expert QAFrontier expert level
CurrentDisplay only
HLE w/o tools 2026 · updated September 2, 2026
MMLU-Pro (Arcee)
MMLU-Pro first-party comparison snapshot
A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.
Professional academic QA10-way multiple choiceProfessional level
CurrentDisplay only
MMLU-Pro (Arcee) 2026 · updated September 2, 2026
MMLU-Redux
MMLU-Redux
A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.
Broad academic QAMultiple choice questionsAdvanced general knowledge
CurrentDisplay only
MMLU-Redux 2026 · updated September 2, 2026
MMMLU
MMMLU
A multilingual MMLU-style benchmark reported in provider evaluation tables.
Multilingual academic QAExact matchBroad multilingual knowledge
CurrentDisplay only
MMMLU 2026 · updated September 2, 2026
C-Eval
C-Eval
A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.
Chinese academic and professional examsMultiple choice questionsHigh school to professional level
StaleDisplay only
C-Eval 2023 · updated September 2, 2026
CMMLU
Chinese Massive Multitask Language Understanding
A Chinese multitask academic benchmark reported in DeepSeek-V4 base-model evaluations.
Chinese academic QAExact matchBroad Chinese knowledge
CurrentDisplay only
CMMLU 2026 · updated September 2, 2026
MultiLoKo
MultiLoKo
A multilingual/localized knowledge benchmark reported in DeepSeek-V4 base-model evaluations.
Localized multilingual knowledge questionsExact matchMultilingual knowledge
CurrentDisplay only
MultiLoKo 2026 · updated September 2, 2026
FACTS Parametric
FACTS Parametric
A parametric factuality benchmark reported in DeepSeek-V4 base-model evaluations.
Parametric factual recallExact matchFactual accuracy focused
CurrentDisplay only
FACTS Parametric 2026 · updated September 2, 2026
TriviaQA
TriviaQA
A reading and trivia question-answering benchmark reported in DeepSeek-V4 base-model evaluations.
Trivia and reading-comprehension QAExact matchGeneral factual QA
CurrentDisplay only
TriviaQA 2026 · updated September 2, 2026
FinanceArena
FinanceArena — FinanceQA Assumption-Based
An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.
Professional financial-analysis questionsOpen-ended financial QA with exact-match gradingProfessional finance reasoning
CurrentDisplay only
FinanceArena 2025 · updated September 2, 2026
Multilingual(13 benchmarks)
GMMLU
Global MMLU
MMLU-style knowledge evaluation across 42 high- and low-resource languages.
Knowledge questions across 42 languagesAverage accuracyMultilingual knowledge
RefreshingDisplay only
GMMLU 2024 · updated September 2, 2026
MILU
Multi-task Indic Language Understanding Benchmark
Culturally grounded knowledge comprehension across ten Indic languages and English.
Knowledge tasks across 11 languagesAverage accuracyMultilingual Indic knowledge
RefreshingDisplay only
MILU 2024 · updated September 2, 2026
AA Global-MMLU-Lite
Artificial Analysis Global-MMLU-Lite
An independently evaluated multilingual knowledge result from Artificial Analysis.
Multilingual knowledge questionsAccuracyMultilingual professional knowledge
CurrentDisplay only
AA Global-MMLU-Lite 2026 · updated September 2, 2026
MGSM
Multilingual Grade School Math
A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.
250 problems × 11 languagesMath word problemsGrade school math, multilingual
StaleDisplay only
MGSM 2022 · updated September 2, 2026
MMLU-ProX
MMLU-ProX
A multilingual extension of professional-level academic evaluation across many languages.
Multilingual professional QAMultilingual multiple choiceProfessional multilingual
CurrentWeighted 100%
MMLU-ProX 2025 · updated September 2, 2026
NOVA-63
NOVA-63
A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.
Broad multilingual evaluationCross-lingual benchmarkBroad multilingual capability
CurrentDisplay only
NOVA-63 2026 · updated September 2, 2026
INCLUDE
INCLUDE
A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.
Cross-lingual understandingMultilingual benchmarkBroad multilingual capability
CurrentDisplay only
INCLUDE 2026 · updated September 2, 2026
PolyMath
PolyMath
A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.
Multilingual math problemsCross-lingual mathematical reasoningAdvanced multilingual reasoning
CurrentDisplay only
PolyMath 2026 · updated September 2, 2026
VWT2k-lite
VWT2k-lite
A lighter multilingual benchmark slice published in provider tables for broad cross-lingual transfer and understanding.
Multilingual transfer tasksCross-lingual benchmarkBroad multilingual capability
CurrentDisplay only
VWT2k-lite 2026 · updated September 2, 2026
MAXIFE
MAXIFE
A multilingual instruction-following and understanding benchmark row published in Qwen's launch comparisons.
Multilingual instruction followingCross-lingual benchmarkAdvanced multilingual instruction following
CurrentDisplay only
MAXIFE 2026 · updated September 2, 2026
SWE Multilingual
SWE-bench Multilingual
A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.
300 problems across 9 languagesMulti-language code patch generationProfessional multilingual software engineering
CurrentDisplay only
SWE Multilingual 2025 · updated September 2, 2026
NanoBEIR Multilingual
NanoBEIR Multilingual Extended
A display-only multilingual retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using NDCG@10 across 11 languages.
Multilingual document retrievalNDCG@10 averageMultilingual retrieval
CurrentDisplay only
NanoBEIR Multilingual 2026 · updated September 2, 2026
MKQA-11
MKQA-11 multilingual retrieval
A display-only multilingual QA retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using Recall@20 across 11 languages.
Cross-lingual open-domain QA retrievalRecall@20 averageMultilingual retrieval
CurrentDisplay only
MKQA-11 2026 · updated September 2, 2026
Instruction Following(4 benchmarks)
IFEval
Instruction-Following Eval
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
541 prompts across 25 instruction typesConstrained generationInstruction precision
StaleWeighted 35%
IFEval 2023 · updated September 2, 2026
IFBench
Instruction Following Benchmark
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
58
CurrentWeighted 65%
IFBench 2025 · updated September 2, 2026
AA-IFBench
Artificial Analysis IFBench
A display-only Artificial Analysis IFBench score.
Verifiable instruction constraintsConstraint satisfaction accuracyInstruction precision
CurrentDisplay only
AA-IFBench 2026 · updated September 2, 2026
SOB Value Acc
Structured Output Benchmark Value Accuracy
A structured-output benchmark from Interfaze measuring whether extracted JSON leaf values exactly match verified ground truth.
Structured output extractionValue accuracyProduction structured-output reliability
CurrentDisplay only
SOB Value Acc 2026 · updated September 2, 2026
Mathematics(32 benchmarks)
IMO 2026
International Mathematical Olympiad 2026
Proof-based olympiad performance on all six IMO 2026 problems.
6 proof-based problemsOfficial-style proof scoreInternational olympiad mathematics
CurrentDisplay only
IMO 2026 2026 · updated September 2, 2026
RiemannBench (no tools)
RiemannBench without tools
Research-level mathematics problems with unique programmatically verified closed-form answers.
25 private research-level mathematics problemsAccuracy without toolsResearch mathematics
CurrentDisplay only
RiemannBench (no tools) 2026 · updated September 2, 2026
RiemannBench (tools)
RiemannBench with tools
Research-level mathematics problems with unique programmatically verified closed-form answers and tool access.
25 private research-level mathematics problemsAccuracy with toolsResearch mathematics
CurrentDisplay only
RiemannBench (tools) 2026 · updated September 2, 2026
ArXivMath Jun. 2026 (no tools)
ArXivMath June 2026 without tools
Final-answer research mathematics problems drawn from recent arXiv abstracts.
49 recent research-mathematics problemsFinal-answer accuracyResearch mathematics
CurrentDisplay only
ArXivMath Jun. 2026 (no tools) 2026 · updated September 2, 2026
ArXivMath Jun. 2026 (tools)
ArXivMath June 2026 with tools
Final-answer research mathematics problems drawn from recent arXiv abstracts with tool access.
49 recent research-mathematics problemsFinal-answer accuracy with toolsResearch mathematics
CurrentDisplay only
ArXivMath Jun. 2026 (tools) 2026 · updated September 2, 2026
AA AIME 2025
Artificial Analysis AIME 2025
An independently evaluated AIME 2025 result from Artificial Analysis.
30 AIME 2025 problemsAccuracyOlympiad mathematics
CurrentDisplay only
AA AIME 2025 2026 · updated September 2, 2026
AA MATH-500
Artificial Analysis MATH-500
An independently evaluated MATH-500 result from Artificial Analysis.
500 competition mathematics problemsAccuracyHigh school to undergraduate mathematics
CurrentDisplay only
AA MATH-500 2026 · updated September 2, 2026
AIME 2023
American Invitational Mathematics Examination 2023
A 15-question, 3-hour examination where each answer is an integer from 000 to 999. Serves as the intermediate step between AMC 10/12 and the USA Mathematical Olympiad (USAMO).
15 problemsInteger answers 000-999High school olympiad level
StaleDisplay only
AIME 2023 2023 · updated September 2, 2026
AIME 2024
American Invitational Mathematics Examination 2024
The 2024 edition of AIME, maintaining the same format of 15 challenging mathematics problems with integer answers from 000 to 999.
15 problemsInteger answers 000-999High school olympiad level
RefreshingDisplay only
AIME 2024 2024 · updated September 2, 2026
AIME 2025
American Invitational Mathematics Examination 2025
The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.
15 problemsInteger answers 000-999High school olympiad level
CurrentDisplay only
AIME 2025 · updated September 2, 2026
GSM8K
Grade School Math 8K
A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
Grade-school math word problemsExact matchGrade-school math
CurrentDisplay only
GSM8K 2026 · updated September 2, 2026
MATH
MATH
A competition-style mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
Competition math problemsExact matchAdvanced math reasoning
CurrentDisplay only
MATH 2026 · updated September 2, 2026
CMath
CMath
A Chinese mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.
Chinese math problemsExact matchMath reasoning
CurrentDisplay only
CMath 2026 · updated September 2, 2026
AIME25 (Arcee)
AIME25 first-party comparison snapshot
A display-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart.
15 problemsInteger answers 000-999High school olympiad level
CurrentDisplay only
AIME25 (Arcee) 2026 · updated September 2, 2026
HMMT Feb 2023
Harvard-MIT Mathematics Tournament February 2023
A prestigious high school mathematics competition hosted jointly by Harvard and MIT, featuring challenging problems across various mathematical disciplines.
Tournament problemsCompetition mathematicsHigh school olympiad level
StaleDisplay only
HMMT Feb 2023 2023 · updated September 2, 2026
HMMT Feb 2024
Harvard-MIT Mathematics Tournament February 2024
The 2024 February edition of the Harvard-MIT Mathematics Tournament, continuing the tradition of challenging high school mathematics competition.
Tournament problemsCompetition mathematicsHigh school olympiad level
RefreshingDisplay only
HMMT Feb 2024 2024 · updated September 2, 2026
HMMT Feb 2025
Harvard-MIT Mathematics Tournament February 2025
The most recent February edition of the Harvard-MIT Mathematics Tournament, featuring the latest challenging problems in competitive mathematics.
Tournament problemsCompetition mathematicsHigh school olympiad level
CurrentDisplay only
HMMT Feb 2025 2025 · updated September 2, 2026
BRUMO 2025
Bulgarian Mathematical Olympiad 2025
A challenging mathematical olympiad competition featuring problems that test advanced mathematical reasoning and problem-solving skills at the olympiad level.
Olympiad problemsMathematical olympiadMathematical olympiad level
CurrentDisplay only
BRUMO 2025 2025 · updated September 2, 2026
MATH-500
MATH-500 Problem Set
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
500 problemsFree-form mathematical answersHigh school to undergraduate
StaleDisplay only
MATH-500 2021 · updated September 2, 2026
AIME26
AIME 2026
A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.
Competition math problemsShort-answer mathematicsOlympiad-style mathematics
CurrentWeighted 25%
AIME26 2026 · updated September 2, 2026
IPhO 2025 (Theory)
International Physics Olympiad 2025 (Theory)
The three official theory problems from the 2025 International Physics Olympiad, scored with blinded human evaluation.
3 olympiad theory problemsPhysics olympiad theoryInternational olympiad physics
CurrentDisplay only
IPhO 2025 (Theory) 2026 · updated September 2, 2026
HMMT Feb 2025
Harvard-MIT Mathematics Tournament February 2025
A February 2025 HMMT slice used in exact-value provider tables for advanced contest-math reasoning.
Competition math problemsContest mathematicsOlympiad-style mathematics
CurrentDisplay only
HMMT Feb 2025 2025 · updated September 2, 2026
HMMT Nov 2025
Harvard-MIT Mathematics Tournament November 2025
A November 2025 HMMT slice for high-end mathematical reasoning comparisons.
Competition math problemsContest mathematicsOlympiad-style mathematics
CurrentDisplay only
HMMT Nov 2025 2025 · updated September 2, 2026
HMMT Feb 2026
Harvard-MIT Mathematics Tournament February 2026
A February 2026 HMMT slice used in newer frontier-model math comparisons.
Competition math problemsContest mathematicsOlympiad-style mathematics
CurrentWeighted 25%
HMMT Feb 2026 2026 · updated September 2, 2026
IMOAnswerBench
IMOAnswerBench
A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Advanced mathematical answer generationPass@1 math benchmarkOlympiad-level mathematics
CurrentDisplay only
IMOAnswerBench 2026 · updated September 2, 2026
Apex
Apex
A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning
CurrentDisplay only
Apex 2026 · updated September 2, 2026
Apex Shortlist
Apex Shortlist
A shortlist subset of the Apex mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.
Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning
CurrentDisplay only
Apex Shortlist 2026 · updated September 2, 2026
MMAnswerBench
MMAnswerBench
A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.
Multimodal math questionsVisual and structured mathematical QAAdvanced mathematical reasoning
CurrentDisplay only
MMAnswerBench 2026 · updated September 2, 2026
FrontierMath (legacy)
FrontierMath legacy aggregate
Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.
Historical aggregateOpen-ended mathematical reasoning with tool accessResearch-level mathematics
RefreshingDisplay only
FrontierMath (legacy) 2024 · updated September 2, 2026
FrontierMath v2 (Tiers 1-3)
FrontierMath v2 Tiers 1-3
Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.
295 private advanced mathematics problemsPython-enabled iterative mathematical problem solvingFrom olympiad-plus to early research mathematics
CurrentWeighted 30%
FrontierMath v2 (Tiers 1-3) 2026 · updated September 2, 2026
FrontierMath v2 (Tier 4)
FrontierMath v2 Tier 4
Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.
43 private extreme-difficulty mathematics problemsPython-enabled iterative mathematical problem solvingResearch-level mathematics requiring hours or days of expert work
CurrentWeighted 10%
FrontierMath v2 (Tier 4) 2026 · updated September 2, 2026
USAMO 2026
United States of America Mathematical Olympiad 2026
The premier US mathematical olympiad competition, featuring proof-based problems that require deep mathematical insight and rigorous argumentation at the highest competition level.
6 proof-based problemsMathematical proof constructionInternational olympiad level
CurrentWeighted 10%
USAMO 2026 2026 · updated September 2, 2026
Korean(8 benchmarks)
KMMLU
Korean Massive Multitask Language Understanding
Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.
35,030 questionsMultiple choice questionsElementary to professional level in Korean
RefreshingDisplay only
KMMLU 2024 · updated September 2, 2026
KMMLU-Hard
KMMLU-Hard
A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.
~5,000 questionsMultiple choice questionsAdvanced Korean reasoning
CurrentDisplay only
KMMLU-Hard 2025 · updated September 2, 2026
KMMLU-Redux
KMMLU-Redux
Cleaned KMMLU from national technical qualification exams, with errors removed, decontaminated, and deduplicated.
~3,500 questionsTechnical multiple choiceIndustrial/technical
RefreshingDisplay only
KMMLU-Redux · updated September 2, 2026
KMMLU-Pro
KMMLU-Pro
Korean National Professional Licensure exams evaluating professional-grade knowledge.
~2,500 questionsProfessional licensure examsProfessional
RefreshingDisplay only
KMMLU-Pro · updated September 2, 2026
CLIcK
Cultural and Linguistic Intelligence in Korean
Evaluates Korean culture and linguistics.
1,995 questionsCultural/linguistic QAKorean cultural nuances
RefreshingDisplay only
CLIcK · updated September 2, 2026
KoBALT
Korean Benchmark for Advanced Linguistic Tasks
Evaluates advanced Korean linguistic competence.
Linguistics questionsAdvanced linguisticsAdvanced linguistic phenomena
RefreshingDisplay only
KoBALT · updated September 2, 2026
Korean CSAT
College Scholastic Ability Test (수능)
The Korean SAT exam.
Multi-subject examStandardized testHigh school to college level
RefreshingDisplay only
Korean CSAT · updated September 2, 2026
HRM8K
HAE-RAE Math 8K
Korean mathematical reasoning (high-school to Olympiad level).
8,011 instancesMath word problemsOlympiad level
RefreshingDisplay only
HRM8K · updated September 2, 2026
External benchmark mirrors(70 benchmarks)
Mirrored from third-party leaderboards; listed by source, not by capability.
KindBench
KindBench Psychological Safety Benchmark
A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.
16 multi-turn scenarios, 72 criteriaJudge-scored behavioral audit with human reviewAdversarial psychological-safety evaluation
CurrentDisplay only
KindBench v0.1.0 · updated September 2, 2026
LiveBench
LiveBench
A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.
23 objective tasks across 7 categoriesMean of category averagesBroad frontier-model evaluation
RefreshingDisplay only
LiveBench 2024 · updated September 2, 2026
Vals Index
Vals Index v2
Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.
Finance, coding, spreadsheet modeling, code migration, and legal-work componentsComposite scorePrivate economic-work benchmark composite
CurrentDisplay only
Vals Index 2026 · updated September 2, 2026
Web Search Index
Vals Web Search Index
A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.
Finance Agent Benchmark v2 and Legal Research Benchmark tasksAccuracy by model and search-tool combinationProfessional web research with controlled search-tool variants
CurrentDisplay only
Web Search Index 2026 · updated September 2, 2026
Time Horizon Index: KSP
Vals Time Horizon Index: Kerbal Space Program
A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.
30 progressively harder Kerbal Space Program missionsMission-ladder progress with partial creditLong-horizon autonomous computer use and planning
CurrentDisplay only
Time Horizon Index: KSP 2026 · updated September 2, 2026
Vals Multimodal Index
Vals Multimodal Index v1.2
Vals AI multimodal composite across finance, coding, education, and mortgage-tax task families.
Finance, coding, education, and mortgage-tax componentsComposite scorePrivate multimodal economic-work benchmark composite
CurrentDisplay only
Vals Multimodal Index 2026 · updated September 2, 2026
Finance Agent v1.1
Vals Finance Agent v1.1
An archived Vals benchmark covering retrieval, numerical reasoning, financial modeling, market analysis, earnings, trends, and adjustments.
11 financial analyst task viewsAccuracy with task-level breakdownsProfessional financial analysis
CurrentDisplay only
Finance Agent v1.1 2026 · updated September 2, 2026
ReverseEngBench
Vals ReverseEngBench
A contamination-resistant agent benchmark for reverse engineering real-world binaries.
Real-world binary reverse-engineering tasksFully solved rate and capability scoreAgentic reverse engineering
CurrentDisplay only
ReverseEngBench 2026 · updated September 2, 2026
Vals Terminal-Bench 1.0 mirror
Vals-hosted Terminal-Bench 1.0 mirror
A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.
Terminal tasks split by easy, medium, and hard difficultyAccuracy scoreTerminal-agent execution
CurrentDisplay only
Vals Terminal-Bench 1.0 mirror 2026 · updated September 2, 2026
CorpFin v2
Vals CorpFin v2
Vals AI private benchmark for understanding long-context credit agreements.
Credit-agreement understanding tasksAccuracy scoreProfessional finance document reasoning
CurrentDisplay only
CorpFin v2 2026 · updated September 2, 2026
MedCode
Vals MedCode
Vals AI healthcare benchmark for whether models can support the medical billing process.
Medical billing support tasksAccuracy scoreProfessional healthcare administration
CurrentDisplay only
MedCode 2026 · updated September 2, 2026
MedScribe
Vals MedScribe
Vals AI healthcare benchmark for whether models can support doctors with administrative work.
Medical administrative support tasksAccuracy scoreProfessional healthcare administration
CurrentDisplay only
MedScribe 2026 · updated September 2, 2026
MortgageTax
Vals MortgageTax
Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.
Mortgage and tax extraction tasksAccuracy scoreProfessional mortgage-tax document reasoning
CurrentDisplay only
MortgageTax 2026 · updated September 2, 2026
ProofBench
Vals ProofBench
Vals AI automated theorem-proving benchmark.
Automated theorem provingAccuracy scoreFormal proof reasoning
CurrentDisplay only
ProofBench 2026 · updated September 2, 2026
LegalBench
Vals LegalBench
Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.
Legal reasoning task viewsAccuracy scoreProfessional legal reasoning
CurrentDisplay only
LegalBench 2026 · updated September 2, 2026
CaseLaw v2
Vals CaseLaw v2
Vals AI private question-answer benchmark over Canadian court cases.
Canadian case-law question answeringAccuracy scoreProfessional legal retrieval and reasoning
CurrentDisplay only
CaseLaw v2 2026 · updated September 2, 2026
DeepSWE
DeepSWE
A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.
113 software engineering tasks across 91 repositories and 5 languagesPass@1 with confidence interval, cost, time, and token metadataLong-horizon software engineering
CurrentDisplay only
DeepSWE 2026 · updated September 2, 2026
SWE-Marathon
SWE-Marathon
A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.
20 multi-hour software engineering tasksTask resolution and trajectory reviewUltra-long-horizon software engineering
CurrentDisplay only
SWE-Marathon 2026 · updated September 2, 2026
ExploitBench
ExploitBench v8-bench
A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.
V8 exploit synthesis runsCapability coverage percentage over 16 flagsBrowser exploitation and cybersecurity
CurrentDisplay only
ExploitBench 2026 · updated September 2, 2026
ACE solved
ACE Cyber Range Challenges Solved
Number of advanced cyber-range challenges solved in the joint NIST CAISI and UK AISI preliminary evaluation.
41 advanced cyber-range challengesChallenges solvedAdvanced cyber operations
CurrentDisplay only
ACE solved 2026 · updated September 2, 2026
The Last Ones steps
The Last Ones Average Progress
Average step reached on a 32-step long-horizon cyber range.
32-step long-horizon cyber rangeAverage step reachedLong-horizon cyber operations
CurrentDisplay only
The Last Ones steps 2026 · updated September 2, 2026
The Last Ones completion
The Last Ones Cyber Range Completion Rate
Share of runs that completed the 32-step cyber range within the 100-million-token limit.
10 long-horizon runsCompletion rateLong-horizon cyber operations
CurrentDisplay only
The Last Ones completion 2026 · updated September 2, 2026
ACCR standard
Advanced Cyber Completion Rate — Standard Access
Share of approved advanced-cyber requests completed rather than refused under standard GPT-5.6 Sol safeguards.
Internal advanced-cyber request setCompletion rateCyber access and safeguards
CurrentDisplay only
ACCR standard 2026 · updated September 2, 2026
ACCR Daybreak Blue
Advanced Cyber Completion Rate — Daybreak Blue
Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Sol under Daybreak Blue safeguards.
Internal advanced-cyber request setCompletion rateCyber access and safeguards
CurrentDisplay only
ACCR Daybreak Blue 2026 · updated September 2, 2026
ACCR Daybreak Red
Advanced Cyber Completion Rate — Daybreak Red
Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Cyber under Daybreak Red access.
Internal advanced-cyber request setCompletion rateCyber access and safeguards
CurrentDisplay only
ACCR Daybreak Red 2026 · updated September 2, 2026
SEC-Bench Pro
SEC-Bench Pro
Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.
Security engineering tasksSuccess rateAdvanced cybersecurity
CurrentDisplay only
SEC-Bench Pro 2026 · updated September 2, 2026
FrontierCyber
FrontierCyber
Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.
197 dynamic cyber tasksTasks solvedEasy through elite cyber operations
CurrentDisplay only
FrontierCyber 2026 · updated September 2, 2026
CyScenarioBench success
CyScenarioBench Average Success Rate
Average success rate across realistic, long-horizon cybersecurity scenarios.
11 cyber scenariosAverage success rateLong-horizon cybersecurity
CurrentDisplay only
CyScenarioBench success 2026 · updated September 2, 2026
CyScenarioBench solved
CyScenarioBench Scenarios Ever Solved
Number of CyScenarioBench scenarios completed successfully in at least one run.
11 cyber scenariosScenarios solved at least onceLong-horizon cybersecurity
CurrentDisplay only
CyScenarioBench solved 2026 · updated September 2, 2026
Atomic network attacks
Atomic Network Attack Simulation
Irregular's domain-level evaluation of network attack simulation capability.
Atomic cyber tasksDomain averageNetwork attack simulation
CurrentDisplay only
Atomic network attacks 2026 · updated September 2, 2026
Atomic vulnerability research
Atomic Vulnerability Research and Exploitation
Irregular's domain-level evaluation of vulnerability research and exploitation capability.
Atomic cyber tasksDomain averageVulnerability research and exploitation
CurrentDisplay only
Atomic vulnerability research 2026 · updated September 2, 2026
Atomic evasion
Atomic Evasion
Irregular's domain-level evaluation of cybersecurity evasion capability.
Atomic cyber tasksDomain averageCybersecurity evasion
CurrentDisplay only
Atomic evasion 2026 · updated September 2, 2026
SCONE post-cutoff success
SCONE Post-Cutoff Exploit Success
Share of SCONE smart-contract vulnerabilities exploited on the 12-task post-cutoff set.
12 post-cutoff smart-contract vulnerabilitiesBest@8 exploit success rateSmart-contract exploitation
CurrentDisplay only
SCONE post-cutoff success 2026 · updated September 2, 2026
SCONE simulated revenue
SCONE Post-Cutoff Simulated Exploit Revenue
Simulated value captured across the SCONE post-cutoff smart-contract set.
12 post-cutoff smart-contract vulnerabilitiesSimulated USD millions, Best@8Smart-contract exploitation
CurrentDisplay only
SCONE simulated revenue 2026 · updated September 2, 2026
Firefox 147 exploits
Firefox 147 Working Exploit Rate
Share of patched Firefox 147 JavaScript-engine targets for which the model produced a working arbitrary-code-execution exploit.
250 Firefox 147 vulnerability trialsWorking arbitrary-code-execution rateBrowser exploit development
CurrentDisplay only
Firefox 147 exploits 2026 · updated September 2, 2026
Anthropic OSS-Fuzz crash
Anthropic OSS-Fuzz Any-Crash Rate
Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.
Approximately 830 OSS-Fuzz entry pointsAny-crash rateVulnerability discovery
CurrentDisplay only
Anthropic OSS-Fuzz crash 2026 · updated September 2, 2026
Anthropic OSS-Fuzz write primitive
Anthropic OSS-Fuzz Write-Primitive-or-Higher Rate
Share of evaluated OSS-Fuzz entry points where the model achieved a write primitive or stronger result.
Approximately 830 OSS-Fuzz entry pointsWrite-primitive-or-higher rateVulnerability exploitation
CurrentDisplay only
Anthropic OSS-Fuzz write primitive 2026 · updated September 2, 2026
CVE-Bench zero-day
CVE-Bench v1 Zero-Day Black-Box Evaluation
OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.
40 critical CVEsPass@1 over three rolloutsBlack-box vulnerability exploitation
CurrentDisplay only
CVE-Bench zero-day 2026 · updated September 2, 2026
GBA-Eval
GBA-Eval
An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.
27 emulator test casesOverall emulator scoreLong-horizon systems programming
CurrentDisplay only
GBA-Eval 2026 · updated September 2, 2026
CAIS Text Leaderboard
CAIS AI Dashboard Text Capabilities Index
A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.
HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuestsAverage component scoreComposite frontier text capability
CurrentDisplay only
CAIS Text Leaderboard 2025 · updated September 2, 2026
VoxelBench Text
VoxelBench Text-Prompt Leaderboard
A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.
Live text prompts for 3D voxel constructionGlicko-2 rating from blind pairwise votes3D spatial construction and visual quality
CurrentDisplay only
Live VoxelBench Glicko-2 · updated September 2, 2026
VoxelBench Image
VoxelBench Image-Prompt Leaderboard
A live human-preference benchmark where multimodal models build voxel structures from image references and voters compare anonymous results produced from the same prompt.
Live image-reference prompts for 3D voxel constructionGlicko-2 rating from blind pairwise votesVisual grounding, 3D construction, and aesthetic quality
CurrentDisplay only
Live VoxelBench Glicko-2 · updated September 2, 2026
WeirdML
WeirdML v2
A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.
17 novel ML engineering tasksAverage accuracy across tasksNovel dataset modeling and iterative debugging
CurrentDisplay only
WeirdML 2026 · updated September 2, 2026
ALE-Bench
Agents Last Exam
A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.
152 ALE-V1 professional workflow tasks across 13 top-level domainsPass rate, partial-credit score, cost, token, and duration metadataReal-world agentic workflows
CurrentDisplay only
ALE-Bench 2026 · updated September 2, 2026
RuneScape-Bench
RuneBench / runescape-bench
An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.
16 RuneScape skill-training tasksAverage log XP-rate scoreAgentic gameplay automation
CurrentDisplay only
RuneScape-Bench 2026 · updated September 2, 2026
Toloka Arena
Toloka Arena
An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.
Private simulated enterprise workflowspass^5 arena scoreAgentic workflow reliability
CurrentDisplay only
Toloka Arena 2026 · updated September 2, 2026
Vals SWE-bench mirror
Vals-hosted SWE-bench mirror
Vals AI hosted SWE-bench view for solving production software engineering tasks.
Software engineering issue-resolution tasksAccuracy scoreProduction software engineering
CurrentDisplay only
Vals SWE-bench mirror 2026 · updated September 2, 2026
Vals Terminal-Bench 2.0 mirror
Vals-hosted Terminal-Bench 2.0 mirror
Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.
Terminal task difficulty splitsAccuracy scoreTerminal-based agent execution
CurrentDisplay only
Vals Terminal-Bench 2.0 mirror 2026 · updated September 2, 2026
Vals LiveCodeBench mirror
Vals-hosted LiveCodeBench mirror
Vals AI implementation of LiveCodeBench with easy, medium, and hard task splits.
Coding problem difficulty splitsAccuracy scoreContamination-resistant coding problems
CurrentDisplay only
Vals LiveCodeBench mirror 2026 · updated September 2, 2026
Vals GPQA Diamond mirror
Vals-hosted GPQA Diamond mirror
Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.
GPQA Diamond task splitsAccuracy scoreGraduate science reasoning
CurrentDisplay only
Vals GPQA Diamond mirror 2026 · updated September 2, 2026
Vals MMLU-Pro mirror
Vals-hosted MMLU-Pro mirror
Vals AI hosted MMLU-Pro view with subject-level task splits.
MMLU-Pro subject splitsAccuracy scoreProfessional academic reasoning
CurrentDisplay only
Vals MMLU-Pro mirror 2026 · updated September 2, 2026
EMB
Vals EMB
Evaluating agents on Excel-based financial modeling tasks
Excel-based financial modeling tasksAccuracy scoreProfessional finance modeling
CurrentDisplay only
EMB 2026 · updated September 2, 2026
CyberBench
Vals CyberBench
Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and stop crashing after the fix?
OSS-Fuzz PoC and patch-verification tasksAccuracy scoreAutonomous cybersecurity exploit reproduction
CurrentDisplay only
CyberBench 2026 · updated September 2, 2026
TaxEval v2
Vals TaxEval v2
A Vals-created set of questions and responses to tax questions
Tax question answering and response evaluationAccuracy scoreProfessional tax reasoning
CurrentDisplay only
TaxEval v2 2026 · updated September 2, 2026
Harvey's Legal Agent Benchmark
Vals Harvey's Legal Agent Benchmark
Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools
Legal agent work across documents, spreadsheets, presentations, and filesAccuracy scoreProfessional legal workflow automation
CurrentDisplay only
Harvey's Legal Agent Benchmark 2026 · updated September 2, 2026
Terminal-Bench 2.1
Vals Terminal-Bench 2.1
State-of-the-art set of difficult terminal-based tasks
Terminal-based task executionAccuracy scoreFrontier terminal-agent execution
CurrentDisplay only
Terminal-Bench 2.1 2026 · updated September 2, 2026
Code Migration
Vals Code Migration
Can language models reimplement real-world programs in another language?
Real-world program reimplementation in another languageAccuracy scoreProduction code migration
CurrentDisplay only
Code Migration 2026 · updated September 2, 2026
Legal Research Bench
Vals Legal Research Bench
Evaluating agents on legal research tasks across diverse areas of US law
US-law legal research tasksAccuracy scoreProfessional legal research
CurrentDisplay only
Legal Research Bench 2026 · updated September 2, 2026
MedQA
Vals MedQA
Evaluating language model bias in medical questions.
Medical question answeringAccuracy scoreMedical knowledge and bias evaluation
CurrentDisplay only
MedQA 2026 · updated September 2, 2026
AIME
Vals AIME
Challenging national math exam given to top high-school students
AIME math problemsAccuracy scoreCompetition math
CurrentDisplay only
AIME 2026 · updated September 2, 2026
MATH 500
Vals MATH 500
Academic math benchmark on probability, algebra, and trigonometry
MATH 500 academic math problemsAccuracy scoreAdvanced academic math
CurrentDisplay only
MATH 500 2026 · updated September 2, 2026
MGSM
Vals MGSM
A multilingual benchmark for mathematical questions.
Multilingual grade-school math questionsAccuracy scoreMultilingual mathematical reasoning
CurrentDisplay only
MGSM 2026 · updated September 2, 2026
MMMU
Vals MMMU
Multimodal Multi-task Benchmark
Multimodal academic task suiteAccuracy scoreMultimodal college-level reasoning
CurrentDisplay only
MMMU 2026 · updated September 2, 2026
SAGE
Vals SAGE
Student Assessment with Generative Evaluation
Student assessment with generative evaluationAccuracy scoreEducation assessment reasoning
CurrentDisplay only
SAGE 2026 · updated September 2, 2026
IOI
Vals IOI
Based on the International Olympiad in Informatics
International Olympiad in Informatics-style programming tasksAccuracy scoreOlympiad programming
CurrentDisplay only
IOI 2026 · updated September 2, 2026
ProgramBench
Vals ProgramBench
Can language models rebuild programs from scratch?
Program reconstruction tasksAccuracy scoreCleanroom software engineering
CurrentDisplay only
ProgramBench 2026 · updated September 2, 2026
SkillsBench
Vals SkillsBench
How important are skills for agents?
Agent skill-importance tasksAccuracy scoreAgent skill evaluation
CurrentDisplay only
SkillsBench 2026 · updated September 2, 2026
Agent Poker Bench
Vals Agent Poker Bench
Which model can make the most money playing poker?
Poker-playing agent trialsAccuracy scoreStrategic game-agent decision making
CurrentDisplay only
Agent Poker Bench 2026 · updated September 2, 2026
Public Benefits Bench v1.1
Vals Public Benefits Bench v1.1
Can AI help people navigate SNAP benefits?
SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation
CurrentDisplay only
Public Benefits Bench v1.1 2026 · updated September 2, 2026
Public Benefits Bench v1
Vals Public Benefits Bench v1
Can AI help people navigate SNAP benefits?
SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation
CurrentDisplay only
Public Benefits Bench v1 2026 · updated September 2, 2026