Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

AI Benchmarks Directory

10 categories · 416 evaluations

BenchLM tracks 416 AI benchmarks across coding, agentic, reasoning, math, knowledge, multimodal, instruction following, and multilingual tasks. Each record links the evaluation to its ranked leaderboard and September 2026 data.

2026

DRACO

Data Research and Analysis with Complex Operations

Agentic data-analysis tasks scored against per-task rubrics at a 980K-token budget.

Agentic data research and analysis tasksNormalized rubric scoreProfessional data analysis

CurrentDisplay only

DRACO 2026 · updated September 2, 2026

2026

BrowseComp (10-agent, prerelease)

Multi-Agent BrowseComp — 10-agent team prerelease configuration

BrowseComp accuracy from ten collaborating Opus 5 agents on a pre-release model and unreleased effort configuration.

BrowseComp web-research tasks10-agent team accuracyLong-horizon web research

CurrentDisplay only

BrowseComp (10-agent, prerelease) 2026 · updated September 2, 2026

2026

MCP-Atlas claim coverage

MCP-Atlas mean claim coverage

Average coverage of required claims in answers produced during real-world MCP tool-use workflows.

Production-like multi-server MCP workflowsMean claim coverageReal-world tool use

CurrentDisplay only

MCP-Atlas claim coverage 2026 · updated September 2, 2026

2026

LAB all-pass (Anthropic harness)

Legal Agent Benchmark all-pass rate — Anthropic harness

Strict task success requiring every expert-written legal-work rubric criterion to pass.

1,235 legal-agent tasksAll-criteria pass rateProfessional legal work

CurrentDisplay only

LAB all-pass (Anthropic harness) 2026 · updated September 2, 2026

2026

LAB criterion-pass (Anthropic harness)

Legal Agent Benchmark mean criterion-pass rate — Anthropic harness

Mean fraction of expert-written rubric criteria passed across legal-agent tasks.

1,235 legal-agent tasksMean criterion-pass rateProfessional legal work

CurrentDisplay only

LAB criterion-pass (Anthropic harness) 2026 · updated September 2, 2026

2026

LAB all-pass (Harvey held-out)

Legal Agent Benchmark all-pass rate — Harvey held-out set

Harvey AI's strict held-out task success rate requiring every rubric criterion to pass.

Harvey-held-out legal-agent tasksAll-criteria pass rateProfessional legal work

CurrentDisplay only

LAB all-pass (Harvey held-out) 2026 · updated September 2, 2026

2026

LAB criterion-pass (Harvey held-out)

Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set

Harvey AI's mean criterion-level score on its held-out legal-agent evaluation.

Harvey-held-out legal-agent tasksMean criterion-pass rateProfessional legal work

CurrentDisplay only

LAB criterion-pass (Harvey held-out) 2026 · updated September 2, 2026

2026

Toolathlon Verified Pass@3

Toolathlon Verified Pass@3

Fraction of Toolathlon Verified tasks solved in at least one of three trials.

108 verified real-world tool-use tasksPass@3Long-horizon application tool use

CurrentDisplay only

Toolathlon Verified Pass@3 2026 · updated September 2, 2026

2026

Toolathlon Verified Pass³

Toolathlon Verified Pass cubed

Fraction of Toolathlon Verified tasks solved in all three independent trials.

108 verified real-world tool-use tasksAll-three-trials pass rateLong-horizon application tool use

CurrentDisplay only

Toolathlon Verified Pass³ 2026 · updated September 2, 2026

2026

Toolathlon Verified avg. turns

Toolathlon Verified average assistant turns

Average assistant turns per Toolathlon Verified trajectory.

108 verified real-world tool-use tasksAverage trajectory lengthLong-horizon application tool use

CurrentDisplay only

Toolathlon Verified avg. turns 2026 · updated September 2, 2026

2026

Terminal-Bench 3.0

Terminal-Bench 3.0

A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance, engineering, math, and science tasks.

74 professional computer-work tasks across 7 domainsTask completion rateFrontier autonomous knowledge work

CurrentDisplay only

Terminal-Bench 3.0 · updated September 2, 2026

2026

Terminal-Bench 4.0

Terminal-Bench 4.0

The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing unstable tasks, and removing tasks that no longer separate frontier systems.

66 professional computer-work tasksTask completion rate across 5 trials per taskFrontier autonomous knowledge work

CurrentDisplay only

Terminal-Bench 4.0.0 · updated September 2, 2026

2026

Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1

A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth, mathematical, and engineering sciences.

70 expert-curated scientific research workflowsResolution rate across 3 trials per taskFrontier scientific research workflows

CurrentDisplay only

Terminal-Bench-Science 0.1.0 · updated September 2, 2026

2026

EQ-Bench 4

EQ-Bench 4

A multi-turn benchmark of emotional and social intelligence using synthetic personas and pairwise LLM judging.

120 multi-turn persona scenariosPairwise EloEmotional and social intelligence

CurrentDisplay only

EQ-Bench 4 2026 · updated September 2, 2026

2026

AutoCAD-Bench

AutoCAD-Bench

A Markov Studios computer-use benchmark that asks agents to produce 2D drawings and 3D models in AutoCAD.

21 2D drawing tasks and 29 3D modeling tasksTask completion rate at a 75-point rubric thresholdProfessional CAD computer use

CurrentDisplay only

AutoCAD-Bench 2026 · updated September 2, 2026

2026

Design Arena Agentic Web Dev

Design Arena Agentic Web Dev Elo

A display-only Elo rating from blinded comparisons of multi-file web applications built by coding agents.

Multi-file web application developmentElo from blinded human preferencesAgentic frontend development

CurrentDisplay only

Design Arena Agentic Web Dev 2026 · updated September 2, 2026

2026

AA Briefcase

Artificial Analysis Briefcase

An independently evaluated professional-work benchmark reported as Elo.

Professional knowledge-work tasksEloProfessional work

CurrentDisplay only

AA Briefcase 2026 · updated September 2, 2026

2026

AA AutomationBench

Artificial Analysis AutomationBench

An independently evaluated automation benchmark from Artificial Analysis.

Business-process automation tasksTask success rateAgentic automation

CurrentDisplay only

AA AutomationBench 2026 · updated September 2, 2026

2026

AA EnterpriseOps-Gym

Artificial Analysis EnterpriseOps-Gym

An independently evaluated enterprise-operations benchmark from Artificial Analysis.

Enterprise operations workflowsTask success rateEnterprise agent operations

CurrentDisplay only

AA EnterpriseOps-Gym 2026 · updated September 2, 2026

2026

AA Harvey LAB

Artificial Analysis Harvey LAB-AA

An independently evaluated legal-agent benchmark from Artificial Analysis.

Legal agent tasksTask success rateProfessional legal work

CurrentDisplay only

AA Harvey LAB 2026 · updated September 2, 2026

2026

AA ITBench

Artificial Analysis ITBench-AA

An independently evaluated IT-operations benchmark from Artificial Analysis.

IT incident-response tasksTask success rateEnterprise IT operations

CurrentDisplay only

AA ITBench 2026 · updated September 2, 2026

2026

AA Tau3 Banking

Artificial Analysis Tau3-Banking

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

Banking tool-use workflowsTask success rateAgentic banking workflows

CurrentDisplay only

AA Tau3 Banking 2026 · updated September 2, 2026

2026

Terminal-Bench 2.0

Terminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. Models must inspect files, run commands, edit code, and recover from errors over multi-step workflows.

Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering

CurrentWeighted 38%

Terminal-Bench 2 · updated September 2, 2026

2026

Terminal-Bench 2.1

Terminal-Bench 2.1 (provider run)

A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.

Terminal-based software-agent tasksInteractive task success rateProfessional software engineering

CurrentDisplay only

Terminal-Bench 2.1 2026 · updated September 2, 2026

2025

BrowseComp

BrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

Research questions requiring browsingWeb search and evidence synthesisHard web research

CurrentWeighted 28%

BrowseComp 2026 · updated September 2, 2026

2026

HLE w/ tools

Humanity's Last Exam with tools

Tool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.

Expert questions with tool usePass@1Frontier tool-augmented reasoning

CurrentDisplay only

HLE w/ tools 2026 · updated September 2, 2026

2026

GDPval-AA

GDPval-AA

An agentic real-world work-task evaluation reported as an Elo score in DeepSeek-V4 thinking-mode evaluations.

Agentic real-world work tasksEloProfessional agentic workflows

CurrentDisplay only

GDPval-AA 2026 · updated September 2, 2026

2026

GDPval-AA

GDPval-AA normalized

A display-only Artificial Analysis normalized score for economically valuable tasks.

Economically valuable tasksNormalized scoreProfessional agentic workflows

CurrentDisplay only

GDPval-AA 2026 · updated September 2, 2026

2026

AA Agentic Index

Artificial Analysis Agentic Index

A display-only Artificial Analysis agentic index.

Cross-benchmark agentic indexAggregated model scoreDisplay-only external reference

CurrentDisplay only

AA Agentic Index 2026 · updated September 2, 2026

2026

APEX-Agents-AA

APEX-Agents-AA

Artificial Analysis' implementation of the APEX-Agents benchmark for long-horizon professional-services agent tasks.

452 professional-services agent tasksPass@1Long-horizon workplace agent tasks

CurrentDisplay only

APEX-Agents-AA 2026 · updated September 2, 2026

2026

Gert Labs

Gert Labs Composite Game Benchmark

A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.

Novel game environmentsComposite game leaderboardAgentic coding and decision-making

CurrentDisplay only

Gert Labs 2026 · updated September 2, 2026

2025

OSWorld-Verified

OSWorld-Verified

OSWorld-Verified is the July 2025 repaired release of OSWorld's real-computer evaluation. It measures whether a model-agent system can finish desktop and web tasks from configured starting states, with success checked by execution-based evaluators.

369 real-world computer tasks (361 when eight Google Drive tasks are excluded)Execution-based interactive task successMulti-step desktop and cross-application workflows

CurrentWeighted 34%

OSWorld Verified · updated September 2, 2026

2026

OSWorld 2.0

OSWorld 2.0

A long-horizon computer-use benchmark covering realistic workflows across everyday and professional desktop tasks.

108 long-horizon computer-use workflowsInteractive computer-use evaluationLong-horizon professional workflows

CurrentDisplay only

OSWorld 2.0 2026 · updated September 2, 2026

2026

CyberGym

CyberGym

A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.

1,507 vulnerability analysis instancesVulnerability reproduction and PoC generationReal-world cybersecurity

CurrentDisplay only

CyberGym 2026 · updated September 2, 2026

2026

CWE-Bench

CWE-Bench

An external benchmark for evaluating whether coding agents can produce correct patches for real-world software vulnerabilities.

Real-world vulnerability patching tasksPass@1Automated vulnerability remediation

CurrentDisplay only

CWE-Bench 2026 · updated September 2, 2026

2026

CTI-REALM

CTI-REALM

A cybersecurity benchmark that measures whether an agent can turn raw threat-intelligence reports into working detection rules.

Threat-intelligence-to-detection-rule workflowsSuccess rateProfessional cyber threat detection

CurrentDisplay only

CTI-REALM 2026 · updated September 2, 2026

2025

Cybench

Cybench

A cybersecurity benchmark of professional Capture the Flag tasks for measuring autonomous cyber agent capability and risk.

40 professional CTF tasksCybersecurity agent task completionProfessional cybersecurity

CurrentDisplay only

Cybench 2025 · updated September 2, 2026

2026

ExploitGym

ExploitGym

A controlled benchmark for evaluating whether AI agents can extend vulnerability-triggering inputs into working exploits.

898 exploitation tasksWorking exploit generationAdvanced cybersecurity exploitation

CurrentDisplay only

ExploitGym 2026 · updated September 2, 2026

2026

JobBench

JobBench

An occupational agent benchmark for professional workflows that workers say they most want delegated to AI.

130 tasks across 35 occupationsAgentic workplace deliverablesProfessional multi-source workflows

CurrentDisplay only

JobBench 2026 · updated September 2, 2026

2026

BrowseComp-VL

BrowseComp-VL

A vision-language browsing benchmark for multimodal web research and tool-use workflows.

Multimodal browsing tasksVision-language web research evaluationMultimodal browser-agent

CurrentDisplay only

BrowseComp-VL 2026 · updated September 2, 2026

2026

OSWorld

OSWorld

A computer-use benchmark for GUI task completion across the broader OSWorld task suite.

Computer-use tasksInteractive GUI evaluationBroad computer-use suite

CurrentDisplay only

OSWorld 2026 · updated September 2, 2026

2026

AndroidWorld

AndroidWorld

A mobile GUI agent benchmark for completing Android app workflows and on-device tasks.

Android app workflowsInteractive mobile-agent evaluationComplex mobile task completion

CurrentDisplay only

AndroidWorld 2026 · updated September 2, 2026

2026

WebVoyager

WebVoyager

A browser-agent benchmark for completing multi-step workflows on live websites.

Live website workflowsInteractive browser-agent evaluationMulti-step web navigation

CurrentDisplay only

WebVoyager 2026 · updated September 2, 2026

2026

MCP Atlas

MCP Atlas

A benchmark for tool-calling over Model Context Protocol integrations and external tools.

Tool-integrated agent tasksInteractive tool-calling evaluationAdvanced tool use

CurrentDisplay only

MCP Atlas 2026 · updated September 2, 2026

2026

Kimi Claw 24/7

Kimi Claw 24/7 Bench

A Moonshot AI internal long-horizon agent benchmark for persistent professional coworking tasks.

17 professional scenarios, 610 evaluation pointsAverage pass rate across repeated OpenClaw runsLong-horizon agentic work

CurrentDisplay only

Kimi Claw 24/7 2026 · updated September 2, 2026

2026

MCP Mark Verified

MCPMark-Verified

A human-verified edition of MCPMark for MCP tool use across Notion, GitHub, Filesystem, Postgres, and Playwright server environments.

MCP tool-use tasks across five server environmentsInteractive MCP task completionAdvanced tool use

CurrentDisplay only

MCP Mark Verified 2026 · updated September 2, 2026

2026

Toolathlon

Toolathlon

A tool-use benchmark focused on selecting, sequencing, and completing tasks with external tools.

Multi-tool workflowsInteractive tool-calling evaluationAdvanced tool use

CurrentDisplay only

Toolathlon 2026 · updated September 2, 2026

2026

Toolathlon-Verified

Toolathlon-Verified

A verified tool-use benchmark variant for completing multi-step workflows with external tools.

Verified multi-tool workflowsInteractive tool-use scoreAdvanced tool use

CurrentDisplay only

Toolathlon-Verified 2026 · updated September 2, 2026

2026

AutomationBench

AutomationBench

An agent benchmark for completing automation workflows in reproducible task environments.

600 public automation tasksAgent task-completion scoreLong-horizon automation

CurrentDisplay only

AutomationBench 2026 · updated September 2, 2026

2026

Agents' Last Exam

Agents' Last Exam

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Agent tasksProvider-reported task scoreAdvanced agentic work

CurrentDisplay only

Agents' Last Exam 2026 · updated September 2, 2026

2026

APEX-Agents

APEX-Agents

A professional-services agent benchmark covering long-horizon knowledge-work tasks.

Professional-services agent tasksAgent task-completion scoreLong-horizon professional work

CurrentDisplay only

APEX-Agents 2026 · updated September 2, 2026

2026

SpreadsheetBench 2

SpreadsheetBench 2

A spreadsheet-focused benchmark for agentic analysis and editing workflows.

Spreadsheet analysis and editing tasksAgent task-completion scoreProfessional spreadsheet work

CurrentDisplay only

SpreadsheetBench 2 2026 · updated September 2, 2026

2026

DECK-Bench

DECK-Bench (Internal)

Moonshot AI's internal benchmark for presentation and deck-production workflows.

Internal presentation workflowsInternal evaluation scoreProfessional presentation creation

CurrentDisplay only

DECK-Bench 2026 · updated September 2, 2026

2026

ZClawBench

ZClawBench

A Z.AI benchmark for OpenClaw-style agent workflows spanning information search, office work, data analysis, development and operations, automation, and security.

OpenClaw agent workflowsEnd-to-end agent benchmarkBroad productivity and operations workflows

CurrentDisplay only

ZClawBench 2026 · updated September 2, 2026

2025

τ²-bench results

τ²-Bench Tool-Agent-User Evaluation

This route is a sourced ledger for published τ²-bench results. Most current rows come from Artificial Analysis's telecom implementation, while named provider rows can use telecom, airline, retail, or aggregate setups.

Airline, retail, and telecom customer-service task setsPublished domain success or pass^k resultsDual-control customer-service workflows

CurrentDisplay only

τ²-Bench 2026 · updated September 2, 2026

2026

DeepSearchQA

DeepSearchQA

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

Agentic browsing and list-answer questionsSearch / open / find browser-agent evaluationAgentic web research

CurrentDisplay only

DeepSearchQA 2026 · updated September 2, 2026

2025

τ²-bench Airline

τ²-Bench Airline Domain

τ²-bench Airline tests conversational agents on airline customer-service tasks governed by domain policy and database-changing tools.

Airline customer-service tasksDomain success under a published trial policyPolicy-constrained airline support workflows

CurrentDisplay only

τ²-bench Airline 2025 · updated September 2, 2026

2026

PinchBench

PinchBench

An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

23 OpenClaw agent tasksAverage success rate from official runsLong-horizon agent workflows

CurrentDisplay only

PinchBench 2026 · updated September 2, 2026

2025

OpenHands Index

OpenHands Index

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIAMacro-average across five coding-agent categoriesReal-world software engineering agent tasks

CurrentDisplay only

OpenHands Index 2025 · updated September 2, 2026

2026

SWE-Atlas Refactoring

SWE-Atlas Refactoring

A Scale SWE-Atlas software-engineering agent benchmark focused on refactoring tasks.

SWE-Atlas refactoring tasksRefactoring score with confidence intervalsReal-world software-engineering agent tasks

CurrentDisplay only

SWE-Atlas Refactoring 2026 · updated September 2, 2026

2026

SWE Refactor Bench

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Tests whether coding agents can complete long-horizon, whole-repository stack migrations while preserving the original program's behavior.

20 whole-repository stack migrationsMigration audit, frozen behavioral checks, and agentic verification6- to 30-hour autonomous repository migrations

CurrentDisplay only

SWE Refactor Bench 2026 · updated September 2, 2026

2026

AI4AI-Bench

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Tests whether coding agents can improve the training algorithm inside an existing AI research codebase, then survive a sealed training run and held-out evaluation.

10 AI training-algorithm design tasksFour-hour code rewrite followed by sealed training and held-out evaluationEnd-to-end AI research and algorithm design

CurrentDisplay only

AI4AI-Bench v1.5 · updated September 2, 2026

2026

InferenceBench

InferenceBench

A benchmark for open-ended LLM inference optimization by AI agents. Agents receive a base model, one H100, and a fixed time budget to build a valid OpenAI-compatible inference server that improves serving speed.

4 inference-serving optimization scenariosTwo-hour autonomous CLI agent runOpen-ended ML systems engineering

CurrentDisplay only

InferenceBench 2026 · updated September 2, 2026

2026

EdgeBench

EdgeBench

A ByteDance Seed benchmark of 134 real-world, day-scale tasks that measures how autonomous agents learn from environment feedback over 12+ hour interaction horizons, spanning scientific and ML, systems and software engineering, optimization, knowledge, formal, and game domains.

134 tasks (51 public) across 6 domainsLong-horizon interactive agent evaluationDay-scale expert tasks

CurrentDisplay only

EdgeBench 2026 · updated September 2, 2026

2026

BFCL v4

Berkeley Function Calling Leaderboard v4

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

Function-calling tasksTool invocation and schema evaluationAdvanced tool use

CurrentDisplay only

BFCL v4 2026 · updated September 2, 2026

2026

MLE-Bench Lite

MLE-Bench Lite

A lightweight machine-learning competition benchmark that measures whether models can iteratively train, evaluate, and improve ML systems in low-resource settings.

Low-resource ML competitionsAutonomous iterative ML optimizationAgentic machine learning

CurrentDisplay only

MLE-Bench Lite 2026 · updated September 2, 2026

2026

MM-ClawBench

MM-ClawBench

An OpenClaw-derived agent benchmark covering practical work and life tasks such as office document delivery, research, planning, and code maintenance.

OpenClaw-style real-world tasksAgent workflow evaluationBroad real-world agentic execution

CurrentDisplay only

MM-ClawBench 2026 · updated September 2, 2026

2026

Claw-Eval

Claw-Eval

A transparent real-world autonomous-agent benchmark with 300 human-verified tasks, 2,159 rubric items, and Pass^3 scoring across general, multi-turn, and native multimodal agent tasks.

300 tasks, 2,159 rubricsEnd-to-end autonomous-agent evaluation with Pass^3 scoringReal-world general, multi-turn, and native multimodal agent execution

CurrentDisplay only

Claw-Eval 2026 · updated September 2, 2026

2026

ResearchClawBench

ResearchClawBench

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

40 tasks across 10 scientific domainsEnd-to-end autonomous research evaluation with RADS scoringScientific research re-discovery

CurrentDisplay only

ResearchClawBench 2026 · updated September 2, 2026

2026

QwenClawBench

QwenClawBench

Qwen's internal OpenClaw-style benchmark for measuring broad real-world agent performance across practical productivity and research tasks.

Real-world agent workflowsEnd-to-end agent evaluationBroad real-world agentic execution

CurrentDisplay only

QwenClawBench 2026 · updated September 2, 2026

2026

QwenWebBench

QwenWebBench

A Qwen benchmark for artifact and webpage generation quality reported as an Elo-style rating.

Web artifacts and interactive deliverablesElo-style artifact benchmarkArtifact generation

CurrentDisplay only

QwenWebBench 2026 · updated September 2, 2026

2026

τ³-bench results

τ³-Bench Tool-Agent-User Evaluation

τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.

Corrected customer-service tasks plus knowledge and voice evaluation modesPublished domain or average success resultsLong-horizon, multimodal, and knowledge-aware tool use

CurrentDisplay only

τ³-bench results 2026 · updated September 2, 2026

2025

VITA-Bench

VITA-Bench

An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.

Interactive consumer-service agent tasksEnd-to-end interactive agent evaluationLong-horizon real-world workflows

CurrentDisplay only

VITA-Bench 2025 · updated September 2, 2026

2026

DeepPlanning

DeepPlanning

A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

Travel planning and constrained shoppingLong-horizon planning benchmarkConstrained agent planning

CurrentDisplay only

DeepPlanning 2026 · updated September 2, 2026

2026

MCP-Tasks

MCP-Tasks

A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.

MCP-integrated tool tasksInteractive tool-use evaluationAdvanced MCP workflows

CurrentDisplay only

MCP-Tasks 2026 · updated September 2, 2026

2026

WideResearch

WideResearch

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

Open-ended research tasksMulti-source research evaluationBroad research-agent workflows

CurrentDisplay only

WideResearch 2026 · updated September 2, 2026

2026

CoWorkBench

CoWorkBench

Qwen's internal benchmark for long-horizon professional work across computer science, finance, law, medicine, and other productivity domains.

Long-horizon professional workflowsProvider-run agent scoreCross-domain professional work

CurrentDisplay only

CoWorkBench 2026 · updated September 2, 2026

2026

MobileWorld

MobileWorld

A mobile-use agent benchmark for completing interactive tasks in smartphone environments.

Interactive mobile-device workflowsMobile agent task scoreLong-horizon mobile computer use

CurrentDisplay only

MobileWorld 2026 · updated September 2, 2026

2024

GAIA

General AI Assistants

GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

466

RefreshingDisplay only

GAIA 2024 · updated September 2, 2026

2024

TAU-bench

Tool-Agent-User Benchmark

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

Airline and retail task sets in the archived 2024 releaseDomain-specific pass^1 through pass^4 task successPolicy-constrained, multi-turn customer service

RefreshingDisplay only

TAU-bench 2024 · updated September 2, 2026

2024

WebArena

WebArena Web Agent Benchmark

WebArena tests whether a browser-agent system can complete 812 long-horizon tasks inside self-hosted replicas of functional websites. It checks the requested end state, so a result reflects the model, agent scaffold, browser interface, action budget, and evaluator together—not the base model alone.

812 long-horizon browser tasksEnd-state task successStateful multi-site browser work

RefreshingDisplay only

WebArena 2024 · updated September 2, 2026

2025

WebArena-Verified

WebArena-Verified Browser Agent Benchmark

WebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task descriptions, reference answers, and evaluators, and replaces nondeterministic judging with deterministic checks where possible.

812 verified tasks; separate 258-task Hard subsetDeterministic end-state task successAudited stateful browser work

CurrentDisplay only

WebArena-Verified 2025 · updated September 2, 2026

2026

MEWC

Multi-Environment Web Challenge

A benchmark that evaluates AI agents on multi-environment web challenges, testing navigation and task completion across diverse live web environments.

Web-agent tasksBrowser task completionOpen-web agent workflows

CurrentDisplay only

MEWC 2026 · updated September 2, 2026

2026

Finance Agent v2

Finance Agent v2

Vals AI benchmark for realistic financial analyst agent tasks across qualitative analysis, quantitative analysis, market work, comparables, precedents, earnings, disclosure, and modeling.

Financial analyst task categoriesMean score across repeated runsProfessional expert-task agent workflow

CurrentDisplay only

Finance Agent v2 2026 · updated September 2, 2026

2025

Market-Bench

Market-Bench

A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.

3 quantitative-trading strategiesBacktester implementation scored by mean absolute errorMarket simulation and quantitative coding

CurrentDisplay only

Market-Bench 2025 · updated September 2, 2026

2026

GDPval rubrics

GDPval rubrics

A display-only provider-table GDPval rubric score for economically valuable work tasks.

Economically valuable work tasksRubric scoreProfessional agentic workflows

CurrentDisplay only

GDPval rubrics 2026 · updated September 2, 2026

2026

BankerToolBench

BankerToolBench

A display-only provider benchmark for finance-oriented tool-use and agent workflows.

Finance and banking tool-use tasksTask success rateProfessional finance-agent workflows

CurrentDisplay only

BankerToolBench 2026 · updated September 2, 2026

2026

ProgramBench (episode 1)

ProgramBench hidden-test pass rate after episode 1

Program-reconstruction hidden-test pass rate after the first of five sequential long-context episodes.

166 golden program-reconstruction tasksHidden-test pass rate after episode 1Long-context clean-room software engineering

CurrentDisplay only

ProgramBench (episode 1) 2026 · updated September 2, 2026

2026

AA LiveCodeBench

Artificial Analysis LiveCodeBench

An independently evaluated LiveCodeBench result from Artificial Analysis.

Contamination-resistant coding tasksPass rateCompetitive programming

CurrentDisplay only

AA LiveCodeBench 2026 · updated September 2, 2026

2026

AA Terminal-Bench 2.1

Artificial Analysis Terminal-Bench v2.1

An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.

Terminal-based agent tasksTask success rateAgentic software engineering

CurrentDisplay only

AA Terminal-Bench 2.1 2026 · updated September 2, 2026

2026

Terminal-Bench 2.1

Terminal-Bench 2.1 (provider run)

A provider-run Terminal-Bench 2.1 result stored separately from the repository's Terminal-Bench 2.0 lane.

Terminal-based software-agent tasksInteractive task success rateProfessional software engineering

CurrentDisplay only

Terminal-Bench 2.1 2026 · updated September 2, 2026

2021

HumanEval

Evaluating Large Language Models Trained on Code

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

164 problemsPython function generationIntroductory to intermediate programming

StaleSaturatedDisplay only

HumanEval · updated September 2, 2026

2026

BigCodeBench

BigCodeBench

A code-generation benchmark reported in DeepSeek-V4 base-model evaluations.

Code generation tasksPass@1Software engineering

CurrentDisplay only

BigCodeBench 2026 · updated September 2, 2026

2026

Codeforces

Codeforces Rating

Competitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.

Competitive programming contestsRatingElite competitive programming

CurrentDisplay only

Codeforces 2026 · updated September 2, 2026

2026

Terminal-Bench 2.0

Terminal-Bench 2.0

A benchmark for agentic software engineering tasks executed in real terminal environments. DeepSeek reports it in the agentic section, while BenchLM also mirrors it in coding for models that publish it as a developer-task signal.

Terminal-based software tasksInteractive CLI agent evaluationProfessional software engineering

CurrentDisplay only

Terminal-Bench 2 · updated September 2, 2026

2024

SWE-bench Verified

Software Engineering Benchmark Verified

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

500 verified issuesCode patch generationProfessional software engineering

RefreshingWeighted 16%

SWE-bench Verified 2024 · updated September 2, 2026

2026

SWE-Rebench

SWE-Rebench

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

Fresh GitHub issues (rolling window)Code patch generationProfessional software engineering

CurrentWeighted 20%

Rolling 2026 window · updated September 2, 2026

2024

LiveCodeBench

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.

Continuously updated contest problemsCompetitive-programming evaluationCompetitive programming level

CurrentWeighted 38%

Rolling 2026 set · updated September 2, 2026

2026

LiveCodeBench v6

LiveCodeBench v6

LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6 results from being mixed into older or rolling LiveCodeBench windows.

Fresh programming problemsProvider-published v6 competitive programming resultsCompetitive programming level

CurrentDisplay only

LiveCodeBench v6 2026 · updated September 2, 2026

2025

LiveCodeBench v5

LiveCodeBench v5

LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside the rolling weighted lane.

July 2024 to May 2025 release windowProvider-published v5 competitive programming resultCompetitive programming level

CurrentDisplay only

LiveCodeBench v5 2025 · updated September 2, 2026

2026

LiveCodeBench Pass@1-COT

LiveCodeBench Pass@1 with Chain-of-Thought

This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.

DeepSeek-V4 report evaluation windowPass@1-COT competitive programming resultsCompetitive programming level

CurrentDisplay only

LiveCodeBench Pass@1-COT 2026 · updated September 2, 2026

2025

LiveCodeBench Pro

LiveCodeBench Pro

A harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with quarter-specific public leaderboards and difficulty-aware reporting.

Quarter-specific contest programming setsCompetitive programmingHigh-end contest programming

CurrentDisplay only

LiveCodeBench Pro 2025 · updated September 2, 2026

2026

FLTEval

FLTEval

A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

FLT project pull requestsLean 4 repository task completionFormal verification / proof engineering

CurrentDisplay only

FLTEval 2026 · updated September 2, 2026

2025

SWE-bench Pro

SWE-bench Pro

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

1,865 repository problemsRepository task completionLong-horizon professional engineering

CurrentWeighted 10%

SWE-bench Pro 2025 · updated September 2, 2026

2026

Senior SWE-Bench

Senior SWE-Bench

A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

Senior-level repository tasksAgentic software-engineering evaluationProfessional senior engineering

CurrentDisplay only

Senior SWE-Bench v2026.06 · updated September 2, 2026

2026

VulcanBench v3

VulcanBench v3

An open software-engineering benchmark built from real merged post-cutoff pull requests across Python, Rust, TypeScript, JavaScript, and Go repositories.

23 post-cutoff repository tasks in the v3 reportPass@1 with low, medium, and high effortProfessional multi-file software engineering

CurrentDisplay only

VulcanBench v3 2026 · updated September 2, 2026

2026

OpenHarmony Bench

OpenHarmony Bench v1.0

An app-level coding benchmark that asks DevEco Code configurations to implement observable behavior in buildable OpenHarmony ArkTS applications.

153 app-development and bug-fix tasksTask completion through DevEco CodeEnd-to-end OpenHarmony application development

CurrentDisplay only

OpenHarmony Bench 2026 · updated September 2, 2026

2026

VulcanBench CII v1

VulcanBench Coding Intelligence Index v1

A post-cutoff software-engineering benchmark with hidden functional tests and regression guards, reported for vendor coding-agent harnesses.

38 validated post-cutoff repository tasksPass@1 with vendor coding-agent harnessesMid-band frontier software engineering

CurrentDisplay only

VulcanBench CII v1 2026 · updated September 2, 2026

2026

FrontierCode 1.1 Main

FrontierCode 1.1 Main

Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

100 private Main tasks (150 in Extended)Repository task completion with maintainer rubricsFrontier coding-agent quality

CurrentDisplay only

FrontierCode 1.1 Main · updated September 2, 2026

2026

FrontierCode 1.1 Extended

FrontierCode 1.1 Extended

Cognition's 150-task Extended subset of the FrontierCode 1.1 software-engineering benchmark.

150 private software-engineering tasksRepository task completion with maintainer rubricsFrontier coding-agent quality

CurrentDisplay only

FrontierCode 1.1 Extended · updated September 2, 2026

2026

IDE-Bench

IDE-Bench

An 80-task software-engineering benchmark across eight repositories that tests whether autonomous IDE agents can explore, edit, run, and verify code changes end to end.

80 tasks across 8 repositoriesAutonomous IDE-agent task completion (pass@1)End-to-end software engineering

CurrentDisplay only

IDE-Bench 2026 · updated September 2, 2026

2025

App-Bench

App-Bench

A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.

6 full-stack app-building tasksBest-of-three one-shot feature completionProduction-style full-stack application generation

CurrentDisplay only

App-Bench 2025 · updated September 2, 2026

2026

SWE Multilingual

SWE Multilingual

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

Multilingual software-engineering tasksRepository task completionProfessional software engineering

CurrentDisplay only

SWE Multilingual 2026 · updated September 2, 2026

2025

SWE Multimodal

SWE-bench Multimodal

A multimodal variant of SWE-bench that adds visual context such as screenshots and design mockups to software engineering issue descriptions.

Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding

CurrentDisplay only

SWE Multimodal 2025 · updated September 2, 2026

2026

CursorBench

CursorBench

Cursor's current first-party benchmark for ambiguous, multi-file coding-agent tasks from real Cursor sessions.

Harder long-horizon agentic coding tasksCursor agent-loop evaluationProfessional agentic software engineering

CurrentDisplay only

CursorBench 2026 · updated September 2, 2026

2026

Multi-SWE Bench

Multi-SWE Bench

A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.

Multi-language repo tasksRepository task completionProfessional software engineering

CurrentDisplay only

Multi-SWE Bench 2026 · updated September 2, 2026

2026

VIBE-Pro

VIBE-Pro

A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.

Full project delivery tasksRepository-level implementation benchmarkEnd-to-end software delivery

CurrentDisplay only

VIBE-Pro 2026 · updated September 2, 2026

2026

Vibe Code Bench

Vibe Code Bench v1.1

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

End-to-end web application buildsFull-stack app implementation benchmarkEnd-to-end software delivery

CurrentDisplay only

Vibe Code Bench 2026 · updated September 2, 2026

2026

ProgramBench

ProgramBench: Can Language Models Rebuild Programs From Scratch?

A cleanroom software-engineering benchmark where agents receive only a compiled executable and documentation, then must architect and implement a complete codebase that reproduces the original program's behavior.

200 program reconstruction tasksCleanroom executable reimplementationFull-repository software architecture

CurrentDisplay only

ProgramBench 2026 · updated September 2, 2026

2026

PostTrain Bench

PostTrain Bench

A software-engineering benchmark for post-training infrastructure and implementation tasks, evaluated through the official Harbor implementation.

Post-training software-engineering tasksHarbor agent evaluationFrontier software engineering

CurrentDisplay only

PostTrain Bench 2026 · updated September 2, 2026

2026

FrontierSWE

FrontierSWE

An ultra-long-horizon software-engineering benchmark with open-ended implementation, performance, and research tasks designed to challenge frontier coding agents.

17 ultra-long-horizon engineering and research tasksMean@5, best@5, average rank, and dominanceUltra-long-horizon frontier software engineering

CurrentDisplay only

FrontierSWE 2026 · updated September 2, 2026

2026

FrontierSWE v2

FrontierSWE v2

A 34-task expansion of FrontierSWE for ultra-long-horizon engineering and research work that remains far from saturation.

34 ultra-long-horizon engineering and research tasksFive-trial mean task score (Mean@5), 0-100Ultra-long-horizon frontier software engineering

CurrentDisplay only

FrontierSWE v2 2026 · updated September 2, 2026

2026

SWE-Atlas Codebase QnA

SWE-Atlas Codebase QnA

A code-comprehension benchmark covering production repositories across Go, Python, C, and TypeScript.

124 codebase questions across 11 repositoriesMean pass@1Production codebase comprehension

CurrentDisplay only

SWE-Atlas Codebase QnA 2026 · updated September 2, 2026

2026

Bug Hunt Bench

Bug Hunt Bench

A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.

105 planted bugs across two production TypeScript repositoriesStrict planted bugs fixedBlind production-repository bug finding and repair

CurrentDisplay only

Bug Hunt Bench 2026 · updated September 2, 2026

2026

Kimi Code Bench v2

Kimi Code Bench v2

A Moonshot AI internal coding-agent benchmark for realistic software-engineering tasks across mainstream programming languages and production technology stacks.

Realistic coding-agent tasksCoding-agent pass rateProduction software engineering

CurrentDisplay only

Kimi Code Bench v2 2026 · updated September 2, 2026

2026

MLS-Bench Lite

MLS-Bench Lite

A 30-task subset of MLS-Bench that evaluates whether AI systems can invent generalizable and scalable machine-learning methods.

30 machine-learning research tasksAgentic ML task evaluationML research and systems engineering

CurrentDisplay only

MLS-Bench Lite 2026 · updated September 2, 2026

2026

PaperBench

PaperBench

A research-reproduction benchmark that asks agents to recreate the contributions of AI papers from the paper alone.

AI research-paper reproductionLong-horizon agent evaluationFrontier autonomous research and engineering

CurrentDisplay only

PaperBench 2026 · updated September 2, 2026

2026

QwenReactBench

QwenReactBench

Qwen's internal benchmark for building and rendering bilingual React projects across seven categories.

Bilingual React project constructionBradley-Terry/Elo ratingProduction frontend development

CurrentDisplay only

QwenReactBench 2026 · updated September 2, 2026

2026

NL2Repo

NL2Repo

A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

Natural language to repository tasksRepository understanding benchmarkSystem-level software comprehension

CurrentDisplay only

NL2Repo 2026 · updated September 2, 2026

2026

DSBench-FullStack

DeepSeek DSBench FullStack

DeepSeek's internal full-stack coding-agent benchmark.

Internal full-stack coding-agent tasksProvider-reported scoreFull-stack software engineering

CurrentDisplay only

DSBench-FullStack 2026 · updated September 2, 2026

2026

DSBench-Hard

DeepSeek DSBench Hard

DeepSeek's internal hard coding-agent benchmark.

Internal hard coding-agent tasksProvider-reported scoreAdvanced coding-agent challenges

CurrentDisplay only

DSBench-Hard 2026 · updated September 2, 2026

2026

React Native Evals

React Native Evals

An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.

React Native app implementation tasksFramework-specific app development evaluationProduction mobile app engineering

CurrentDisplay only

React Native Evals 2026 · updated September 2, 2026

2026

ReactBench

ReactBench v1

A coding-agent benchmark for realistic React work, with rubrics that check production concerns such as performance, accessibility, correctness, and code quality.

51 production React tasksPass@1 weighted rubric scoreProduction frontend engineering

CurrentDisplay only

ReactBench 2026 · updated September 2, 2026

2026

KernelBench

KernelBench Hard H100

An agentic GPU-kernel benchmark that measures how much of the hardware roofline a model's correct, audit-clean kernels reach on six demanding CUDA and Triton problems.

6 GPU-kernel optimization problemsMean peak fraction of hardware roofline over valid cellsAgentic GPU systems engineering

CurrentDisplay only

KernelBench 2026 · updated September 2, 2026

2026

Next.js Evals

AI Agent Evaluations for Next.js

A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.

24 Next.js code generation and migration tasksAgent task completion with withheld Vitest assertionsFramework-specific web application engineering

CurrentDisplay only

Next.js Evals 2026 · updated September 2, 2026

2026

SWE-bench Verified*

SWE-bench Verified (mini-swe-agent-v2)

A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.

Repository task completionAgent scaffold benchmarkProfessional software engineering

CurrentDisplay only

SWE-bench Verified* 2026 · updated September 2, 2026

2024

Spider 2.0-Lite

Spider 2.0-Lite

A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison.

Text-to-SQL queriesExecution accuracyEnterprise text-to-SQL

RefreshingDisplay only

Spider 2.0-Lite 2024 · updated September 2, 2026

2024

SciCode

Scientific Code Benchmark

SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.

80

RefreshingWeighted 16%

SciCode 2024 · updated September 2, 2026

2026

AA Coding Index

Artificial Analysis Coding Index

A display-only Artificial Analysis coding index.

Cross-benchmark coding indexAggregated model scoreDisplay-only external reference

CurrentDisplay only

AA Coding Index 2026 · updated September 2, 2026

2026

AA Coding Agents

Artificial Analysis Coding Agent Index

A display-only Artificial Analysis leaderboard for coding-agent systems, combining agent harnesses, host models, and execution settings across software-engineering benchmarks.

Composite over DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnAAverage pass@1 indexReal-world coding-agent workflows

CurrentDisplay only

AA Coding Agents 2026 · updated September 2, 2026

2026

AA-SciCode

Artificial Analysis SciCode

A display-only Artificial Analysis SciCode score.

Scientific coding subproblemsTask success rateScientific programming

CurrentDisplay only

AA-SciCode 2026 · updated September 2, 2026

2026

Terminal-Bench Hard

Terminal-Bench Hard

A display-only Artificial Analysis coding metric for agentic coding and terminal use on a harder Terminal-Bench slice.

Agentic coding and terminal tasksTask success rateProfessional software engineering

CurrentDisplay only

Terminal-Bench Hard 2026 · updated September 2, 2026

2026

VIBE V2

VIBE V2

A display-only MiniMax provider benchmark for end-to-end coding-agent and product-building tasks.

End-to-end coding-agent tasksTask success rateFrontier coding-agent workflows

CurrentDisplay only

VIBE V2 2026 · updated September 2, 2026

2026

SVG-Bench

SVG-Bench

A display-only provider benchmark for generating or manipulating SVG outputs from natural-language requirements.

SVG generation and editing tasksTask success rateVisual coding and structured graphics generation

CurrentDisplay only

SVG-Bench 2026 · updated September 2, 2026

2026

KernelBench Hard

KernelBench Hard

A display-only benchmark for difficult GPU kernel implementation and optimization tasks.

Hard GPU kernel coding tasksTask success rateSpecialized systems programming

CurrentDisplay only

KernelBench Hard 2026 · updated September 2, 2026

2026

GameDevBench

GameDevBench

Evaluates coding agents on 333 multimodal game-development tasks in Godot, spanning 2D graphics, 3D graphics, user interfaces, and gameplay logic.

333 tasks from 88 tutorialsPass@1 on the full task set with 95% confidence intervalsMultimodal game development in Godot 4.4.1

CurrentDisplay only

GameDevBench 2026 · updated September 2, 2026

2026

EdgeBench

EdgeBench

A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

Systems and software-engineering tasksTime-budgeted agent learning curvesLong-horizon engineering

CurrentDisplay only

EdgeBench 2026 · updated September 2, 2026

Reasoning(29 benchmarks)

View leaderboard
2026

ARC-AGI-1

ARC-AGI-1 Semi-Private Evaluation

ARC Prize fluid-intelligence benchmark using novel visual grid transformations.

Semi-private ARC-AGI-1 evaluation setVerified accuracyAbstract visual reasoning

CurrentDisplay only

ARC-AGI-1 2026 · updated September 2, 2026

2023

MuSR

Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

Multi-step reasoningNarrative-based reasoningComplex reasoning tasks

StaleDisplay only

MuSR 2023 · updated September 2, 2026

2022

BBH

BIG-Bench Hard

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

23 tasksMixed reasoning tasksAdvanced reasoning

StaleSaturatedDisplay only

BBH 2022 · updated September 2, 2026

2026

DROP

Discrete Reasoning Over Paragraphs

A reading-comprehension benchmark requiring discrete reasoning over paragraphs, reported in DeepSeek-V4 base-model evaluations.

Paragraph reasoning questionsF1Reading and numerical reasoning

CurrentDisplay only

DROP 2026 · updated September 2, 2026

2026

HellaSwag

HellaSwag

A commonsense natural-language inference benchmark reported in DeepSeek-V4 base-model evaluations.

Commonsense completion questionsExact matchCommonsense reasoning

CurrentDisplay only

HellaSwag 2026 · updated September 2, 2026

2026

WinoGrande

WinoGrande

A commonsense coreference benchmark reported in DeepSeek-V4 base-model evaluations.

Coreference resolution questionsExact matchCommonsense reasoning

CurrentDisplay only

WinoGrande 2026 · updated September 2, 2026

2026

CLUEWSC

CLUEWSC

A Chinese Winograd Schema Challenge benchmark reported in DeepSeek-V4 base-model evaluations.

Chinese coreference questionsExact matchChinese commonsense reasoning

CurrentDisplay only

CLUEWSC 2026 · updated September 2, 2026

2026

LisanBench

LisanBench

A word-chain reasoning benchmark that tests planning, recall, constraint following, and vocabulary depth by asking models to extend non-repeating edit-distance-1 chains.

50 starting words × 3 trialsDifficulty-weighted word-chain reasoningOpen-ended lexical planning

CurrentDisplay only

LisanBench 2026 · updated September 2, 2026

Conceptual Reasoning

Conceptual Reasoning Benchmark

Tests whether model judgments rank argumentative critiques in the same order as expert human ratings across philosophy, AI alignment, and other concept-heavy texts.

224 texts and 608 within-text critique pairsAverage pairwise-ranking loss against expert ratingsFuzzy, expert-rated argumentative reasoning

CurrentDisplay only

Early results snapshot · updated September 2, 2026

2026

Pencil Puzzle Bench

Pencil Puzzle Bench

A multi-step verifiable reasoning benchmark that evaluates whether models can solve pencil puzzles with unique solutions.

300 evaluation puzzlesDirect and agentic puzzle solve rateMulti-step verifiable reasoning

CurrentDisplay only

Pencil Puzzle Bench 2026 · updated September 2, 2026

2025

LongBench v2

LongBench v2

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

Long-context tasksExtended-context retrieval and reasoningHard long-context

CurrentWeighted 38%

LongBench v2 2025 · updated September 2, 2026

2025

MRCRv2

MRCRv2

A long-context benchmark for memory, retrieval, and multi-round coherence over large contexts.

Long-context retrievalMulti-round long-context evaluationHard long-context

CurrentWeighted 31%

MRCRv2 2025 · updated September 2, 2026

2026

MRCR v2 64K-128K

OpenAI MRCR v2 8-needle 64K-128K

MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.

8-needle retrieval tasksLong-context retrievalLong-context reasoning

CurrentDisplay only

MRCR v2 64K-128K 2026 · updated September 2, 2026

2026

MRCR v2 128K-256K

OpenAI MRCR v2 8-needle 128K-256K

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

8-needle retrieval tasksVery-long-context retrievalVery long-context reasoning

CurrentDisplay only

MRCR v2 128K-256K 2026 · updated September 2, 2026

2026

MRCR v2 256K-512K

OpenAI MRCR v2 8-needle 256K-512K

MRCR v2 slice focused on retrieval across 256K-512K-token contexts.

100 eight-needle retrieval examplesMean sequence-matcher ratioVery long-context retrieval

CurrentDisplay only

MRCR v2 256K-512K 2026 · updated September 2, 2026

2026

MRCR v2 512K-1M

OpenAI MRCR v2 8-needle 512K-1M

MRCR v2 slice focused on retrieval across 512K-1M-token contexts.

100 eight-needle retrieval examplesMean sequence-matcher ratioMillion-token retrieval

CurrentDisplay only

MRCR v2 512K-1M 2026 · updated September 2, 2026

2026

Graphwalks BFS 128K

Graphwalks BFS 0K-128K

Long-context graph traversal benchmark using breadth-first search tasks.

Graph traversal tasksLong-context graph reasoningAlgorithmic long-context reasoning

CurrentDisplay only

Graphwalks BFS 128K 2026 · updated September 2, 2026

2026

Graphwalks Parents 128K

Graphwalks parents 0-128K

Long-context benchmark for recovering parent relationships inside graph tasks.

Graph parent-retrieval tasksLong-context graph reasoningAlgorithmic long-context reasoning

CurrentDisplay only

Graphwalks Parents 128K 2026 · updated September 2, 2026

2026

MRCR 1M

MRCR 1M

A million-token MRCR long-context retrieval benchmark reported in DeepSeek-V4 model evaluations.

Million-token retrievalLong-context retrieval MMRMillion-token long context

CurrentDisplay only

MRCR 1M 2026 · updated September 2, 2026

2026

CorpusQA 1M

CorpusQA 1M

A million-token CorpusQA long-context question-answering benchmark reported in DeepSeek-V4 model evaluations.

Million-token corpus question answeringLong-context QA accuracyMillion-token long context

CurrentDisplay only

CorpusQA 1M 2026 · updated September 2, 2026

2025

ARC-AGI-2

Abstraction and Reasoning Corpus for AGI v2

A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.

Visual pattern completion and abstract reasoningGrid transformation puzzles with novel rulesExpert-level — hardest public reasoning benchmark

CurrentWeighted 31%

ARC-AGI 2 · updated September 2, 2026

2026

ARC-AGI-3

Abstraction and Reasoning Corpus for AGI v3

An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics through action and feedback.

Interactive game-like tasks with hidden rulesAgentic task completion under a capped evaluation budgetFrontier agentic reasoning

CurrentDisplay only

ARC-AGI 3 · updated September 2, 2026

2026

GeneBench-Pro

GeneBench-Pro

A multistage statistical-reasoning benchmark for genomics and biological-data analysis agents.

129 genomics statistical-analysis workflowsEval-level pass rate across dependent analysis decisionsLong-horizon scientific reasoning

CurrentDisplay only

GeneBench-Pro · updated September 2, 2026

2026

AI-Needle

AI-Needle

A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.

Long-context retrievalNeedle-in-a-haystack recallLong-context memory

CurrentDisplay only

AI-Needle 2026 · updated September 2, 2026

2023

GPQA Diamond

GPQA Diamond

The hardest subset of GPQA featuring the most challenging graduate-level science questions. Sometimes reported separately from the standard GPQA benchmark.

Expert-level science questionsMultiple choice questionsGraduate-level scientific reasoning

StaleDisplay only

GPQA Diamond 2023 · updated September 2, 2026

2026

AA-LCR

Artificial Analysis Long Context Reasoning

A display-only Artificial Analysis long-context reasoning evaluation.

Long-context reasoning tasksAccuracyLong-context reasoning

CurrentDisplay only

AA-LCR 2026 · updated September 2, 2026

2026

CritPt

Critical Physics Tasks

A display-only Artificial Analysis metric for research-level physics reasoning.

Research-level physics questionsAccuracyResearch-level physics reasoning

CurrentDisplay only

CritPt 2026 · updated September 2, 2026

2025

BullshitBench v2

BullshitBench v2

A benchmark that tests whether AI models challenge nonsensical, ill-posed, or logically flawed prompts instead of confidently generating incorrect answers. Measures the critical ability to push back on bad input.

Nonsensical and flawed prompts across multiple domainsPrompt challenge and refusal evaluationRobustness and critical reasoning

CurrentDisplay only

BullshitBench v2 2025 · updated September 2, 2026

2024

WildBench

WildBench

An automated evaluation framework using 1,000+ real-world user tasks covering reasoning, planning, coding, and creative writing. Highly correlated with Chatbot Arena human preference rankings.

1,024 real-world tasksReal-world task evaluationDiverse real-world scenarios

RefreshingDisplay only

WildBench 2024 · updated September 2, 2026

Multimodal & Grounded(65 benchmarks)

View leaderboard
2026

Chartography (no tools)

Chartography without tools

Professional chart understanding across 100 specialized chart types with expert-set answer tolerances.

100 specialized chart typesAccuracy without toolsProfessional chart reasoning

CurrentDisplay only

Chartography (no tools) 2026 · updated September 2, 2026

2026

Chartography (tools)

Chartography with image and code tools

Professional chart understanding with a container, standard libraries, and image cropping.

100 specialized chart typesAccuracy with toolsProfessional chart reasoning

CurrentDisplay only

Chartography (tools) 2026 · updated September 2, 2026

2026

BenchCAD Vision2Code (no tools)

BenchCAD Vision2Code voxel IoU without tools

Generates CadQuery code from multi-view renders and scores geometric similarity by voxel intersection-over-union.

1,000-file Vision2Code subsetVoxel IoUProgrammatic CAD generation

CurrentDisplay only

BenchCAD Vision2Code (no tools) 2026 · updated September 2, 2026

2026

BenchCAD Vision2Code (tools)

BenchCAD Vision2Code voxel IoU with tools

Generates CadQuery code from multi-view renders with image inspection and code-execution tools.

1,000-file Vision2Code subsetVoxel IoU with toolsProgrammatic CAD generation

CurrentDisplay only

BenchCAD Vision2Code (tools) 2026 · updated September 2, 2026

2026

GDP.pdf (no tools)

GDP.pdf mean criteria pass rate without tools

Professional document understanding over 100 real-world PDFs from ten domains.

100 professional document promptsMean criteria pass rateProfessional document reasoning

CurrentDisplay only

GDP.pdf (no tools) 2026 · updated September 2, 2026

2026

GDP.pdf (tools)

GDP.pdf mean criteria pass rate with tools

Professional document understanding with a container, standard libraries, and image cropping.

100 professional document promptsMean criteria pass rate with toolsProfessional document reasoning

CurrentDisplay only

GDP.pdf (tools) 2026 · updated September 2, 2026

2026

OfficeQA

OfficeQA

Grounded numerical reasoning over a corpus of historical U.S. Treasury Bulletin documents.

Historical Treasury Bulletin questionsAgentic grounded QA accuracyProfessional document reasoning

CurrentDisplay only

OfficeQA 2026 · updated September 2, 2026

2024

MMMU

Massive Multi-discipline Multimodal Understanding

A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.

Multimodal academic reasoningImage + text question answeringFrontier multimodal

RefreshingDisplay only

MMMU 2024 · updated September 2, 2026

2024

MMMU-Pro

Massive Multi-discipline Multimodal Understanding Pro

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Multimodal academic reasoningImage + text question answeringFrontier multimodal

RefreshingWeighted 45%

MMMU-Pro 2024 · updated September 2, 2026

2026

AA-MMMU-Pro

Artificial Analysis MMMU-Pro

A display-only Artificial Analysis MMMU-Pro score.

Multimodal academic reasoningImage + text question answeringFrontier multimodal

CurrentDisplay only

AA-MMMU-Pro 2026 · updated September 2, 2026

2025

OCRBench V2

OCRBench V2

A native OCR benchmark for reading text from images across multilingual scripts, low-quality scans, handwriting, structured layouts, charts, and screenshots.

Image OCR tasksAccuracyNative visual text understanding

CurrentDisplay only

OCRBench V2 2025 · updated September 2, 2026

2025

olmOCR

olmOCR-Bench

An end-to-end document understanding benchmark over long, layout-rich PDFs with tables, equations, headers, footnotes, and multi-column flows.

Layout-rich PDF understandingMean accuracyComplex document processing

CurrentDisplay only

olmOCR 2025 · updated September 2, 2026

2026

VoxPopuli WER

VoxPopuli-Cleaned-AA Word Error Rate

A speech-recognition benchmark on the cleaned Artificial Analysis VoxPopuli subset, reported as word error rate where lower is better.

Speech-to-text transcriptionWord error rateAudio speech recognition

CurrentDisplay only

VoxPopuli WER 2026 · updated September 2, 2026

2026

Design Arena Website

Design Arena Website Elo

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Website generation comparisonsEloDesign and website generation

CurrentDisplay only

Design Arena Website 2026 · updated September 2, 2026

2026

OfficeQA Pro

OfficeQA Pro

A benchmark for grounded reasoning over office-style documents, spreadsheets, charts, and business artifacts.

Document and spreadsheet tasksGrounded QA over office artifactsEnterprise grounded reasoning

CurrentWeighted 30%

OfficeQA Pro 2026 · updated September 2, 2026

2026

MathVision w/ Python

MathVision with Python

A tool-augmented MathVision variant that permits Python during visual mathematics reasoning.

Visual mathematics problems with PythonImage and mathematics reasoning with toolsAdvanced multimodal mathematics

CurrentDisplay only

MathVision w/ Python 2026 · updated September 2, 2026

2026

BabyVision w/ Python

BabyVision with Python

A Python-assisted BabyVision evaluation for fine-grained visual perception and grounded reasoning.

Visual perception tasks with PythonTool-augmented multimodal scoreFine-grained visual perception

CurrentDisplay only

BabyVision w/ Python 2026 · updated September 2, 2026

2026

ZeroBench w/ Python

ZeroBench_main with Python

A Python-assisted ZeroBench_main evaluation reported as pass@5.

Visual reasoning questions with PythonPass@5Tool-augmented visual reasoning

CurrentDisplay only

ZeroBench w/ Python 2026 · updated September 2, 2026

2026

WorldVQA ForceAnswer

WorldVQA ForceAnswer

A forced-answer WorldVQA variant for atomic visual world knowledge.

Atomic visual world-knowledge questionsForced-answer visual QAFine-grained visual knowledge

CurrentDisplay only

WorldVQA ForceAnswer 2026 · updated September 2, 2026

2026

OmniDocBench

OmniDocBench

A document-understanding benchmark for parsing and reasoning over complex document layouts.

Complex document-understanding tasksDocument-understanding scoreGrounded document reasoning

CurrentDisplay only

OmniDocBench 2026 · updated September 2, 2026

2026

PerceptionBench

PerceptionBench (Internal)

Moonshot AI's internal benchmark for atomic visual perception capabilities.

Internal atomic visual-perception tasksInternal evaluation scoreFine-grained visual perception

CurrentDisplay only

PerceptionBench 2026 · updated September 2, 2026

2026

MMMU-Pro w/ Python

MMMU-Pro with Python

Tool-augmented MMMU-Pro variant that allows Python assistance during multimodal reasoning.

Multimodal academic reasoningImage + text question answering with PythonFrontier multimodal

CurrentDisplay only

MMMU-Pro w/ Python 2026 · updated September 2, 2026

2026

OmniDocBench 1.5

OmniDocBench 1.5

A document understanding benchmark used in frontier-model comparison tables to measure extraction and grounded reasoning quality on complex documents.

Document understanding tasksDocument understanding benchmarkGrounded document reasoning

CurrentDisplay only

OmniDocBench 1.5 2026 · updated September 2, 2026

2026

Liquid Extract JSON Validity

Liquid image-to-JSON extraction JSON validity

A display-only Liquid AI extraction metric measuring the share of image-to-JSON outputs that parse as strict JSON.

Image-to-JSON extractionStrict JSON parseability rateStructured visual extraction

CurrentDisplay only

Liquid Extract JSON Validity 2026 · updated September 2, 2026

2026

Liquid Extract F1

Liquid image-to-JSON extraction schema consistency F1

A display-only Liquid AI extraction metric measuring field-name agreement between requested schema fields and extracted JSON fields.

Image-to-JSON extractionSchema field F1Structured visual extraction

CurrentDisplay only

Liquid Extract F1 2026 · updated September 2, 2026

2026

Liquid Extract VLM Judge

Liquid image-to-JSON extraction VLM judge score

A display-only Liquid AI extraction metric measuring judged agreement between extracted values and the source image.

Image-to-JSON extractionVLM-judged extraction accuracyStructured visual extraction

CurrentDisplay only

Liquid Extract VLM Judge 2026 · updated September 2, 2026

2026

RealWorldQA

RealWorldQA

A grounded visual QA benchmark focused on answering practical questions about real-world images and scenes.

Real-world visual question answeringImage-grounded QAGeneral visual reasoning

CurrentDisplay only

RealWorldQA 2026 · updated September 2, 2026

2026

Video-MME (with subtitle)

Video-MME with subtitle

A video understanding benchmark that allows subtitle access when answering multimodal questions about videos.

Video understandingVideo QA with subtitle contextMultimodal video reasoning

CurrentDisplay only

Video-MME (with subtitle) 2026 · updated September 2, 2026

2026

Video-MME (w/o subtitle)

Video-MME without subtitle

A stricter Video-MME setting that removes subtitle help and tests video understanding from visual and audio context alone.

Video understandingVideo QA without subtitle contextMultimodal video reasoning

CurrentDisplay only

Video-MME (w/o subtitle) 2026 · updated September 2, 2026

2024

Video-MME

Video-MME

A comprehensive benchmark for multimodal large language models on video understanding, covering temporal reasoning, perception, and question answering over videos.

Video understandingVideo QA and analysisBroad multimodal video reasoning

RefreshingDisplay only

Video-MME 2024 · updated September 2, 2026

2026

MathVision

MathVision

A visual mathematics benchmark that tests whether a model can solve math problems grounded in diagrams, equations, figures, and other visual inputs.

Visually grounded math problemsImage + math reasoningAdvanced multimodal mathematics

CurrentDisplay only

MathVision 2026 · updated September 2, 2026

2026

We-Math

We-Math

A multimodal math benchmark for visually grounded mathematical reasoning and answer generation.

Visually grounded math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics

CurrentDisplay only

We-Math 2026 · updated September 2, 2026

2026

DynaMath

DynaMath

A multimodal benchmark for dynamic mathematical reasoning over visual and structured inputs.

Dynamic visual math problemsMultimodal mathematical reasoningAdvanced multimodal mathematics

CurrentDisplay only

DynaMath 2026 · updated September 2, 2026

2026

MStar

MStar

A general visual question-answering benchmark used in provider tables for real-image reasoning quality.

Real-image visual QAImage-grounded QAGeneral visual reasoning

CurrentDisplay only

MStar 2026 · updated September 2, 2026

2026

ChatCVQA

ChatCVQA

A conversational visual QA benchmark that tests multi-turn grounded answering over images and documents.

Conversational visual QAMulti-turn image-grounded QAConversational multimodal reasoning

CurrentDisplay only

ChatCVQA 2026 · updated September 2, 2026

2026

MMLongBench-Doc

MMLongBench-Doc

A long-document multimodal benchmark for grounded reasoning over extended document contexts.

Long document understandingDocument-grounded reasoningLong-context document reasoning

CurrentDisplay only

MMLongBench-Doc 2026 · updated September 2, 2026

2026

CC-OCR

CC-OCR

An OCR-focused benchmark for reading and extracting text from visually complex documents and images.

Optical character recognitionText extraction from images and documentsDocument reading

CurrentDisplay only

CC-OCR 2026 · updated September 2, 2026

2026

AI2D_TEST

AI2D test split

A diagram understanding benchmark focused on scientific and educational visual question answering.

Diagram understandingDiagram-grounded QAStructured visual reasoning

CurrentDisplay only

AI2D_TEST 2026 · updated September 2, 2026

2026

CountBench

CountBench

A visual counting benchmark that tests whether a model can count objects and entities reliably in complex scenes.

Visual counting tasksImage-grounded countingFine-grained visual perception

CurrentDisplay only

CountBench 2026 · updated September 2, 2026

2026

RefCOCO (avg)

RefCOCO average

A referring-expression grounding benchmark averaged across RefCOCO variants to test whether a model can localize described objects correctly.

Referring-expression groundingGrounded visual localizationFine-grained visual grounding

CurrentDisplay only

RefCOCO (avg) 2026 · updated September 2, 2026

2026

ODINW13

ODINW13

A visual detection and grounding benchmark slice used to compare zero-shot object understanding across diverse domains.

Out-of-distribution object understandingDetection and groundingRobust visual grounding

CurrentDisplay only

ODINW13 2026 · updated September 2, 2026

2026

ERQA

ERQA

A grounded visual reasoning benchmark focused on evidence-based question answering over real images.

Evidence-based visual QAGrounded image reasoningGrounded multimodal reasoning

CurrentDisplay only

ERQA 2026 · updated September 2, 2026

2026

VideoMMMU

VideoMMMU

A video extension of MMMU-style multimodal reasoning over expert questions grounded in temporal media.

Video-grounded expert reasoningVideo + text reasoningFrontier multimodal video reasoning

CurrentDisplay only

VideoMMMU 2026 · updated September 2, 2026

2026

MLVU (M-Avg)

MLVU mean average

A multi-task video understanding benchmark averaged across MLVU categories.

General video understandingVideo QA and understandingBroad multimodal video reasoning

CurrentDisplay only

MLVU (M-Avg) 2026 · updated September 2, 2026

2026

LVBench

LVBench

A long-video understanding benchmark for retrieving and reasoning over information distributed across extended video inputs.

Long-form video question answeringLong-video understanding scoreExtended temporal reasoning

CurrentDisplay only

LVBench 2026 · updated September 2, 2026

2026

MMVU

Multimodal Multi-disciplinary Video Understanding

A benchmark for evaluating multimodal models on video understanding tasks across multiple disciplines, emphasizing temporal reasoning and comprehension over video content.

Video understandingVideo reasoning benchmarkMulti-disciplinary multimodal video reasoning

CurrentDisplay only

MMVU 2026 · updated September 2, 2026

2025

ScreenSpot Pro

ScreenSpot Pro

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

1,581 grounding instructionsStatic interface element localizationProfessional GUI grounding

CurrentDisplay only

ScreenSpot Pro 2025 · updated September 2, 2026

2026

TIR-Bench

TIR-Bench

A visual agent benchmark for interface reasoning and task execution over screenshots or software surfaces.

Visual agent and interface reasoningScreenshot-grounded task reasoningComputer-use visual reasoning

CurrentDisplay only

TIR-Bench 2026 · updated September 2, 2026

2026

GDPval-AA

GDPval-AA

An evaluation focused on professional domain expertise and task delivery quality in office-style knowledge work.

Professional office deliveryELO-style office benchmarkProfessional knowledge work

CurrentDisplay only

GDPval-AA 2026 · updated September 2, 2026

2026

MedXpertQA (MM)

MedXpertQA Multimodal

A multimodal medical multiple-choice benchmark covering clinical images such as X-rays, histology, and dermatology.

2,000 multimodal medical questionsMedical visual MCQClinical multimodal reasoning

CurrentDisplay only

MedXpertQA (MM) 2026 · updated September 2, 2026

2026

ZeroBench

ZeroBench

A multi-step visual reasoning benchmark with pass@5 reporting and optional tool use.

100 visual reasoning questionsMulti-step visual reasoningTool-augmented visual reasoning

CurrentDisplay only

ZeroBench 2026 · updated September 2, 2026

2026

Design2Code

Design2Code

A multimodal coding benchmark for turning visual designs into working frontend implementations.

Design-to-code tasksVisual input to frontend implementationMultimodal coding

CurrentDisplay only

Design2Code 2026 · updated September 2, 2026

2026

Flame-VLM-Code

Flame-VLM-Code

A vision-language coding benchmark for generating correct code from visual and multimodal inputs.

Multimodal coding tasksVision-language code generationMultimodal coding

CurrentDisplay only

Flame-VLM-Code 2026 · updated September 2, 2026

2026

Vision2Web

Vision2Web

A benchmark for converting visual references into functional web implementations.

Screenshot-to-web tasksVisual reference to web implementationMultimodal web generation

CurrentDisplay only

Vision2Web 2026 · updated September 2, 2026

2026

ImageMining

ImageMining

A multimodal retrieval and extraction benchmark over image-heavy task settings.

Visual retrieval tasksImage-grounded retrieval and extractionMultimodal retrieval

CurrentDisplay only

ImageMining 2026 · updated September 2, 2026

2026

MMSearch

MMSearch

A multimodal search benchmark for retrieval and grounded answering across mixed-media inputs.

Multimodal search tasksMixed-media retrieval and grounded answeringMultimodal search

CurrentDisplay only

MMSearch 2026 · updated September 2, 2026

2026

MMSearch-Plus

MMSearch-Plus

A harder MMSearch variant for multimodal retrieval and grounded tool-use workflows.

Hard multimodal search tasksAdvanced mixed-media retrieval benchmarkAdvanced multimodal search

CurrentDisplay only

MMSearch-Plus 2026 · updated September 2, 2026

2026

SimpleVQA

SimpleVQA

A visual question answering benchmark focused on straightforward image-grounded understanding.

Visual QA tasksImage-grounded question answeringGeneral visual understanding

CurrentDisplay only

SimpleVQA 2026 · updated September 2, 2026

2026

Facts-VLM

Facts-VLM

A grounded multimodal factuality benchmark for evidence-linked answer correctness.

Grounded factuality tasksEvidence-linked multimodal factualityGrounded multimodal factuality

CurrentDisplay only

Facts-VLM 2026 · updated September 2, 2026

2026

V*

V*

A vision-centric benchmark for high-level multimodal reasoning and perception quality.

Frontier multimodal reasoning tasksVision-centric reasoning benchmarkFrontier multimodal

CurrentDisplay only

V* 2026 · updated September 2, 2026

2024

CharXiv

CharXiv Reasoning

A scientific chart reasoning benchmark that tests whether models can understand, interpret, and reason about complex scientific visualizations including plots, diagrams, and data charts.

Scientific chart reasoningChart understanding and reasoningScientific visualization reasoning

RefreshingWeighted 25%

CharXiv 2024 · updated September 2, 2026

2024

CharXiv w/o tools

CharXiv Reasoning without tools

Tool-free variant of CharXiv that isolates raw visual reasoning ability without code execution or tool augmentation.

Scientific chart reasoning (tool-free)Chart understanding without toolsScientific visualization reasoning

RefreshingDisplay only

CharXiv w/o tools 2024 · updated September 2, 2026

2026

BabyVision

BabyVision

A multimodal benchmark for fine-grained visual perception and grounded reasoning tasks.

Visual perception tasksMultimodal visual reasoningFine-grained visual perception

CurrentDisplay only

BabyVision 2026 · updated September 2, 2026

2025

SWE-bench Multimodal

SWE-bench Multimodal

A multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software engineering issue descriptions, testing whether models can leverage visual information for code generation.

Multimodal software engineering tasksCode patch generation with visual contextFrontier multimodal coding

CurrentDisplay only

SWE-bench Multimodal 2025 · updated September 2, 2026

2026

Blueprint-Bench 2

Blueprint-Bench 2

An agentic spatial reasoning benchmark reported as a normalized score.

Spatial reasoning from blueprintsNormalized scoreAgentic spatial reasoning

CurrentDisplay only

Blueprint-Bench 2 2026 · updated September 2, 2026

Knowledge(48 benchmarks)

View leaderboard
2026

HealthBench (raw)

HealthBench raw score

Raw score on realistic multi-turn healthcare conversations graded against expert-written rubrics.

5,000 multi-turn patient conversationsRaw rubric scoreRealistic healthcare conversations

CurrentDisplay only

HealthBench (raw) 2026 · updated September 2, 2026

2026

HealthBench (length-adjusted)

HealthBench length-adjusted score

HealthBench score after applying a verbosity penalty to model responses.

5,000 multi-turn patient conversationsLength-adjusted rubric scoreRealistic healthcare conversations

CurrentDisplay only

HealthBench (length-adjusted) 2026 · updated September 2, 2026

2026

HealthBench Professional (raw)

HealthBench Professional raw score

Raw score on physician-authored clinical consult, documentation, and research conversations.

525 physician-authored conversationsRaw rubric scoreProfessional clinical tasks

CurrentDisplay only

HealthBench Professional (raw) 2026 · updated September 2, 2026

2026

BioMysteryBench (human-solvable)

BioMysteryBench Human Solvable

Computational biology challenges that independent human experts were able to solve.

Human-solvable computational biology investigationsTask scoreExpert computational biology

CurrentDisplay only

BioMysteryBench (human-solvable) 2026 · updated September 2, 2026

2026

BioMysteryBench (human-difficult)

BioMysteryBench Human Difficult

Computational biology challenges with objective answers that remained unsolved by independent human experts.

Human-difficult computational biology investigationsTask scoreFrontier computational biology

CurrentDisplay only

BioMysteryBench (human-difficult) 2026 · updated September 2, 2026

2026

SpatialBench Verified

LatchBio SpatialBench Verified

Analysis of spatial transcriptomics data across externally validated biological problems.

115 externally validated spatial transcriptomics problemsTask scoreProfessional bioinformatics

CurrentDisplay only

SpatialBench Verified 2026 · updated September 2, 2026

2026

SingleCellBench

LatchBio SingleCellBench

Single-cell RNA sequencing analysis tasks spanning common bioinformatics workflows.

195 single-cell RNA sequencing problemsTask scoreProfessional bioinformatics

CurrentDisplay only

SingleCellBench 2026 · updated September 2, 2026

2026

ProteinGym Hard

ProteinGym Hard

Predicts mutation effects by ranking mutant protein sequences against wild type and comparing against laboratory measurements.

Hard protein mutation-effect ranking tasksRank correlationComputational protein science

CurrentDisplay only

ProteinGym Hard 2026 · updated September 2, 2026

2026

Protein Design

Anthropic Protein Design evaluation

Generates novel protein sequences under family, topology, globularity, and structural-motif constraints.

Constrained protein-sequence design tasksComposite scoreComputational protein design

CurrentDisplay only

Protein Design 2026 · updated September 2, 2026

2026

Organic chemistry V2

Anthropic Organic Chemistry V2 evaluation

Chemistry tasks covering spectroscopy, synthesis planning, reaction prediction, and chemical structure images.

Organic chemistry reasoning tasksTask scoreExpert organic chemistry

CurrentDisplay only

Organic chemistry V2 2026 · updated September 2, 2026

2026

Protocols (troubleshooting)

Molecular Biology Protocols Troubleshooting

Detects and fixes errors in molecular-biology protocols using document, code, and web-search tools.

Molecular-biology protocol troubleshootingTask scoreExpert laboratory protocols

CurrentDisplay only

Protocols (troubleshooting) 2026 · updated September 2, 2026

2026

Protocols (understanding)

Benchling Molecular Biology Protocols Understanding

Extends online molecular-biology protocols in additional directions using document and web-search tools.

Molecular-biology protocol extensionTask scoreExpert laboratory protocols

CurrentDisplay only

Protocols (understanding) 2026 · updated September 2, 2026

2026

AA Openness Index

Artificial Analysis Openness Index

A display-only Artificial Analysis model-openness index.

Model openness assessmentIndex scoreDisplay-only external reference

CurrentDisplay only

AA Openness Index 2026 · updated September 2, 2026

2026

AA MMLU-Pro

Artificial Analysis MMLU-Pro

An independently evaluated MMLU-Pro result from Artificial Analysis.

Professional multi-subject questionsAccuracyProfessional knowledge and reasoning

CurrentDisplay only

AA MMLU-Pro 2026 · updated September 2, 2026

2020

MMLU

Massive Multitask Language Understanding

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

57 subjectsMultiple choice questionsElementary to professional level

StaleSaturatedDisplay only

MMLU · updated September 2, 2026

2023

GPQA

Graduate-Level Google-Proof Q&A

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.

448 questionsMultiple choice questionsGraduate level

RefreshingWeighted 7%

GPQA Diamond · updated September 2, 2026

2026

GPQA-D

GPQA Diamond

A display-only GPQA Diamond reference from provider comparison charts.

Graduate-level science questionsMultiple choice questionsGraduate level

CurrentDisplay only

GPQA-D 2026 · updated September 2, 2026

2025

SuperGPQA

SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines

An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.

285 disciplinesMultiple choice questionsGraduate level

CurrentWeighted 7%

SuperGPQA 2025 · updated September 2, 2026

2024

MMLU-Pro

Massive Multitask Language Understanding Professional

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Multiple subjects10-way multiple choiceProfessional level

RefreshingWeighted 30%

MMLU-Pro · updated September 2, 2026

2026

AGIEval

AGIEval

A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.

General academic and professional exam questionsExact matchGeneral knowledge

CurrentDisplay only

AGIEval 2026 · updated September 2, 2026

2025

HLE

Humanity's Last Exam

An expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol differences visible because tool-assisted and closed-book HLE runs answer different questions.

Expert-level questionsOpen-ended and multiple choiceFrontier expert level

CurrentWeighted 45%

Humanity's Last Exam · updated September 2, 2026

2026

HLE-Verified

HLE-Verified

A verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions before evaluating frontier expert reasoning.

1,811 verified or revised expert questionsFull-set accuracyFrontier multidisciplinary expert reasoning

CurrentDisplay only

HLE-Verified 2026 · updated September 2, 2026

2026

LABBench2

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

A benchmark of realistic biology-research tasks involving literature, figures, tables, databases, and bioinformatics files.

Nearly 1,900 biology-research tasksAggregate accuracyReal-world biology research

CurrentDisplay only

LABBench2 2026 · updated September 2, 2026

2026

FrontierScience

FrontierScience

A benchmark for research-level scientific reasoning, designed to separate frontier models on difficult science tasks that mix domain knowledge with deep reasoning.

Research-level science tasksScientific reasoning benchmarkResearch frontier

CurrentDisplay only

FrontierScience 2026 · updated September 2, 2026

2026

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index

A display-only intelligence index published by Artificial Analysis that aggregates provider-reported and benchmark-derived signals into a single model-level score.

Cross-benchmark intelligence indexAggregated model scoreDisplay-only external reference

CurrentDisplay only

Artificial Analysis Intelligence Index 2026 · updated September 2, 2026

2026

AA-GPQA Diamond

Artificial Analysis GPQA Diamond

A display-only Artificial Analysis GPQA Diamond score.

Graduate-level science questionsAccuracyGraduate-level science reasoning

CurrentDisplay only

AA-GPQA Diamond 2026 · updated September 2, 2026

2026

AA-HLE

Artificial Analysis Humanity's Last Exam

A display-only Artificial Analysis Humanity's Last Exam score.

Expert-level questionsAccuracyFrontier expert reasoning

CurrentDisplay only

AA-HLE 2026 · updated September 2, 2026

2026

AA-Omniscience Index

Artificial Analysis Omniscience Index

A display-only Artificial Analysis factual knowledge index.

Knowledge questionsIndex scoreBroad factual knowledge

CurrentDisplay only

AA-Omniscience Index 2026 · updated September 2, 2026

2026

AA-Omniscience Accuracy

Artificial Analysis Omniscience Accuracy

A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.

Knowledge questionsAccuracyBroad knowledge

CurrentDisplay only

AA-Omniscience Accuracy 2026 · updated September 2, 2026

2026

AA-Omniscience Hallucination Rate

Artificial Analysis Omniscience Hallucination Rate

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

Knowledge questionsHallucination rateFactuality

CurrentDisplay only

AA-Omniscience Hallucination Rate 2026 · updated September 2, 2026

2024

SimpleQA

Measuring Short-Form Factuality in Large Language Models

A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.

Factual questionsShort-form Q&AFactual accuracy focused

RefreshingWeighted 11%

SimpleQA 2024 · updated September 2, 2026

2026

Chinese-SimpleQA

Chinese-SimpleQA

A Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.

Chinese factual questionsShort-form factual QAFactual accuracy focused

CurrentDisplay only

Chinese-SimpleQA 2026 · updated September 2, 2026

2018

OpenBookQA

OpenBookQA

A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.

Elementary science questions4-way multiple choiceElementary science reasoning

StaleDisplay only

OpenBookQA 2018 · updated September 2, 2026

2026

HealthBench Hard

HealthBench Hard

A harder subset of OpenAI's HealthBench for evaluating open-ended medical and health reasoning with rubric-based grading.

1,000 health promptsOpen-ended health evaluationAdvanced health reasoning

CurrentDisplay only

HealthBench Hard 2026 · updated September 2, 2026

2026

HealthBench Professional

HealthBench Professional

An open benchmark for clinician-facing model responses across care consult, writing and documentation, and medical research tasks.

Clinician chat tasksRubric-graded open-ended responsesProfessional clinical workflows

CurrentDisplay only

HealthBench Professional 2026 · updated September 2, 2026

2026

MedXpertQA (Text)

MedXpertQA Text

A medical multiple-choice benchmark spanning many specialties with 10 answer options per question.

2,450 medical multiple-choice questionsMedical MCQProfessional medical knowledge

CurrentDisplay only

MedXpertQA (Text) 2026 · updated September 2, 2026

2026

FrontierScience Research

FrontierScience Research

A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.

Scientific research problemsResearch evaluationFrontier scientific research

CurrentDisplay only

FrontierScience Research 2026 · updated September 2, 2026

2021

TruthfulQA

TruthfulQA

A benchmark designed to measure whether language models produce truthful answers instead of repeating common misconceptions or misleading falsehoods.

Truthfulness and misconception resistanceQuestion answeringHallucination and factuality stress test

StaleDisplay only

TruthfulQA 2021 · updated September 2, 2026

2026

HLE w/o tools

Humanity's Last Exam without tools

Tool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.

Expert-level questionsTool-free expert QAFrontier expert level

CurrentDisplay only

HLE w/o tools 2026 · updated September 2, 2026

2026

MMLU-Pro (Arcee)

MMLU-Pro first-party comparison snapshot

A display-only MMLU-Pro reference from Arcee AI's Trinity-Large-Thinking launch chart.

Professional academic QA10-way multiple choiceProfessional level

CurrentDisplay only

MMLU-Pro (Arcee) 2026 · updated September 2, 2026

2026

MMLU-Redux

MMLU-Redux

A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.

Broad academic QAMultiple choice questionsAdvanced general knowledge

CurrentDisplay only

MMLU-Redux 2026 · updated September 2, 2026

2026

MMMLU

MMMLU

A multilingual MMLU-style benchmark reported in provider evaluation tables.

Multilingual academic QAExact matchBroad multilingual knowledge

CurrentDisplay only

MMMLU 2026 · updated September 2, 2026

2023

C-Eval

C-Eval

A Chinese-language academic and professional benchmark spanning humanities, social science, STEM, and applied subjects.

Chinese academic and professional examsMultiple choice questionsHigh school to professional level

StaleDisplay only

C-Eval 2023 · updated September 2, 2026

2026

CMMLU

Chinese Massive Multitask Language Understanding

A Chinese multitask academic benchmark reported in DeepSeek-V4 base-model evaluations.

Chinese academic QAExact matchBroad Chinese knowledge

CurrentDisplay only

CMMLU 2026 · updated September 2, 2026

2026

MultiLoKo

MultiLoKo

A multilingual/localized knowledge benchmark reported in DeepSeek-V4 base-model evaluations.

Localized multilingual knowledge questionsExact matchMultilingual knowledge

CurrentDisplay only

MultiLoKo 2026 · updated September 2, 2026

2026

FACTS Parametric

FACTS Parametric

A parametric factuality benchmark reported in DeepSeek-V4 base-model evaluations.

Parametric factual recallExact matchFactual accuracy focused

CurrentDisplay only

FACTS Parametric 2026 · updated September 2, 2026

2026

TriviaQA

TriviaQA

A reading and trivia question-answering benchmark reported in DeepSeek-V4 base-model evaluations.

Trivia and reading-comprehension QAExact matchGeneral factual QA

CurrentDisplay only

TriviaQA 2026 · updated September 2, 2026

2025

FinanceArena

FinanceArena — FinanceQA Assumption-Based

An AfterQuery benchmark of open-ended financial analysis that requires models to read financial data, make assumptions, and return exact answers.

Professional financial-analysis questionsOpen-ended financial QA with exact-match gradingProfessional finance reasoning

CurrentDisplay only

FinanceArena 2025 · updated September 2, 2026

Multilingual(13 benchmarks)

View leaderboard
2024

GMMLU

Global MMLU

MMLU-style knowledge evaluation across 42 high- and low-resource languages.

Knowledge questions across 42 languagesAverage accuracyMultilingual knowledge

RefreshingDisplay only

GMMLU 2024 · updated September 2, 2026

2024

MILU

Multi-task Indic Language Understanding Benchmark

Culturally grounded knowledge comprehension across ten Indic languages and English.

Knowledge tasks across 11 languagesAverage accuracyMultilingual Indic knowledge

RefreshingDisplay only

MILU 2024 · updated September 2, 2026

2026

AA Global-MMLU-Lite

Artificial Analysis Global-MMLU-Lite

An independently evaluated multilingual knowledge result from Artificial Analysis.

Multilingual knowledge questionsAccuracyMultilingual professional knowledge

CurrentDisplay only

AA Global-MMLU-Lite 2026 · updated September 2, 2026

2022

MGSM

Multilingual Grade School Math

A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.

250 problems × 11 languagesMath word problemsGrade school math, multilingual

StaleDisplay only

MGSM 2022 · updated September 2, 2026

2025

MMLU-ProX

MMLU-ProX

A multilingual extension of professional-level academic evaluation across many languages.

Multilingual professional QAMultilingual multiple choiceProfessional multilingual

CurrentWeighted 100%

MMLU-ProX 2025 · updated September 2, 2026

2026

NOVA-63

NOVA-63

A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.

Broad multilingual evaluationCross-lingual benchmarkBroad multilingual capability

CurrentDisplay only

NOVA-63 2026 · updated September 2, 2026

2026

INCLUDE

INCLUDE

A multilingual benchmark used in provider tables to measure inclusive language coverage and cross-lingual understanding beyond common high-resource languages.

Cross-lingual understandingMultilingual benchmarkBroad multilingual capability

CurrentDisplay only

INCLUDE 2026 · updated September 2, 2026

2026

PolyMath

PolyMath

A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

Multilingual math problemsCross-lingual mathematical reasoningAdvanced multilingual reasoning

CurrentDisplay only

PolyMath 2026 · updated September 2, 2026

2026

VWT2k-lite

VWT2k-lite

A lighter multilingual benchmark slice published in provider tables for broad cross-lingual transfer and understanding.

Multilingual transfer tasksCross-lingual benchmarkBroad multilingual capability

CurrentDisplay only

VWT2k-lite 2026 · updated September 2, 2026

2026

MAXIFE

MAXIFE

A multilingual instruction-following and understanding benchmark row published in Qwen's launch comparisons.

Multilingual instruction followingCross-lingual benchmarkAdvanced multilingual instruction following

CurrentDisplay only

MAXIFE 2026 · updated September 2, 2026

2025

SWE Multilingual

SWE-bench Multilingual

A multilingual extension of SWE-bench covering 300 problems across 9 programming languages, testing code generation and bug fixing beyond Python.

300 problems across 9 languagesMulti-language code patch generationProfessional multilingual software engineering

CurrentDisplay only

SWE Multilingual 2025 · updated September 2, 2026

2026

NanoBEIR Multilingual

NanoBEIR Multilingual Extended

A display-only multilingual retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using NDCG@10 across 11 languages.

Multilingual document retrievalNDCG@10 averageMultilingual retrieval

CurrentDisplay only

NanoBEIR Multilingual 2026 · updated September 2, 2026

2026

MKQA-11

MKQA-11 multilingual retrieval

A display-only multilingual QA retrieval benchmark reported by Liquid AI for LFM2.5 retriever models, using Recall@20 across 11 languages.

Cross-lingual open-domain QA retrievalRecall@20 averageMultilingual retrieval

CurrentDisplay only

MKQA-11 2026 · updated September 2, 2026

Instruction Following(4 benchmarks)

View leaderboard

Mathematics(32 benchmarks)

View leaderboard
2026

IMO 2026

International Mathematical Olympiad 2026

Proof-based olympiad performance on all six IMO 2026 problems.

6 proof-based problemsOfficial-style proof scoreInternational olympiad mathematics

CurrentDisplay only

IMO 2026 2026 · updated September 2, 2026

2026

RiemannBench (no tools)

RiemannBench without tools

Research-level mathematics problems with unique programmatically verified closed-form answers.

25 private research-level mathematics problemsAccuracy without toolsResearch mathematics

CurrentDisplay only

RiemannBench (no tools) 2026 · updated September 2, 2026

2026

RiemannBench (tools)

RiemannBench with tools

Research-level mathematics problems with unique programmatically verified closed-form answers and tool access.

25 private research-level mathematics problemsAccuracy with toolsResearch mathematics

CurrentDisplay only

RiemannBench (tools) 2026 · updated September 2, 2026

2026

ArXivMath Jun. 2026 (no tools)

ArXivMath June 2026 without tools

Final-answer research mathematics problems drawn from recent arXiv abstracts.

49 recent research-mathematics problemsFinal-answer accuracyResearch mathematics

CurrentDisplay only

ArXivMath Jun. 2026 (no tools) 2026 · updated September 2, 2026

2026

ArXivMath Jun. 2026 (tools)

ArXivMath June 2026 with tools

Final-answer research mathematics problems drawn from recent arXiv abstracts with tool access.

49 recent research-mathematics problemsFinal-answer accuracy with toolsResearch mathematics

CurrentDisplay only

ArXivMath Jun. 2026 (tools) 2026 · updated September 2, 2026

2026

AA AIME 2025

Artificial Analysis AIME 2025

An independently evaluated AIME 2025 result from Artificial Analysis.

30 AIME 2025 problemsAccuracyOlympiad mathematics

CurrentDisplay only

AA AIME 2025 2026 · updated September 2, 2026

2026

AA MATH-500

Artificial Analysis MATH-500

An independently evaluated MATH-500 result from Artificial Analysis.

500 competition mathematics problemsAccuracyHigh school to undergraduate mathematics

CurrentDisplay only

AA MATH-500 2026 · updated September 2, 2026

2023

AIME 2023

American Invitational Mathematics Examination 2023

A 15-question, 3-hour examination where each answer is an integer from 000 to 999. Serves as the intermediate step between AMC 10/12 and the USA Mathematical Olympiad (USAMO).

15 problemsInteger answers 000-999High school olympiad level

StaleDisplay only

AIME 2023 2023 · updated September 2, 2026

2024

AIME 2024

American Invitational Mathematics Examination 2024

The 2024 edition of AIME, maintaining the same format of 15 challenging mathematics problems with integer answers from 000 to 999.

15 problemsInteger answers 000-999High school olympiad level

RefreshingDisplay only

AIME 2024 2024 · updated September 2, 2026

2025

AIME 2025

American Invitational Mathematics Examination 2025

The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.

15 problemsInteger answers 000-999High school olympiad level

CurrentDisplay only

AIME 2025 · updated September 2, 2026

2026

GSM8K

Grade School Math 8K

A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

Grade-school math word problemsExact matchGrade-school math

CurrentDisplay only

GSM8K 2026 · updated September 2, 2026

2026

MATH

MATH

A competition-style mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

Competition math problemsExact matchAdvanced math reasoning

CurrentDisplay only

MATH 2026 · updated September 2, 2026

2026

CMath

CMath

A Chinese mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

Chinese math problemsExact matchMath reasoning

CurrentDisplay only

CMath 2026 · updated September 2, 2026

2026

AIME25 (Arcee)

AIME25 first-party comparison snapshot

A display-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart.

15 problemsInteger answers 000-999High school olympiad level

CurrentDisplay only

AIME25 (Arcee) 2026 · updated September 2, 2026

2023

HMMT Feb 2023

Harvard-MIT Mathematics Tournament February 2023

A prestigious high school mathematics competition hosted jointly by Harvard and MIT, featuring challenging problems across various mathematical disciplines.

Tournament problemsCompetition mathematicsHigh school olympiad level

StaleDisplay only

HMMT Feb 2023 2023 · updated September 2, 2026

2024

HMMT Feb 2024

Harvard-MIT Mathematics Tournament February 2024

The 2024 February edition of the Harvard-MIT Mathematics Tournament, continuing the tradition of challenging high school mathematics competition.

Tournament problemsCompetition mathematicsHigh school olympiad level

RefreshingDisplay only

HMMT Feb 2024 2024 · updated September 2, 2026

2025

HMMT Feb 2025

Harvard-MIT Mathematics Tournament February 2025

The most recent February edition of the Harvard-MIT Mathematics Tournament, featuring the latest challenging problems in competitive mathematics.

Tournament problemsCompetition mathematicsHigh school olympiad level

CurrentDisplay only

HMMT Feb 2025 2025 · updated September 2, 2026

2025

BRUMO 2025

Bulgarian Mathematical Olympiad 2025

A challenging mathematical olympiad competition featuring problems that test advanced mathematical reasoning and problem-solving skills at the olympiad level.

Olympiad problemsMathematical olympiadMathematical olympiad level

CurrentDisplay only

BRUMO 2025 2025 · updated September 2, 2026

2021

MATH-500

MATH-500 Problem Set

A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

500 problemsFree-form mathematical answersHigh school to undergraduate

StaleDisplay only

MATH-500 2021 · updated September 2, 2026

2026

AIME26

AIME 2026

A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.

Competition math problemsShort-answer mathematicsOlympiad-style mathematics

CurrentWeighted 25%

AIME26 2026 · updated September 2, 2026

2026

IPhO 2025 (Theory)

International Physics Olympiad 2025 (Theory)

The three official theory problems from the 2025 International Physics Olympiad, scored with blinded human evaluation.

3 olympiad theory problemsPhysics olympiad theoryInternational olympiad physics

CurrentDisplay only

IPhO 2025 (Theory) 2026 · updated September 2, 2026

2025

HMMT Feb 2025

Harvard-MIT Mathematics Tournament February 2025

A February 2025 HMMT slice used in exact-value provider tables for advanced contest-math reasoning.

Competition math problemsContest mathematicsOlympiad-style mathematics

CurrentDisplay only

HMMT Feb 2025 2025 · updated September 2, 2026

2025

HMMT Nov 2025

Harvard-MIT Mathematics Tournament November 2025

A November 2025 HMMT slice for high-end mathematical reasoning comparisons.

Competition math problemsContest mathematicsOlympiad-style mathematics

CurrentDisplay only

HMMT Nov 2025 2025 · updated September 2, 2026

2026

HMMT Feb 2026

Harvard-MIT Mathematics Tournament February 2026

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Competition math problemsContest mathematicsOlympiad-style mathematics

CurrentWeighted 25%

HMMT Feb 2026 2026 · updated September 2, 2026

2026

IMOAnswerBench

IMOAnswerBench

A challenging mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

Advanced mathematical answer generationPass@1 math benchmarkOlympiad-level mathematics

CurrentDisplay only

IMOAnswerBench 2026 · updated September 2, 2026

2026

Apex

Apex

A high-difficulty mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning

CurrentDisplay only

Apex 2026 · updated September 2, 2026

2026

Apex Shortlist

Apex Shortlist

A shortlist subset of the Apex mathematical reasoning benchmark reported in DeepSeek-V4 model evaluations.

Advanced mathematical reasoningPass@1 math benchmarkFrontier math reasoning

CurrentDisplay only

Apex Shortlist 2026 · updated September 2, 2026

2026

MMAnswerBench

MMAnswerBench

A multimodal mathematical reasoning benchmark that tests whether models can answer visually grounded math questions correctly.

Multimodal math questionsVisual and structured mathematical QAAdvanced mathematical reasoning

CurrentDisplay only

MMAnswerBench 2026 · updated September 2, 2026

2024

FrontierMath (legacy)

FrontierMath legacy aggregate

Legacy FrontierMath values retained for historical model pages. This field is not used in current rankings because it can mix prior benchmark versions and slices.

Historical aggregateOpen-ended mathematical reasoning with tool accessResearch-level mathematics

RefreshingDisplay only

FrontierMath (legacy) 2024 · updated September 2, 2026

2026

FrontierMath v2 (Tiers 1-3)

FrontierMath v2 Tiers 1-3

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

295 private advanced mathematics problemsPython-enabled iterative mathematical problem solvingFrom olympiad-plus to early research mathematics

CurrentWeighted 30%

FrontierMath v2 (Tiers 1-3) 2026 · updated September 2, 2026

2026

FrontierMath v2 (Tier 4)

FrontierMath v2 Tier 4

Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.

43 private extreme-difficulty mathematics problemsPython-enabled iterative mathematical problem solvingResearch-level mathematics requiring hours or days of expert work

CurrentWeighted 10%

FrontierMath v2 (Tier 4) 2026 · updated September 2, 2026

2026

USAMO 2026

United States of America Mathematical Olympiad 2026

The premier US mathematical olympiad competition, featuring proof-based problems that require deep mathematical insight and rigorous argumentation at the highest competition level.

6 proof-based problemsMathematical proof constructionInternational olympiad level

CurrentWeighted 10%

USAMO 2026 2026 · updated September 2, 2026

2024

KMMLU

Korean Massive Multitask Language Understanding

Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.

35,030 questionsMultiple choice questionsElementary to professional level in Korean

RefreshingDisplay only

KMMLU 2024 · updated September 2, 2026

2025

KMMLU-Hard

KMMLU-Hard

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

~5,000 questionsMultiple choice questionsAdvanced Korean reasoning

CurrentDisplay only

KMMLU-Hard 2025 · updated September 2, 2026

KMMLU-Redux

KMMLU-Redux

Cleaned KMMLU from national technical qualification exams, with errors removed, decontaminated, and deduplicated.

~3,500 questionsTechnical multiple choiceIndustrial/technical

RefreshingDisplay only

KMMLU-Redux · updated September 2, 2026

KMMLU-Pro

KMMLU-Pro

Korean National Professional Licensure exams evaluating professional-grade knowledge.

~2,500 questionsProfessional licensure examsProfessional

RefreshingDisplay only

KMMLU-Pro · updated September 2, 2026

CLIcK

Cultural and Linguistic Intelligence in Korean

Evaluates Korean culture and linguistics.

1,995 questionsCultural/linguistic QAKorean cultural nuances

RefreshingDisplay only

CLIcK · updated September 2, 2026

KoBALT

Korean Benchmark for Advanced Linguistic Tasks

Evaluates advanced Korean linguistic competence.

Linguistics questionsAdvanced linguisticsAdvanced linguistic phenomena

RefreshingDisplay only

KoBALT · updated September 2, 2026

Korean CSAT

College Scholastic Ability Test (수능)

The Korean SAT exam.

Multi-subject examStandardized testHigh school to college level

RefreshingDisplay only

Korean CSAT · updated September 2, 2026

HRM8K

HAE-RAE Math 8K

Korean mathematical reasoning (high-school to Olympiad level).

8,011 instancesMath word problemsOlympiad level

RefreshingDisplay only

HRM8K · updated September 2, 2026

External benchmark mirrors(70 benchmarks)

Mirrored from third-party leaderboards; listed by source, not by capability.

2026

KindBench

KindBench Psychological Safety Benchmark

A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.

16 multi-turn scenarios, 72 criteriaJudge-scored behavioral audit with human reviewAdversarial psychological-safety evaluation

CurrentDisplay only

KindBench v0.1.0 · updated September 2, 2026

2024

LiveBench

LiveBench

A frequently refreshed benchmark with objective scoring across reasoning, coding, agentic coding, mathematics, data analysis, language, and instruction following.

23 objective tasks across 7 categoriesMean of category averagesBroad frontier-model evaluation

RefreshingDisplay only

LiveBench 2024 · updated September 2, 2026

2026

Vals Index

Vals Index v2

Vals AI composite benchmark across professional finance, coding, modeling, and legal-work tasks, including Finance Agent v2, EMB, Terminal-Bench 2.1, Vibe Code Bench, Code Migration, Legal Research, and HLAB.

Finance, coding, spreadsheet modeling, code migration, and legal-work componentsComposite scorePrivate economic-work benchmark composite

CurrentDisplay only

Vals Index 2026 · updated September 2, 2026

2026

Web Search Index

Vals Web Search Index

A Vals AI comparison of native provider search and Exa across finance-analysis and legal-research tasks.

Finance Agent Benchmark v2 and Legal Research Benchmark tasksAccuracy by model and search-tool combinationProfessional web research with controlled search-tool variants

CurrentDisplay only

Web Search Index 2026 · updated September 2, 2026

2026

Time Horizon Index: KSP

Vals Time Horizon Index: Kerbal Space Program

A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.

30 progressively harder Kerbal Space Program missionsMission-ladder progress with partial creditLong-horizon autonomous computer use and planning

CurrentDisplay only

Time Horizon Index: KSP 2026 · updated September 2, 2026

2026

Vals Multimodal Index

Vals Multimodal Index v1.2

Vals AI multimodal composite across finance, coding, education, and mortgage-tax task families.

Finance, coding, education, and mortgage-tax componentsComposite scorePrivate multimodal economic-work benchmark composite

CurrentDisplay only

Vals Multimodal Index 2026 · updated September 2, 2026

2026

Finance Agent v1.1

Vals Finance Agent v1.1

An archived Vals benchmark covering retrieval, numerical reasoning, financial modeling, market analysis, earnings, trends, and adjustments.

11 financial analyst task viewsAccuracy with task-level breakdownsProfessional financial analysis

CurrentDisplay only

Finance Agent v1.1 2026 · updated September 2, 2026

2026

ReverseEngBench

Vals ReverseEngBench

A contamination-resistant agent benchmark for reverse engineering real-world binaries.

Real-world binary reverse-engineering tasksFully solved rate and capability scoreAgentic reverse engineering

CurrentDisplay only

ReverseEngBench 2026 · updated September 2, 2026

2026

Vals Terminal-Bench 1.0 mirror

Vals-hosted Terminal-Bench 1.0 mirror

A Vals-hosted view of Terminal-Bench 1.0 with easy, medium, and hard task splits.

Terminal tasks split by easy, medium, and hard difficultyAccuracy scoreTerminal-agent execution

CurrentDisplay only

Vals Terminal-Bench 1.0 mirror 2026 · updated September 2, 2026

2026

CorpFin v2

Vals CorpFin v2

Vals AI private benchmark for understanding long-context credit agreements.

Credit-agreement understanding tasksAccuracy scoreProfessional finance document reasoning

CurrentDisplay only

CorpFin v2 2026 · updated September 2, 2026

2026

MedCode

Vals MedCode

Vals AI healthcare benchmark for whether models can support the medical billing process.

Medical billing support tasksAccuracy scoreProfessional healthcare administration

CurrentDisplay only

MedCode 2026 · updated September 2, 2026

2026

MedScribe

Vals MedScribe

Vals AI healthcare benchmark for whether models can support doctors with administrative work.

Medical administrative support tasksAccuracy scoreProfessional healthcare administration

CurrentDisplay only

MedScribe 2026 · updated September 2, 2026

2026

MortgageTax

Vals MortgageTax

Vals AI benchmark for mortgage and tax document reasoning, including semantic and numerical extraction task views.

Mortgage and tax extraction tasksAccuracy scoreProfessional mortgage-tax document reasoning

CurrentDisplay only

MortgageTax 2026 · updated September 2, 2026

2026

ProofBench

Vals ProofBench

Vals AI automated theorem-proving benchmark.

Automated theorem provingAccuracy scoreFormal proof reasoning

CurrentDisplay only

ProofBench 2026 · updated September 2, 2026

2026

LegalBench

Vals LegalBench

Vals AI legal benchmark with issue, rule, conclusion, interpretation, and rhetoric task views.

Legal reasoning task viewsAccuracy scoreProfessional legal reasoning

CurrentDisplay only

LegalBench 2026 · updated September 2, 2026

2026

CaseLaw v2

Vals CaseLaw v2

Vals AI private question-answer benchmark over Canadian court cases.

Canadian case-law question answeringAccuracy scoreProfessional legal retrieval and reasoning

CurrentDisplay only

CaseLaw v2 2026 · updated September 2, 2026

2026

DeepSWE

DeepSWE

A long-horizon software engineering benchmark from Datacurve for measuring frontier coding agents on original tasks drawn from active open-source repositories.

113 software engineering tasks across 91 repositories and 5 languagesPass@1 with confidence interval, cost, time, and token metadataLong-horizon software engineering

CurrentDisplay only

DeepSWE 2026 · updated September 2, 2026

2026

SWE-Marathon

SWE-Marathon

A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.

20 multi-hour software engineering tasksTask resolution and trajectory reviewUltra-long-horizon software engineering

CurrentDisplay only

SWE-Marathon 2026 · updated September 2, 2026

2026

ExploitBench

ExploitBench v8-bench

A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

V8 exploit synthesis runsCapability coverage percentage over 16 flagsBrowser exploitation and cybersecurity

CurrentDisplay only

ExploitBench 2026 · updated September 2, 2026

2026

ACE solved

ACE Cyber Range Challenges Solved

Number of advanced cyber-range challenges solved in the joint NIST CAISI and UK AISI preliminary evaluation.

41 advanced cyber-range challengesChallenges solvedAdvanced cyber operations

CurrentDisplay only

ACE solved 2026 · updated September 2, 2026

2026

The Last Ones steps

The Last Ones Average Progress

Average step reached on a 32-step long-horizon cyber range.

32-step long-horizon cyber rangeAverage step reachedLong-horizon cyber operations

CurrentDisplay only

The Last Ones steps 2026 · updated September 2, 2026

2026

The Last Ones completion

The Last Ones Cyber Range Completion Rate

Share of runs that completed the 32-step cyber range within the 100-million-token limit.

10 long-horizon runsCompletion rateLong-horizon cyber operations

CurrentDisplay only

The Last Ones completion 2026 · updated September 2, 2026

2026

ACCR standard

Advanced Cyber Completion Rate — Standard Access

Share of approved advanced-cyber requests completed rather than refused under standard GPT-5.6 Sol safeguards.

Internal advanced-cyber request setCompletion rateCyber access and safeguards

CurrentDisplay only

ACCR standard 2026 · updated September 2, 2026

2026

ACCR Daybreak Blue

Advanced Cyber Completion Rate — Daybreak Blue

Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Sol under Daybreak Blue safeguards.

Internal advanced-cyber request setCompletion rateCyber access and safeguards

CurrentDisplay only

ACCR Daybreak Blue 2026 · updated September 2, 2026

2026

ACCR Daybreak Red

Advanced Cyber Completion Rate — Daybreak Red

Share of approved advanced-cyber requests completed rather than refused by GPT-5.6 Cyber under Daybreak Red access.

Internal advanced-cyber request setCompletion rateCyber access and safeguards

CurrentDisplay only

ACCR Daybreak Red 2026 · updated September 2, 2026

2026

SEC-Bench Pro

SEC-Bench Pro

Cybersecurity benchmark for agentic vulnerability analysis and exploit-oriented security tasks.

Security engineering tasksSuccess rateAdvanced cybersecurity

CurrentDisplay only

SEC-Bench Pro 2026 · updated September 2, 2026

2026

FrontierCyber

FrontierCyber

Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.

197 dynamic cyber tasksTasks solvedEasy through elite cyber operations

CurrentDisplay only

FrontierCyber 2026 · updated September 2, 2026

2026

CyScenarioBench success

CyScenarioBench Average Success Rate

Average success rate across realistic, long-horizon cybersecurity scenarios.

11 cyber scenariosAverage success rateLong-horizon cybersecurity

CurrentDisplay only

CyScenarioBench success 2026 · updated September 2, 2026

2026

CyScenarioBench solved

CyScenarioBench Scenarios Ever Solved

Number of CyScenarioBench scenarios completed successfully in at least one run.

11 cyber scenariosScenarios solved at least onceLong-horizon cybersecurity

CurrentDisplay only

CyScenarioBench solved 2026 · updated September 2, 2026

2026

Atomic network attacks

Atomic Network Attack Simulation

Irregular's domain-level evaluation of network attack simulation capability.

Atomic cyber tasksDomain averageNetwork attack simulation

CurrentDisplay only

Atomic network attacks 2026 · updated September 2, 2026

2026

Atomic vulnerability research

Atomic Vulnerability Research and Exploitation

Irregular's domain-level evaluation of vulnerability research and exploitation capability.

Atomic cyber tasksDomain averageVulnerability research and exploitation

CurrentDisplay only

Atomic vulnerability research 2026 · updated September 2, 2026

2026

Atomic evasion

Atomic Evasion

Irregular's domain-level evaluation of cybersecurity evasion capability.

Atomic cyber tasksDomain averageCybersecurity evasion

CurrentDisplay only

Atomic evasion 2026 · updated September 2, 2026

2026

SCONE post-cutoff success

SCONE Post-Cutoff Exploit Success

Share of SCONE smart-contract vulnerabilities exploited on the 12-task post-cutoff set.

12 post-cutoff smart-contract vulnerabilitiesBest@8 exploit success rateSmart-contract exploitation

CurrentDisplay only

SCONE post-cutoff success 2026 · updated September 2, 2026

2026

SCONE simulated revenue

SCONE Post-Cutoff Simulated Exploit Revenue

Simulated value captured across the SCONE post-cutoff smart-contract set.

12 post-cutoff smart-contract vulnerabilitiesSimulated USD millions, Best@8Smart-contract exploitation

CurrentDisplay only

SCONE simulated revenue 2026 · updated September 2, 2026

2026

Firefox 147 exploits

Firefox 147 Working Exploit Rate

Share of patched Firefox 147 JavaScript-engine targets for which the model produced a working arbitrary-code-execution exploit.

250 Firefox 147 vulnerability trialsWorking arbitrary-code-execution rateBrowser exploit development

CurrentDisplay only

Firefox 147 exploits 2026 · updated September 2, 2026

2026

Anthropic OSS-Fuzz crash

Anthropic OSS-Fuzz Any-Crash Rate

Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.

Approximately 830 OSS-Fuzz entry pointsAny-crash rateVulnerability discovery

CurrentDisplay only

Anthropic OSS-Fuzz crash 2026 · updated September 2, 2026

2026

Anthropic OSS-Fuzz write primitive

Anthropic OSS-Fuzz Write-Primitive-or-Higher Rate

Share of evaluated OSS-Fuzz entry points where the model achieved a write primitive or stronger result.

Approximately 830 OSS-Fuzz entry pointsWrite-primitive-or-higher rateVulnerability exploitation

CurrentDisplay only

Anthropic OSS-Fuzz write primitive 2026 · updated September 2, 2026

2026

CVE-Bench zero-day

CVE-Bench v1 Zero-Day Black-Box Evaluation

OpenAI's black-box, no-source variant of CVE-Bench v1 across 40 critical vulnerabilities.

40 critical CVEsPass@1 over three rolloutsBlack-box vulnerability exploitation

CurrentDisplay only

CVE-Bench zero-day 2026 · updated September 2, 2026

2026

GBA-Eval

GBA-Eval

An agentic coding benchmark that asks models to build a Game Boy Advance emulator from scratch and grades emulator behavior against procedural, audio, and gameplay tests.

27 emulator test casesOverall emulator scoreLong-horizon systems programming

CurrentDisplay only

GBA-Eval 2026 · updated September 2, 2026

2025

CAIS Text Leaderboard

CAIS AI Dashboard Text Capabilities Index

A Center for AI Safety dashboard view summarizing text capabilities across HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuests.

HLE, ARC-AGI-2, SWE-Bench Pro, and TextQuestsAverage component scoreComposite frontier text capability

CurrentDisplay only

CAIS Text Leaderboard 2025 · updated September 2, 2026

2025

VoxelBench Text

VoxelBench Text-Prompt Leaderboard

A live human-preference benchmark where language models turn text prompts into voxel structures and voters compare anonymous builds from the same prompt.

Live text prompts for 3D voxel constructionGlicko-2 rating from blind pairwise votes3D spatial construction and visual quality

CurrentDisplay only

Live VoxelBench Glicko-2 · updated September 2, 2026

2025

VoxelBench Image

VoxelBench Image-Prompt Leaderboard

A live human-preference benchmark where multimodal models build voxel structures from image references and voters compare anonymous results produced from the same prompt.

Live image-reference prompts for 3D voxel constructionGlicko-2 rating from blind pairwise votesVisual grounding, 3D construction, and aesthetic quality

CurrentDisplay only

Live VoxelBench Glicko-2 · updated September 2, 2026

2026

WeirdML

WeirdML v2

A machine-learning engineering benchmark that tests whether LLMs can train models on novel datasets, write PyTorch code, and improve through iterative feedback.

17 novel ML engineering tasksAverage accuracy across tasksNovel dataset modeling and iterative debugging

CurrentDisplay only

WeirdML 2026 · updated September 2, 2026

2026

ALE-Bench

Agents Last Exam

A benchmark for agentic professional workflows with verifiable success criteria, reporting pass rates and partial scores for model plus agent-harness rows.

152 ALE-V1 professional workflow tasks across 13 top-level domainsPass rate, partial-credit score, cost, token, and duration metadataReal-world agentic workflows

CurrentDisplay only

ALE-Bench 2026 · updated September 2, 2026

2026

RuneScape-Bench

RuneBench / runescape-bench

An agentic coding benchmark where models use a TypeScript SDK to play a RuneScape-like environment and optimize skill-training performance.

16 RuneScape skill-training tasksAverage log XP-rate scoreAgentic gameplay automation

CurrentDisplay only

RuneScape-Bench 2026 · updated September 2, 2026

2026

Toloka Arena

Toloka Arena

An independent agentic-intelligence evaluation from Toloka using private simulated workflows and a pass^5 metric.

Private simulated enterprise workflowspass^5 arena scoreAgentic workflow reliability

CurrentDisplay only

Toloka Arena 2026 · updated September 2, 2026

2026

Vals SWE-bench mirror

Vals-hosted SWE-bench mirror

Vals AI hosted SWE-bench view for solving production software engineering tasks.

Software engineering issue-resolution tasksAccuracy scoreProduction software engineering

CurrentDisplay only

Vals SWE-bench mirror 2026 · updated September 2, 2026

2026

Vals Terminal-Bench 2.0 mirror

Vals-hosted Terminal-Bench 2.0 mirror

Vals AI hosted Terminal-Bench 2.0 view with easy, medium, and hard task splits.

Terminal task difficulty splitsAccuracy scoreTerminal-based agent execution

CurrentDisplay only

Vals Terminal-Bench 2.0 mirror 2026 · updated September 2, 2026

2026

Vals LiveCodeBench mirror

Vals-hosted LiveCodeBench mirror

Vals AI implementation of LiveCodeBench with easy, medium, and hard task splits.

Coding problem difficulty splitsAccuracy scoreContamination-resistant coding problems

CurrentDisplay only

Vals LiveCodeBench mirror 2026 · updated September 2, 2026

2026

Vals GPQA Diamond mirror

Vals-hosted GPQA Diamond mirror

Vals AI hosted GPQA Diamond view with few-shot and zero-shot chain-of-thought task splits.

GPQA Diamond task splitsAccuracy scoreGraduate science reasoning

CurrentDisplay only

Vals GPQA Diamond mirror 2026 · updated September 2, 2026

2026

Vals MMLU-Pro mirror

Vals-hosted MMLU-Pro mirror

Vals AI hosted MMLU-Pro view with subject-level task splits.

MMLU-Pro subject splitsAccuracy scoreProfessional academic reasoning

CurrentDisplay only

Vals MMLU-Pro mirror 2026 · updated September 2, 2026

2026

EMB

Vals EMB

Evaluating agents on Excel-based financial modeling tasks

Excel-based financial modeling tasksAccuracy scoreProfessional finance modeling

CurrentDisplay only

EMB 2026 · updated September 2, 2026

2026

CyberBench

Vals CyberBench

Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and stop crashing after the fix?

OSS-Fuzz PoC and patch-verification tasksAccuracy scoreAutonomous cybersecurity exploit reproduction

CurrentDisplay only

CyberBench 2026 · updated September 2, 2026

2026

TaxEval v2

Vals TaxEval v2

A Vals-created set of questions and responses to tax questions

Tax question answering and response evaluationAccuracy scoreProfessional tax reasoning

CurrentDisplay only

TaxEval v2 2026 · updated September 2, 2026

2026

Harvey's Legal Agent Benchmark

Vals Harvey's Legal Agent Benchmark

Tests an agent's ability to complete legal work using documents, spreadsheets, presentations, and file-system tools

Legal agent work across documents, spreadsheets, presentations, and filesAccuracy scoreProfessional legal workflow automation

CurrentDisplay only

Harvey's Legal Agent Benchmark 2026 · updated September 2, 2026

2026

Terminal-Bench 2.1

Vals Terminal-Bench 2.1

State-of-the-art set of difficult terminal-based tasks

Terminal-based task executionAccuracy scoreFrontier terminal-agent execution

CurrentDisplay only

Terminal-Bench 2.1 2026 · updated September 2, 2026

2026

Code Migration

Vals Code Migration

Can language models reimplement real-world programs in another language?

Real-world program reimplementation in another languageAccuracy scoreProduction code migration

CurrentDisplay only

Code Migration 2026 · updated September 2, 2026

2026

Legal Research Bench

Vals Legal Research Bench

Evaluating agents on legal research tasks across diverse areas of US law

US-law legal research tasksAccuracy scoreProfessional legal research

CurrentDisplay only

Legal Research Bench 2026 · updated September 2, 2026

2026

MedQA

Vals MedQA

Evaluating language model bias in medical questions.

Medical question answeringAccuracy scoreMedical knowledge and bias evaluation

CurrentDisplay only

MedQA 2026 · updated September 2, 2026

2026

AIME

Vals AIME

Challenging national math exam given to top high-school students

AIME math problemsAccuracy scoreCompetition math

CurrentDisplay only

AIME 2026 · updated September 2, 2026

2026

MATH 500

Vals MATH 500

Academic math benchmark on probability, algebra, and trigonometry

MATH 500 academic math problemsAccuracy scoreAdvanced academic math

CurrentDisplay only

MATH 500 2026 · updated September 2, 2026

2026

MGSM

Vals MGSM

A multilingual benchmark for mathematical questions.

Multilingual grade-school math questionsAccuracy scoreMultilingual mathematical reasoning

CurrentDisplay only

MGSM 2026 · updated September 2, 2026

2026

MMMU

Vals MMMU

Multimodal Multi-task Benchmark

Multimodal academic task suiteAccuracy scoreMultimodal college-level reasoning

CurrentDisplay only

MMMU 2026 · updated September 2, 2026

2026

SAGE

Vals SAGE

Student Assessment with Generative Evaluation

Student assessment with generative evaluationAccuracy scoreEducation assessment reasoning

CurrentDisplay only

SAGE 2026 · updated September 2, 2026

2026

IOI

Vals IOI

Based on the International Olympiad in Informatics

International Olympiad in Informatics-style programming tasksAccuracy scoreOlympiad programming

CurrentDisplay only

IOI 2026 · updated September 2, 2026

2026

ProgramBench

Vals ProgramBench

Can language models rebuild programs from scratch?

Program reconstruction tasksAccuracy scoreCleanroom software engineering

CurrentDisplay only

ProgramBench 2026 · updated September 2, 2026

2026

SkillsBench

Vals SkillsBench

How important are skills for agents?

Agent skill-importance tasksAccuracy scoreAgent skill evaluation

CurrentDisplay only

SkillsBench 2026 · updated September 2, 2026

2026

Agent Poker Bench

Vals Agent Poker Bench

Which model can make the most money playing poker?

Poker-playing agent trialsAccuracy scoreStrategic game-agent decision making

CurrentDisplay only

Agent Poker Bench 2026 · updated September 2, 2026

2026

Public Benefits Bench v1.1

Vals Public Benefits Bench v1.1

Can AI help people navigate SNAP benefits?

SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation

CurrentDisplay only

Public Benefits Bench v1.1 2026 · updated September 2, 2026

2026

Public Benefits Bench v1

Vals Public Benefits Bench v1

Can AI help people navigate SNAP benefits?

SNAP public-benefits navigation tasksAccuracy scorePublic-benefits policy navigation

CurrentDisplay only

Public Benefits Bench v1 2026 · updated September 2, 2026