# About BenchLM

> BenchLM tracks 411 AI models across 111 benchmarks in eight categories: agentic, coding, multimodal & grounded, reasoning, knowledge, instruction following, multilingual, and math.

## Why BenchLM Exists

Comparing AI models is harder than it should be. Benchmark rows are scattered across launch posts, model cards, benchmark leaderboards, papers, and screenshots. BenchLM pulls them into one place and tries to make the final score understandable instead of magical.

## Where the Data Comes From

- OpenBench evaluations and benchmark-native public leaderboards
- Official provider tables, model cards, and release announcements
- Cross-referenced public sources when multiple reports exist

If sources conflict, BenchLM prefers the more exact or more controlled evaluation environment.

## How Scoring Works

BenchLM starts from a benchmark backbone. Within each category, weighted benchmark rows are normalized and blended into a category score. The overall score then combines those category scores using fixed category weights. Display-only benchmarks remain visible for context, but do not directly affect the weighted ranking.

| Category | Weight |
|----------|--------|
| Agentic | 22% |
| Coding | 20% |
| Reasoning | 17% |
| Multimodal & Grounded | 12% |
| Knowledge | 12% |
| Multilingual | 7% |
| Instruction Following | 5% |
| Math | 5% |

BenchLM also applies bounded external consensus calibration to some public display scores. That calibration can nudge benchmark-only output when the benchmark backbone clearly misranks frontier rows, but it does not replace the benchmark backbone.

Agentic gets the highest weight because the frontier has shifted from answering questions to completing workflows. Coding and reasoning still matter heavily, while harder, less saturated benchmarks are favored over legacy floor-check rows.

## Verified vs Provisional

The site separates verified and provisional views. Both exclude unresolved manual values and generated benchmark rows from scoring. The provisional view can also admit a benchmark-sparse model when several independent public evaluation families agree; that evidence does not upgrade the model to verified status.

Every score also carries a confidence signal based on how much source-backed benchmark coverage supports it.

## Update Frequency

BenchLM updates when new benchmark rows appear, when evaluation protocols materially change, and when important model releases justify a refresh. The current dataset was last updated on **September 4, 2026**.

## What BenchLM Does Not Track

- BenchLM shows headline API pricing, but not provider-specific contract or region-level pricing
- BenchLM does not run proprietary private evaluations
- BenchLM is best used for public benchmark performance, relative tradeoffs, and shortlist construction

## Who Runs This

BenchLM is built and maintained by [@glevd](https://x.com/glevd). If you spot a data error or want a model or benchmark added, reach out there.

Canonical page: https://benchlm.ai/about
