# How we score models

The methodology combines benchmark results, evidence quality, freshness, pricing, and runtime metadata. Missing tests are not treated as zeroes; sparse models receive wider uncertainty and an Estimated label until stronger evidence arrives.

## Ranking rules

Public rankings use the same evidence-gated scores shown on model and provider pages. Unrounded scores determine order. Display-only or unsupported evidence does not silently affect the ranking.

## Freshness and provenance

The benchmark catalog was last refreshed September 27, 2026. Each benchmark page records its source and verification state.

Last updated: September 27, 2026

Canonical page: https://benchlm.ai/methodology
