Skip to main content

BenchLM Research

LLM benchmark research & analysis

Evidence for choosing and evaluating AI models, grounded in benchmark explainers, model comparisons, and the BenchLM dataset.

68 published articles RSS feed
Two separate benchmark lanes labelled Terminal-Bench 2.0 and 2.1, with the 56.9 and 82.7 scores sitting in different lanes and the 61.8 like-for-like comparison marked in the 2.1 lane

Featured research

DeepSeek V4 Flash Gained 20.9 Points, Not 25.8

DeepSeek-V4-Flash-0731 shipped July 31 with MIT weights and $0.28 output pricing. The Terminal-Bench jump being quoted compares two benchmark versions, and DeepSeek's own table puts it behind Opus 4.8 on all nine rows.

Glevd · August 1, 2026 · 7 min read

Read the article

Latest research

A chronological record of benchmark analysis, model decisions, and evaluation practice.

Showing 67 of 68 articles

Get the BenchLM weekly brief

New analysis, material benchmark and price changes, and one model decision worth revisiting.

Read a sample issue

Join 2,000+ readers.