Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Research deskBenchLM Research

LLM benchmark research & analysis

Evidence for choosing and evaluating AI models, grounded in benchmark explainers, model comparisons, and the BenchLM dataset.

88 published articles RSS feed
Two ink checkmarks on ivory ledger paper, the second crossed by a single diagonal cobalt strike

Featured research

Valid JSON can still be wrong

Structured output guarantees the shape of a response, not the truth in it. Four synthetic pairs from one support-ticket fixture show a schema pass hiding a wrong field, then two filed runs of Claude Sonnet 5 and GPT-5.6 Terra show the same failure class with denominators.

Glevd · September 14, 2026 · 13 min read

Read the article

Latest research

A chronological record of benchmark analysis, model decisions, and evaluation practice.

Showing 87 of 88 articles, plus the featured article above

Get the BenchLM weekly brief

New analysis, material benchmark and price changes, and one model decision worth revisiting.

Read a sample issue

Join 2,000+ readers.