Skip to main content
BenchLM
Weekly brief and archive

Weekly LLM benchmark digest

Jev cannot write. We gave it a job.

Our LLM Selector now uses a model that cannot write a reply. We gave it nineteen questions instead. TypeSafe released Jev on September 15. It reads text and returns choices, scores or yes-or-no probabilities, within the answer space you define. That makes it useful for the small decisions inside an AI workflow, such as routing a ticket or checking a stated criterion. It does not replace a model that has to write the reply.

Read the Jev breakdown →

Why the model is getting attention

The limit

A typed answer can still be wrong. We have not measured Jev’s calibration or speed on our selector traffic. A confidence value does not establish the accuracy of a job classification, and consequential actions still need permissions, confirmation and limits in code.

The boundary that matters

The option list is part of the contract.

Jev cannot select an option outside the list you provide. That bounds the output, but a badly defined list can still produce a bad decision.

Writing remains a separate job.

Jev cannot compose a summary, reply, patch or explanation. Pair it with a language model when a workflow needs both a bounded decision and generated text.

Analysis worth opening

This issue in numbers

19
Fixed questions in the selector’s standard request
14 + 5
Choice and yes-or-no readings in that request
31 + unknown
Catalog jobs available to the primary job choice
$0.042 / M
Jev 1.13 input list price on September 22; output listed at $0

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief