Weekly LLM benchmark digest
Jev cannot write. We gave it a job.
Our LLM Selector now uses a model that cannot write a reply. We gave it nineteen questions instead. TypeSafe released Jev on September 15. It reads text and returns choices, scores or yes-or-no probabilities, within the answer space you define. That makes it useful for the small decisions inside an AI workflow, such as routing a ticket or checking a stated criterion. It does not replace a model that has to write the reply.
Read the Jev breakdown →Why the model is getting attention
The limit
A typed answer can still be wrong. We have not measured Jev’s calibration or speed on our selector traffic. A confidence value does not establish the accuracy of a job classification, and consequential actions still need permissions, confirmation and limits in code.
The boundary that matters
The option list is part of the contract.
Jev cannot select an option outside the list you provide. That bounds the output, but a badly defined list can still produce a bad decision.
Writing remains a separate job.
Jev cannot compose a summary, reply, patch or explanation. Pair it with a language model when a workflow needs both a bounded decision and generated text.
Analysis worth opening
This issue in numbers
- 19
- Fixed questions in the selector’s standard request
- 14 + 5
- Choice and yes-or-no readings in that request
- 31 + unknown
- Catalog jobs available to the primary job choice
- $0.042 / M
- Jev 1.13 input list price on September 22; output listed at $0