Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
The author spent a week measuring Jev: 18,000 judgments across a full HTTP spec for 8 cents, then hand-labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction.
Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵