Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
Evaluation of Jev and nine open checkpoints across 22 tasks shows Jev leads on macro accuracy, but an open 4B model produces better-calibrated probabilities; jev-bench is fully open.
Jev Tests: Accuracy gave one ranking. Calibration gave another. We evaluated Jev and nine open checkpoints across 22 tasks. Jev led on macro accuracy. But an open 4B model produced better-calibrated probabilities. jev-bench is fully open 👇