Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
FLock.io shares a third-party-recorded 68-question accuracy test: FLock's THIS/THAT Model 1.0 scores 94.1%, while Jev (hosted System One service) scores 76.5%. The post calls this a promising result for a specialised decision model, noting complex tasks needing graph search or multi-step arithmetic remain challenging.
In a test of 68 questions recorded by a third party for accuracy: - FLock's THIS / THAT Model 1.0: 94.1% - Jev (hosted System One service): 76.5% A promising result for a specialised decision model. Complex tasks requiring graph search or multi-step arithmetic remain