Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
A third party tested accuracy on 68 questions: FLock's THIS/THAT Model 1.0 scored 94.1% and hosted System One service Jev scored 76.5%; the post notes encouraging results for a specialized decision model, while complex tasks like graph traversal and multi-step arithmetic remain limited.
제3자가 정확도 검증을 위해 기록한 68개 질문 테스트 결과: FLock의 THIS / THAT Model 1.0: 94.1% Jev (호스팅형 System One 서비스): 76.5% 특화된 의사결정 모델로서는 고무적인 결과입니다. 다만 그래프 탐색이나 여러 단계의 산술 연산이 필요한 복잡한 작업은 여전히 한계로 남아 있습니다.