Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
A JevBench comparison from Benchmark Heaven shows Jev 1.13.0 scoring 86.6% public and 36.7% sealed, while GPT-6 Luna reaches 99.6% and 95.5%. Luna still ranks #30 on JevBench, with the cost axis ($0.135 vs $0.040 per 1,000 decisions) and gate logic below 50 as key factors.
GPT-6 Luna (medium) gets nearly everything right: 99.6% public, 95.5% sealed. Jev 1.13.0: 86.6% and 36.7%. Luna still ranks #30 on JevBench. One column explains it: $0.135 per 1,000 decisions vs $0.040, Cost axis 36.0. Below 50, the gate applies. https://benchmarkheaven.com/jev-models/v1.4.1?compare=jev-1.13.0,gpt-6-luna…