Jev beats Gemini 2.5 Flash Lite on classifier eval quality and speed
Malte Ubl ran TypeSafe AI's Jev against an existing classifier eval, where it saturated the eval on quality and was 6x faster than Gemini 2.5 Flash Lite.
Nathan LeClaire suggests that Jev's ability to run many evaluations quickly and concurrently allows breaking apart, mutating, and experimenting with inputs while storing numeric results, potentially enabling significance analysis.
Here's food for thought @simonw one thing unique about jev is that you can do so many evaluations so fast (incl. concurrently) that you could break apart, mutate and experiment with inputs and store the numeric results, this probably would allow some 'significance analysis'