Back to explore

Jev presteert goed op ScopeJudge-benchmark

Dreadnode meldt dat Jev van TypeSafe concurreert met toonaangevende LLM-rechters op de ScopeJudge-benchmark, scope-overtredingen van agenten detecteert voor enkele centen per duizend checks, met gemiddeld 130 ms reactietijd.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev presteert goed op ScopeJudge-benchmark 1
· 26 likesOpen on X