Back to explore

Jev si distingue nel benchmark ScopeJudge

Dreadnode riferisce che Jev di TypeSafe è competitivo con i principali giudici LLM nel benchmark ScopeJudge, rilevando violazioni di ambito degli agenti a pochi centesimi per mille controlli, con risposte medie di 130 ms.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev si distingue nel benchmark ScopeJudge 1
· 26 likesOpen on X