Back to explore

Jev destaca en el benchmark ScopeJudge

Dreadnode informa que Jev de TypeSafe compite con los principales jueces LLM en el benchmark ScopeJudge, detectando violaciones de alcance de agentes a centavos por mil comprobaciones con respuestas de 130 ms.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev destaca en el benchmark ScopeJudge 1
· 26 likesOpen on X