Back to explore

Jev dobrze wypada w benchmarku ScopeJudge

Dreadnode informuje, że Jev od TypeSafe konkuruje z wiodącymi sędziami LLM w benchmarku ScopeJudge, wykrywając naruszenia zakresu agentów za kilka centów na tysiąc sprawdzeń, ze średnim czasem odpowiedzi 130 ms.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev dobrze wypada w benchmarku ScopeJudge 1
· 26 likesOpen on X