Back to explore

Jev добре показує себе в бенчмарку ScopeJudge

Dreadnode повідомляє, що Jev від TypeSafe конкурує з провідними LLM-суддями в бенчмарку ScopeJudge, виявляючи порушення області агентів за кілька центів на тисячу перевірок із середнім відгуком 130 мс.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev добре показує себе в бенчмарку ScopeJudge 1
· 26 likesOpen on X