Back to explore

Jev, ScopeJudge 벤치마크에서 우수한 성과

Dreadnode는 TypeSafe의 Jev가 ScopeJudge 벤치마크에서 주요 LLM 심사 모델과 경쟁하며 에이전트 범위 위반을 탐지하고 평균 130ms 응답, 1000회당 몇 센트 비용이라고 밝혔다.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev, ScopeJudge 벤치마크에서 우수한 성과 1
· 26 likesOpen on X