Back to explore

JevがScopeJudgeベンチマークで好成績

Dreadnodeによると、TypeSafeのJevはScopeJudgeベンチマークで主要なLLM審査モデルと競合し、エージェントのスコープ違反を検出、平均130ミリ秒、1000回あたり数セントで処理する。

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

JevがScopeJudgeベンチマークで好成績 1
· 26 likesOpen on X