Back to explore

Jev يحقق أداءً قوياً في معيار ScopeJudge

ذكرت Dreadnode أن Jev من TypeSafe ينافس نماذج تحكيم LLM الرائدة في معيار ScopeJudge، ويكتشف انتهاكات نطاق الوكلاء بتكلفة سنتات لكل ألف فحص مع استجابة متوسطة 130 مللي ثانية.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev يحقق أداءً قوياً في معيار ScopeJudge 1
· 26 likesOpen on X