返回探索

Jev 在 ScopeJudge 基准上表现亮眼

Dreadnode 称 TypeSafe 的 Jev 在 ScopeJudge 基准上与领先 LLM 评审模型竞争,能捕捉智能体范围违规,平均响应 130 毫秒,每千次检查成本仅几美分。

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev 在 ScopeJudge 基准上表现亮眼 1
· 26 次赞在 X 打开