Back to explore

Jev đạt kết quả tốt trên benchmark ScopeJudge

Dreadnode cho biết Jev của TypeSafe cạnh tranh với các giám khảo LLM hàng đầu trên benchmark ScopeJudge, phát hiện vi phạm phạm vi của tác nhân với chi phí vài xu mỗi nghìn lần kiểm tra, phản hồi trung bình 130 ms.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev đạt kết quả tốt trên benchmark ScopeJudge 1
· 26 likesOpen on X