Back to explore

Jev ทำผลงานได้ดีบน ScopeJudge benchmark

Dreadnode รายงานว่า Jev ของ TypeSafe แข่งขันได้กับผู้ตัดสิน LLM ชั้นนำบน ScopeJudge benchmark ตรวจจับการละเมิดขอบเขตของเอเจนต์ด้วยต้นทุนไม่กี่เซนต์ต่อพันครั้ง ใช้เวลาเฉลี่ย 130 มิลลิวินาที

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev ทำผลงานได้ดีบน ScopeJudge benchmark 1
· 26 likesOpen on X