Jev ชนะ Gemini 2.5 Flash Lite ทั้งด้านคุณภาพและความเร็วในการประเมินตัวจำแนก
Malte Ubl ทดลองใช้ Jev ของ TypeSafe AI กับชุดประเมินตัวจำแนกที่มีอยู่ โดยทำคะแนนคุณภาพได้เต็มชุดประเมินและเร็วกว่า Gemini 2.5 Flash Lite 6 เท่า
ผู้เขียนใช้เวลาหนึ่งสัปดาห์วัด Jev: 18,000 การตัดสินบนสเปก HTTP ฉบับเต็มด้วยค่าใช้จ่าย 8 เซนต์ จากนั้นติดป้าย 59 ประโยคด้วยมือเพื่อตรวจสอบว่าความน่าจะเป็นมีความหมายหรือไม่ ทุกช่วงต่ำกว่าที่คาดการณ์ไว้
Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵