Jev, 분류기 평가에서 품질과 속도 모두 우세
Malte Ubl이 TypeSafe AI의 Jev를 기존 분류기 평가에 적용한 결과, 품질에서 평가를 포화시켰고 Gemini 2.5 Flash Lite보다 6배 빨랐다.
저자는 일주일 동안 Jev를 측정했습니다. 전체 HTTP 명세에 대해 1만 8천 건 판단을 8센트로 수행하고, 확률의 의미를 확인하기 위해 59개 문장을 수동 라벨링했습니다. 모든 구간이 예측보다 낮았습니다.
Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵