Jev 在分类器评测中质量与速度均胜出
Malte Ubl 将 TypeSafe AI 的 Jev 用于现有分类器评测,结果在质量上饱和评测集,速度比 Gemini 2.5 Flash Lite 快 6 倍。
作者用一周时间对 Jev 模型进行测量:在完整 HTTP 规范上完成 1.8 万次判断,成本 8 美分;并手工标注 59 个句子检验概率含义,发现每个分桶的实际结果都低于预测值。
Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵