返回探索

对 Jev 模型进行 1.8 万次判断测量与概率校准验证

作者用一周时间对 Jev 模型进行测量:在完整 HTTP 规范上完成 1.8 万次判断,成本 8 美分;并手工标注 59 个句子检验概率含义,发现每个分桶的实际结果都低于预测值。

Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵

对 Jev 模型进行 1.8 万次判断测量与概率校准验证 1
· 0 次赞在 X 打开