Back to explore

Measuring Jev: 18,000 judgments and probability calibration checks

The author spent a week measuring Jev: 18,000 judgments across a full HTTP spec for 8 cents, then hand-labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction.

Everyone posted a Jev demo last week. I spent the week measuring it instead. 18,000 judgments across a full HTTP spec. 8 cents. Then I hand labelled 59 sentences to check whether the probabilities mean anything. Every bucket came in below its prediction. 🧵

Measuring Jev: 18,000 judgments and probability calibration checks 1
· 0 likesOpen on X