Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
This repo is an open reproduction of TypeSafe AI's Jev model, providing a 150M-parameter non-autoregressive typed decision engine. It scores 0.697 accuracy on the typed-decisions benchmark (Jev: 0.727), with better calibration, and can be trained for free on Colab.
该仓库明确复现 TypeSafe AI 的 Jev 模型,包含完整代码、训练与评估流程,并在 typed-decisions 基准上与 Jev 对比,属于深入研究/复现项目。