Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
A 151M-parameter decision engine based on ModernBERT that outperforms TypeSafe Jev and Laya on LocalLLaMA/typed-decisions with 77.10% accuracy and 0.0636 Brier score, and includes a WebGPU in-browser demo.
该仓库构建并评测了一个基于 ModernBERT 的非自回归决策引擎,在 LocalLLaMA/typed-decisions 基准上与 TypeSafe Jev 1.13.0 对比,受 Jev 架构启发,并提供 WebGPU 可运行 demo,属于对 TypeSafe Jev 模型的深入研究。