Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
An independent evaluation of TypeSafe AI's Jev / System One model, with 50 Chinese-language tests, 143 traceable data points, and reproduction scripts.
该仓库是 Jev 模型的中文独立研究报告,包含 50 条实际 API 调用测试、复现脚本和 143 条可溯源数据表,深度研究了 TypeSafe AI 的 Jev / System One 模型。