Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
An empirical experiment and API research repository for TypeSafe's Jev / System One model, running controlled tests to evaluate model behavior and saving raw evidence.
仓库明确针对 TypeSafe Jev / System One,包含 Jev API 客户端、实验脚本、测试与原始结果,是真实的 API 集成和实证评估研究。