Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
A research repo that wraps TypeSafe's Jev/System One in a TLA+-verified consensus kernel, with noise-floor measurements, calibration probes, and 1,680 chaos-tested pharmacy decisions via the live Jev API.
仓库以 TypeSafe 的 Jev/System One 为核心,通过形式化验证、混沌测试和实际 API 调用深入测试共识内核,包含真实的 Jev API 调用、SDK 使用和评测数据。