Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
Minesweeper experiments comparing TypeSafe Jev (System One), LLMs, hybrid policies, and deterministic solvers, with documented methods and recorded results.
仓库通过调用 TypeSafe Jev(System One)与 LLM、确定性求解器在扫雷任务中进行对比实验,包含可运行基准、文档和记录结果,与 TypeSafe Jev 直接相关。