Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
Code and raw data behind Lunar Terminal: Robocode Tank Royale bots, GapFlap, a token-metering proxy, and a 327-decision randomised trial of TypeSafe Jev.
仓库包含 JevTank、GapFlap、llm-meter 及 327 次随机决策试验,对 TypeSafe Jev 进行实测评估并提供原始日志。