Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
Six recorded chess experiments that call TypeSafe Jev through OpenRouter: independent play, a tactical filter, Stockfish assistance, and a Jev review loop, with full traces, replays, and cost records.
仓库 README 明确说明通过 OpenRouter 调用 TypeSafe Jev 进行六次国际象棋对局实验,包含完整请求/响应、回放和成本记录,并附有运行脚本与测试,属于实质性的 Jev 实验项目。