Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
Connects TypeSafe Jev to Pokémon Showdown, recording 40 battles and 940 API decisions with reproducible results, datasets, and annotated replays. Independent research, not an official benchmark.
该仓库真实连接并调用 TypeSafe Jev API,在 Pokémon Showdown 中进行了 40 场对战、940 次 API 决策实验,提供完整代码、数据、结果与回放,属于实质性的集成与评估项目。