Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
A multimodal JEV-like model fine-tuned from Qwen3.5-0.8B, playing ten browser games from raw pixels in real time, with demo, training code, and evaluation.
该仓库直接引用 TypeSafe Jev/System One,并实现了 JEV-like 多模态游戏模型,提供可运行 demo、训练代码、评估结果,属于实质研究。