Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
AnyJev is an open research library that turns any open LLM into a Jev-style decision model, offering typed decisions, calibrated probabilities, training-free L0/L1 calibration, and benchmark comparisons against Jev.
该仓库是一个实质性的研究项目,实现并评测了Jev风格的决策模型(AnyJev),将任意LLM转换为带类型化决策和可靠概率的模型,并与TypeSafe AI的Jev结果进行对比,属于深入研究Jev/System One思路的技术资源。