Jevals Data: Independent benchmark for TypeSafe Jev
Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.
Three measured experiments on RAG hallucination: quote-checking, TypeSafe's Jev, and IBM's STAIR. 850+ graded questions, raw responses included.
仓库包含对TypeSafe AI Jev模型在RAG场景中的多项实测评估(幻觉抑制、重排序、路由、重试决策),提供详细报告与可复现数据,属于实质性测试研究。