TypeSafe AI 被用于学术引用检查工具
用户将 @typesafeai 加入其引用检查器并给予好评,可在 PaperTrellis 免费试用,标签涉及 Jev 与医学 AI。 链接
Benchmark Heaven 测试显示,v1.4.1 的 Jev-Omni 在 hard tier 准确率 75.0%,高于 Jev 1.13.0 的 74.1%;但在 sealed set、校准和综合评分上均落后,评分 51.34(第 7)对 63.29(第 1)。答案接近,但置信度差距明显。
Jev-Omni, new in v1.4.1, beats Jev 1.13.0 on the hard tier: 75.0% vs 74.1%. Harold checked twice. The rest runs the other way: sealed set 32.1% vs 36.7%, Calibration 64.1 vs 76.3, score 51.34 (#7) vs 63.29 (#1). Close on answers, not on confidence. https://benchmarkheaven.com/jev-models/v1.4.1?compare=jev-1.13.0,jev-omni…