TypeSafe AI 被用于学术引用检查工具
用户将 @typesafeai 加入其引用检查器并给予好评,可在 PaperTrellis 免费试用,标签涉及 Jev 与医学 AI。 链接
同一430例测试中,展示证据时Jev漏过68-92%的虚假代理报告,Laya漏过84-88%;改为直接提问后,Jev仅漏过约1%的错误计数和文件。实验结果表明提问方式的影响大于模型选择。
Same 430-case test, two AI judges. Shown the evidence, Jev passed 68 to 92% of the false agent reports and Laya 84 to 88%. Asked one literal question instead, Jev let through about 1% of wrong counts and files. The question matters more than the model. https://cejel.dev/experiments