TypeSafe AIが引用チェッカーに追加
ユーザーが引用チェッカーに@typesafeaiを追加し高評価。PaperTrellisで無料試用可能。Jevと医療AIのタグ付き。 Link
同じ430ケースのテストで、証拠を示すとJevは68〜92%、Layaは84〜88%の虚偽エージェント報告を見逃した。一方、直接的な質問に変えるとJevは約1%しか誤りを見逃さなかった。モデルよりも質問設計の影響が大きいことを示す。
Same 430-case test, two AI judges. Shown the evidence, Jev passed 68 to 92% of the false agent reports and Laya 84 to 88%. Asked one literal question instead, Jev let through about 1% of wrong counts and files. The question matters more than the model. https://cejel.dev/experiments