TypeSafe AIが引用チェッカーに追加
ユーザーが引用チェッカーに@typesafeaiを追加し高評価。PaperTrellisで無料試用可能。Jevと医療AIのタグ付き。 Link
Benchmark Heaven によると、v1.4.1 の Jev-Omni はハード層で 75.0%、Jev 1.13.0 は 74.1%。ただし、シールドセット(32.1% vs 36.7%)、キャリブレーション(64.1 vs 76.3)、総合スコア(51.34 で7位 vs 63.29 で1位)では劣る。回答は近いが、信頼度に差がある。
Jev-Omni, new in v1.4.1, beats Jev 1.13.0 on the hard tier: 75.0% vs 74.1%. Harold checked twice. The rest runs the other way: sealed set 32.1% vs 36.7%, Calibration 64.1 vs 76.3, score 51.34 (#7) vs 63.29 (#1). Close on answers, not on confidence. https://benchmarkheaven.com/jev-models/v1.4.1?compare=jev-1.13.0,jev-omni…