Jev 在邮件分类基准测试中不敌 Gemini,但仍考虑投产
作者在 1,565 封德语和英语工业供应商商务邮件的 10 类分类基准上测试了 Jev,发现其准确率不及 Gemini,但更关注错误出现的模式,并仍有意将其投入生产。
文章从 Bridge360 元理论模型出发,分析 Jev/System One 模型与 Palantir 相似的局限性,并关联 AI 治理与 AI 安全议题。
Jev/System One model comes short where Palantir comes short: Applying Bridge360 Metatheory Model lens #AIGovernance #AISafety https://agericomontecillodevilla.substack.com/p/jevsystem-one-model-comes-short-where…