Jev 在分类器评测中质量与速度均胜出
Malte Ubl 将 TypeSafe AI 的 Jev 用于现有分类器评测,结果在质量上饱和评测集,速度比 Gemini 2.5 Flash Lite 快 6 倍。
Benchmark Heaven 的 JevBench 对比显示:Jev 1.13.0 在公开/密封测试集得分 86.6% / 36.7%,而 GPT-6 Luna 达到 99.6% / 95.5%。但 Luna 在 JevBench 仍排名第 30,成本轴(每千次决策 0.135 美元 vs 0.040 美元)和门槛逻辑是重要解释因素。
GPT-6 Luna (medium) gets nearly everything right: 99.6% public, 95.5% sealed. Jev 1.13.0: 86.6% and 36.7%. Luna still ranks #30 on JevBench. One column explains it: $0.135 per 1,000 decisions vs $0.040, Cost axis 36.0. Below 50, the gate applies. https://benchmarkheaven.com/jev-models/v1.4.1?compare=jev-1.13.0,gpt-6-luna…