用 Jev 构建代码库分类器
开发者分享用 Jev 构建代码库分类器,认为这可能解决智能体生成过度工程化代码的问题,并询问下一步测试方向。
Braintrust 测试了 Jev 作为评判模型的表现:在 groundedness 评判上快速、低成本且具竞争力,但在数学和代码领域落后于具备推理能力的模型,建议用于特定评估任务而非完全替代 LLM-as-a-judge。
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →