用 Jev 构建代码库分类器
开发者分享用 Jev 构建代码库分类器,认为这可能解决智能体生成过度工程化代码的问题,并询问下一步测试方向。
Sydney Runkle 分享 Sean 和 Daniel 的指南,说明 Jev 作为评判器为何适合在线评估,包含与其他 LLM 的实验对比以及如何在智能体中试用。
jev as a Judge proves to be a cheaper and more precise alternative to LLM as a judge for online evals. great guide from Sean and Daniel on why Jev is great for evals, an experiment vs other LLMs, and how to try this out for your agents!