返回探索

Braintrust 测试 Jev 作为评判模型的表现

Braintrust 测试了 Jev 作为评判模型的表现:在 groundedness 评判上快速、低成本且具竞争力,但在数学和代码领域落后于具备推理能力的模型,建议用于特定评估任务而非完全替代 LLM-as-a-judge。

We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →

Braintrust 测试 Jev 作为评判模型的表现 1
· 26 次赞在 X 打开