Jevでコードベース分類器を構築
開発者がJevでコードベース分類器を構築し、エージェントが生成する過剰設計なコードの解決策になり得ると述べ、次のテスト対象を尋ねている。
BraintrustがJevのジャッジモデルとしての性能を検証。groundedness判定では高速・低コストで競争力がある一方、数学やコード領域では推論モデルに劣る。特定の評価タスク向けで、LLM-as-a-judgeの完全代替ではない。
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →