Budowanie klasyfikatora bazy kodu z Jev
Deweloper dzieli się budową klasyfikatora bazy kodu z Jev, sugerując, że może rozwiązać problem przekombinowanego kodu od agentów, i pyta, co testować dalej.
Braintrust przetestował Jev jako model sędziego: szybki, tani i konkurencyjny w ocenie groundedness, ale ustępuje modelom rozumującym w matematyce i kodzie. Przydatny do niektórych zadań ewaluacyjnych, nie zastępuje w pełni LLM-as-a-judge.
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →