Creare un classificatore di codebase con Jev
Uno sviluppatore condivide la creazione di un classificatore di codebase con Jev, suggerendo che può risolvere codice sovraingegnerizzato dagli agenti, e chiede cosa testare dopo.
Braintrust ha testato Jev come modello giudice: veloce, economico e competitivo per il groundedness, ma indietro rispetto ai modelli di ragionamento in matematica e codice. Utile per certi task di valutazione, non sostituisce del tutto LLM-as-a-judge.
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →