Building a Codebase Classifier with Jev
A developer shares building a codebase classifier with Jev, suggesting it may solve overengineered code from agents, and asks what to test next.
Braintrust tested Jev as a judge model: fast, cheap, and competitive for groundedness judging, but behind reasoning models for math and code domains. Recommended for certain eval tasks, not a full replacement for LLM-as-a-judge.
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more →