Jev로 코드베이스 분류기 구축
개발자가 Jev로 코드베이스 분류기를 구축하고 에이전트가 생성하는 과도하게 설계된 코드 문제를 해결할 수 있다고 언급하며 다음 테스트 대상을 묻습니다.
저자는 텍스트 기반 루브릭을 구조화된 답변으로 분해하는 토이 파이프라인을 실행하여 FrontierCode 스타일 코드 취향 작업에서 Jev와 Codex Luna의 평가 성능을 비교했다.
ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in