Back to explore

코드 취향 작업에서 Jev와 Codex Luna의 평가자 비교

저자는 텍스트 기반 루브릭을 구조화된 답변으로 분해하는 토이 파이프라인을 실행하여 FrontierCode 스타일 코드 취향 작업에서 Jev와 Codex Luna의 평가 성능을 비교했다.

ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in

· 0 likesOpen on X