Back to explore

So sánh Jev và Codex Luna làm giám khảo trong các tác vụ gu mã nguồn

Tác giả đã chạy một pipeline đồ chơi phân rã các rubric dạng văn bản thành câu trả lời có cấu trúc để so sánh Jev và Codex Luna làm giám khảo trong các tác vụ gu mã nguồn kiểu FrontierCode.

ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in

· 0 likesOpen on X