Back to explore

مقارنة Jev وCodex Luna كمقيّمين في مهام ذوق الكود

شغّل المؤلف خط أنابيب تجريبيًا يفكك معايير التقييم النصية إلى إجابات منظمة لمقارنة Jev وCodex Luna كمقيّمين في مهام ذوق الكود بأسلوب FrontierCode.

ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in

· 0 likesOpen on X