Back to explore

Kod zevki görevlerinde Jev ve Codex Luna'nın değerlendirici olarak karşılaştırılması

Yazar, FrontierCode tarzı kod zevki görevlerinde Jev ve Codex Luna'yı değerlendirici olarak karşılaştırmak için metin tabanlı rubrikleri yapılandırılmış yanıtlara ayrıştıran bir oyuncak boru hattı çalıştırdı.

ran a toy pipeline comparing jev vs codex luna as graders on FrontierCode style "code taste" tasks: decomposed text-based rubrics into structured answers (e.g. are all changes necessary? do all changes consistently use the abstractions required by the task?) -> run both models in

· 0 likesOpen on X