Jev rispetto ai modelli che generano testo
Jev non genera testo libero. Un video affiancato confronta decisioni parallele e generazione token per token.

Un utente ha dato a Jev due diff identici etichettati A e B chiedendo quale fosse migliore, scoprendo che ordine degli argomenti ed etichette influenzano il punteggio, con A intorno a 0,95 e risultati che cambiano invertendo l'ordine.
(2/5) hit some wild results. argument order and labels both move the score i gave jev 2 identical diffs (labelled A and B) and asked which is better. the results were shocking: fork A at 0.95 (five runs: 0.95, 0.96, 0.95, 0.96, 0.96) put B first with the same names and A is