Jev kontra modele generujące tekst
Jev nie generuje swobodnego tekstu. Film porównuje równoległe decyzje z generowaniem tekstu token po tokenie.

Użytkownik dał Jevowi dwa identyczne diffy oznaczone A i B i zapytał, który jest lepszy; kolejność argumentów i etykiety wpływają na ocenę, A utrzymuje się około 0,95, a wyniki zmieniają się przy zamianie kolejności.
(2/5) hit some wild results. argument order and labels both move the score i gave jev 2 identical diffs (labelled A and B) and asked which is better. the results were shocking: fork A at 0.95 (five runs: 0.95, 0.96, 0.95, 0.96, 0.96) put B first with the same names and A is