Jev versus text-generating models
Jev does not generate free-form text. A side-by-side video compares parallel decision-making with token-by-token text generation.

A user gave Jev two identical diffs labelled A and B and asked which is better, finding that argument order and labels both move the score, with A around 0.95 and results shifting when order is swapped.
(2/5) hit some wild results. argument order and labels both move the score i gave jev 2 identical diffs (labelled A and B) and asked which is better. the results were shocking: fork A at 0.95 (five runs: 0.95, 0.96, 0.95, 0.96, 0.96) put B first with the same names and A is