Back to explore

Question Design Matters More Than the Model: Jev vs Laya Test

In the same 430-case test, Jev let through 68–92% of false agent reports and Laya 84–88% when shown evidence; after asking one literal question instead, Jev let through only about 1% of wrong counts and files. The question matters more than the model.

Same 430-case test, two AI judges. Shown the evidence, Jev passed 68 to 92% of the false agent reports and Laya 84 to 88%. Asked one literal question instead, Jev let through about 1% of wrong counts and files. The question matters more than the model. https://cejel.dev/experiments

Question Design Matters More Than the Model: Jev vs Laya Test 1
· 0 likesOpen on X