Back to explore

질문 설계가 모델보다 중요: Jev vs Laya 테스트

동일한 430건 테스트에서 증거를 보여줬을 때 Jev는 68~92%, Laya는 84~88%의 허위 에이전트 보고를 통과시켰습니다. 직접적인 질문을 사용하자 Jev는 약 1%만 통과시켰습니다. 질문 설계가 모델보다 중요합니다.

Same 430-case test, two AI judges. Shown the evidence, Jev passed 68 to 92% of the false agent reports and Laya 84 to 88%. Asked one literal question instead, Jev let through about 1% of wrong counts and files. The question matters more than the model. https://cejel.dev/experiments

질문 설계가 모델보다 중요: Jev vs Laya 테스트 1
· 0 likesOpen on X