TypeSafe AI, 인용 검사 도구에 추가
사용자가 인용 검사기에 @typesafeai를 추가하고 호평했으며 PaperTrellis에서 무료로 사용 가능. Jev와 의료 AI 태그 포함. Link
동일한 430건 테스트에서 증거를 보여줬을 때 Jev는 68~92%, Laya는 84~88%의 허위 에이전트 보고를 통과시켰습니다. 직접적인 질문을 사용하자 Jev는 약 1%만 통과시켰습니다. 질문 설계가 모델보다 중요합니다.
Same 430-case test, two AI judges. Shown the evidence, Jev passed 68 to 92% of the false agent reports and Laya 84 to 88%. Asked one literal question instead, Jev let through about 1% of wrong counts and files. The question matters more than the model. https://cejel.dev/experiments