Back to explore

Jev vs Gemini:300のカスタマーリクエスト実測

300のカスタマーリクエストで、有用な what if を尋ねることで両モデルが高速かつ正確になった。Geminiは依然として正確性で勝ったが、Jevはより多くの誤答を検出し、誤って拒否する正答が70%少なく、検証コストは約1/11だった。

6/7 So, across 300 customer requests: Asking the useful “what ifs” together made BOTH models faster and more accurate. Gemini still won on accuracy. Jev caught more wrong answers while wrongly rejecting 70% fewer correct ones. With verification, it cost about 1/11th as much

Jev vs Gemini:300のカスタマーリクエスト実測 1
· 0 likesOpen on X