Back to explore

JevBench がコンテキスト長評価を追加

JevBench ページに、実際の入力長ごとの精度、別の長文ポリシーストレス表示、報告されたコンテキスト制限が追加され、精度が長さとともにどう変化するかを示していますが、長さだけが原因と主張するものではありません。

Context length has its own section on the JevBench page now: accuracy by actual input length, a separate long-policy stress view, and reported context limits. It shows where accuracy moves with length. It does not claim length alone moved it. https://benchmarkheaven.com/jev-models

JevBench がコンテキスト長評価を追加 1
· 0 likesOpen on X