Back to explore

JevがWebMCPベンチマークで100%のタスク完了を達成

JevとMercury 2.5がWebMCPを使用し、ベンチマークで100%のタスクを解決。コード実行を伴うコンピュータ使用のGPT-6 Astraと比較してモデルコストが約112倍低い。

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

JevがWebMCPベンチマークで100%のタスク完了を達成 1
· 2023 likesOpen on X