返回探索

Jev在WebMCP基准上实现100%任务完成率

Jev与Mercury 2.5结合使用WebMCP,在基准测试中解决了100%的任务,模型成本比使用代码执行的GPT-6 Astra低约112倍。

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev在WebMCP基准上实现100%任务完成率 1
· 2023 次赞在 X 打开