返回探索

Jev 在 WebMCP 基准测试中表现亮眼

在 WebMCP 基准测试中,Jev 单独处理原始 DOM 完成 25/49 任务;结合 WebMCP 工具选择与 Mercury 2.5 参数生成后达到 49/49,且模型成本远低于 GPT-6 Astra。

🧪 The WebMCP benchmark by @0xidanlevin is the best stress test so far. Jev alone on the raw DOM: 25/49 tasks. Jev picking WebMCP tools + Mercury 2.5 writing args: 49/49, ~112x lower model cost than GPT-6 Astra. https://webmcp.com/benchmark

Jev 在 WebMCP 基准测试中表现亮眼 1
· 0 次赞在 X 打开