Back to explore

Jev, WebMCP 벤치마크에서 100% 작업 완료 달성

Jev와 Mercury 2.5가 WebMCP를 사용하여 벤치마크에서 100%의 작업을 해결했으며, 코드 실행을 사용하는 GPT-6 Astra보다 모델 비용이 약 112배 낮습니다.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev, WebMCP 벤치마크에서 100% 작업 완료 달성 1
· 2천 likesOpen on X