Back to explore

Jev, WebMCP 벤치마크에서 뛰어난 성과

WebMCP 벤치마크에서 Jev 단독으로 raw DOM을 처리해 25/49개 작업을 완료했습니다. WebMCP 도구 선택과 Mercury 2.5의 인자 생성을 결합하면 49/49를 달성했으며 모델 비용은 GPT-6 Astra보다 약 112배 낮습니다.

🧪 The WebMCP benchmark by @0xidanlevin is the best stress test so far. Jev alone on the raw DOM: 25/49 tasks. Jev picking WebMCP tools + Mercury 2.5 writing args: 49/49, ~112x lower model cost than GPT-6 Astra. https://webmcp.com/benchmark

Jev, WebMCP 벤치마크에서 뛰어난 성과 1
· 0 likesOpen on X