Back to explore

Jev, WebMCP Benchmark'ında %100 Görev Tamamlama Elde Etti

Jev, Mercury 2.5 ile birlikte WebMCP kullanarak bir benchmark'ta görevlerin %100'ünü çözdü; kod yürütmeli computer use kullanan GPT-6 Astra'ya göre yaklaşık 112 kat daha düşük model maliyetiyle.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev, WebMCP Benchmark'ında %100 Görev Tamamlama Elde Etti 1
· 2 B likesOpen on X