Back to explore

Jev uppnår 100 % uppgiftslösning i WebMCP-benchmark

Jev i kombination med Mercury 2.5 som använder WebMCP löste 100 % av uppgifterna i ett benchmark, till cirka 112 gånger lägre modellkostnad än GPT-6 Astra med computer use och kodkörning.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev uppnår 100 % uppgiftslösning i WebMCP-benchmark 1
· 2 tn likesOpen on X