Back to explore

Jev đạt 100% hoàn thành nhiệm vụ trên benchmark WebMCP

Jev kết hợp với Mercury 2.5 sử dụng WebMCP đã giải quyết 100% nhiệm vụ trong một benchmark, với chi phí mô hình thấp hơn khoảng 112 lần so với GPT-6 Astra sử dụng computer use với thực thi mã.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev đạt 100% hoàn thành nhiệm vụ trên benchmark WebMCP 1
· 2 N likesOpen on X