Back to explore

Jev ทำภารกิจสำเร็จ 100% บน WebMCP benchmark

Jev ร่วมกับ Mercury 2.5 โดยใช้ WebMCP แก้ไขงาน 100% ใน benchmark ด้วยต้นทุนโมเดลต่ำกว่าประมาณ 112 เท่าเมื่อเทียบกับ GPT-6 Astra ที่ใช้ computer use พร้อมการรันโค้ด

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev ทำภารกิจสำเร็จ 100% บน WebMCP benchmark 1
· 2K likesOpen on X