Back to explore

Jev raggiunge il 100% di completamento attività sul benchmark WebMCP

Jev combinato con Mercury 2.5 utilizzando WebMCP ha risolto il 100% delle attività in un benchmark, con un costo del modello circa 112 volte inferiore rispetto a GPT-6 Astra che utilizza computer use con esecuzione di codice.

We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using

Jev raggiunge il 100% di completamento attività sul benchmark WebMCP 1
· 2K likesOpen on X