Jev versus modelos de geração de texto
Jev não gera texto livre. Um vídeo lado a lado compara decisões paralelas com geração de texto token por token.

Taro L. Saito observa que, como o Jev, um motor LLM local alcança inferência de 40 ms ao abandonar a geração de texto/JSON e se especializar em julgamento, rodando no DGX Spark.
Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.