Jev face aux modèles générateurs de texte
Jev ne génère pas de texte libre. Une vidéo côte à côte compare la décision parallèle avec la génération token par token.

Taro L. Saito note que, comme Jev, un moteur LLM local atteint une inférence de 40 ms en abandonnant la génération de texte/JSON au profit du jugement, sur DGX Spark.
Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.