Back to explore

Moteur LLM local façon Jev : inférence en 40 ms dédiée au jugement

Taro L. Saito note que, comme Jev, un moteur LLM local atteint une inférence de 40 ms en abandonnant la génération de texte/JSON au profit du jugement, sur DGX Spark.

Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.

· 1,3 k likesOpen on X