Back to explore

Motore LLM locale in stile Jev: inferenza in 40 ms solo per il giudizio

Taro L. Saito osserva che, come Jev, un motore LLM locale raggiunge un'inferenza di 40 ms abbandonando la generazione di testo/JSON e specializzandosi nel giudizio, su DGX Spark.

Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.

· 1,3K likesOpen on X