Back to explore

เอนจิน LLM โลคัลสไตล์ Jev: ประมวลผล 40ms เฉพาะการตัดสิน

Taro L. Saito ระบุว่าเช่นเดียวกับ Jev เอนจิน LLM โลคัลบรรลุการประมวลผล 40ms ด้วยการทิ้งการสร้างข้อความ/JSON และมุ่งเฉพาะการตัดสิน ทำงานบน DGX Spark

Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.

· 1.3K likesOpen on X