Back to explore

Engine LLM cục bộ kiểu Jev: suy luận 40ms chỉ cho phán đoán

Taro L. Saito lưu ý rằng, giống như Jev, một engine LLM cục bộ đạt suy luận 40ms bằng cách từ bỏ tạo văn bản/JSON và chuyên về phán đoán, chạy trên DGX Spark.

Just like Jev, to think that a Local LLM engine capable of inference in 40ms is realized simply by abandoning text/json generation and specializing solely in judgment. It's running on DGX Spark, but anyway, it's incredibly fast.

· 1,3 N likesOpen on X