Back to explore

Jev、スクリーンショットなしでコンピュータ操作を実現

JevはローカルのCoreMLモデルで画面上のボタンやUI要素を分割し、オンデバイスOCRでラベルを読み取ることで、スクリーンショットやLLMなしにコンピュータ操作を実現します。

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all Jev gets.

· 1886 likesOpen on X