Back to explore

Jevに視覚を与える:まずローカルOCR、その後モデルへ

エージェント設計でJevに視覚を組み込むパターンについての議論。画像を直接入力するのではなく、ローカルOCRで構造化テキストやオブジェクトを抽出してからJevに渡す方法。

Keno's breakdown of giving Jev eyes is a small example of a much bigger pattern in agent design. Jev can't see images. So the fix isn't to bolt vision onto a decision model, it's to run local OCR first, pull structured text and objects out of the picture, and only then hand Jev

· 0 likesOpen on X