返回探索

给 Jev 视觉:先本地 OCR 再交给模型

讨论在 agent 设计中将视觉能力赋予 Jev 的模式:不使用图像直接输入,而是先通过本地 OCR 提取结构化文本和对象,再传递给 Jev 处理。

Keno's breakdown of giving Jev eyes is a small example of a much bigger pattern in agent design. Jev can't see images. So the fix isn't to bolt vision onto a decision model, it's to run local OCR first, pull structured text and objects out of the picture, and only then hand Jev

· 0 次赞在 X 打开