発表スレッドからJev技術ブログへ
Diogoの発表スレッドの最後には、公式技術ブログと早期アクセス入口へのリンクがあり、モデル原理や製品説明を確認できます。
Qwen2.5-VLバックボーンでCUB-200-2011を用いて訓練されたマルチモーダルJEVプロトタイプ。画像の質問に対し選択肢確率とバウンディングボックスを出力し、1画像あたり約0.4秒。
Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…