返回探索

多模态 JEV 原型:图像问答与边界框

基于 Qwen2.5-VL 骨干网络在 CUB-200-2011 上训练的多模态 JEV 原型,可对图像问题输出选项概率和边界框,每张图约 0.4 秒。

Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…

· 0 次赞在 X 打开