Back to explore

نموذج JEV متعدد الوسائط: أسئلة الصور مع مربعات التحديد

نموذج JEV متعدد الوسائط مدرب على CUB-200-2011 مع backbone Qwen2.5-VL، يجيب على أسئلة الصور باحتمالات الخيارات ومربعات التحديد بمعدل ~0.4 ثانية لكل صورة.

Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…

· 0 likesOpen on X