Back to explore

Prototipo JEV multimodale: QA su immagini con bounding box

Un prototipo JEV multimodale addestrato su CUB-200-2011 con backbone Qwen2.5-VL, che risponde a domande sulle immagini con probabilità delle opzioni e bounding box a ~0,4 s per immagine.

Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…

· 0 likesOpen on X