Back to explore

Prototype JEV multimodal : questions d'image avec boîtes englobantes

Un prototype JEV multimodal entraîné sur CUB-200-2011 avec un backbone Qwen2.5-VL, répondant aux questions d'image avec probabilités d'option et boîtes englobantes à ~0,4 s par image.

Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…

· 0 likesOpen on X