Back to explore

Multimodaler JEV-Prototyp: Bild-QA mit Bounding Boxes

Ein multimodaler JEV-Prototyp, trainiert auf CUB-200-2011 mit Qwen2.5-VL-Backbone, beantwortet Bildfragen mit Optionswahrscheinlichkeiten und Bounding Boxes bei ~0,4 s pro Bild.

Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…

· 0 likesOpen on X