Jev technical blog and early access
The end of Diogo’s launch thread links to the official technical blog and early-access entry point for further reading.
A multimodal JEV prototype trained on CUB-200-2011 with a Qwen2.5-VL backbone, answering image questions with option probabilities and bounding boxes at ~0.4s per image.
Built a multimodal JEV prototype that answers image questions with option probabilities and bounding boxes. Trained on CUB-200-2011 with Qwen2.5-VL backbone (0.4s/image) https://github.com/tin-xai/multimodal-jev-grounding…