Multimodal foundation model for image and video understanding from Microsoft

https://huggingface.co/microsoft/Mage-VL

Comments