Tencent has released WeMM-Embedding-9B, a universal multimodal embedding model built on the Qwen 3.5 architecture. This model is capable of processing text, images, videos, and visual documents, generating a 4,096-dimensional embedding. It is designed for integration with popular libraries like Transformers and Sentence Transformers, with usage examples provided for both. WeMM-Embedding-9B supports various input types, including interleaved multimodal inputs, and offers detailed instructions for implementation across different platforms such as Google Colab and Kaggle. The model requires specific installations of PyTorch, transformers, and other related utilities to function correctly. AI
IMPACT Enables developers to integrate multimodal embedding capabilities into their applications.
RANK_REASON Model release from a non-frontier lab with usage instructions. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →