Tencent has released WeMM-Embedding-2B, a multimodal embedding model built on the Qwen 3.5 architecture. This model is capable of processing text, images, videos, and visual documents to generate 2,048-dimensional embeddings, though it does not support audio input. It is designed for integration with popular libraries such as Hugging Face's Transformers and sentence-transformers, with usage examples and installation instructions provided. AI
IMPACT Enables new multimodal applications by providing a versatile embedding model for text, images, and video.
RANK_REASON Release of a new multimodal embedding model with usage examples. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →