Tencent has introduced WeMM-Embedding, a new family of universal multimodal embedding models designed to represent diverse content like text, images, and videos in a shared space. Available in 2B, 4B, and 9B variants, these models are trained in two stages and have demonstrated state-of-the-art performance on public benchmarks, with the 2B variant outperforming previous 8B open-source models on MMEB-v2. WeMM-Embedding has already been deployed across various WeChat applications, including recommendation and search services, showing significant improvements in practical performance. AI
IMPACT Enhances multimodal AI capabilities, potentially improving search, recommendation, and agentic systems across various platforms.
RANK_REASON The item describes a new multimodal embedding model family with technical details and benchmark results, released by a major tech company. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →