WeChat has introduced WeMM-Embedding, a new family of universal multimodal embedding models designed to represent text, images, videos, and interleaved multimodal inputs in a shared space. The models, available in 2B, 4B, and 9B variants, were trained in two stages and have demonstrated state-of-the-art performance on public benchmarks, with the 2B variant outperforming an 8B baseline on MMEB-v2. WeMM-Embedding has been deployed across various WeChat applications, including Channels, Official Accounts, and Moments, showing significant gains in recommendation and search functionalities. AI
IMPACT Sets new SOTA on multimodal benchmarks and enhances retrieval/recommendation capabilities within WeChat applications.
RANK_REASON The cluster describes a technical report detailing a new family of multimodal embedding models with benchmark results and release information.
Read on arXiv cs.IR (Information Retrieval) →
- arXiv
- Hugging Face
- MMEB-v2
- Moments
- Official Accounts
- tencent/WeMM-Embedding-2B
- WeChat Channels
- WeChat Vision team
- WeMM-Embedding
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →