Researchers have developed the Douyin Multimodal Embedding (DME) model, a two-stage system designed for efficient and fine-grained multimodal search and recommendation. The model first undergoes large-scale contrastive pre-training to establish a unified embedding space, followed by a second stage that enhances semantic sufficiency through evidence-grounded latent reasoning and cross-conditional reconstruction. This approach allows DME to achieve state-of-the-art results on the MMEB-v2 benchmark, with its 9B variant scoring 78.4. In production on Douyin, DME has demonstrated a 2.92% relative gain in offline evaluations and a 0.1% lifetime gain in online A/B testing for search scenarios. AI
IMPACT This model's dual-stage training approach could influence future multimodal embedding architectures, balancing efficiency with fine-grained discrimination for large-scale applications.
RANK_REASON Technical report detailing a new multimodal embedding model with benchmark results.
- alphaXiv
- CatalyzeX
- DagsHub
- DME
- Douyin
- Douyin Multimodal Embedding Model
- Gotit.pub
- Hugging Face
- MMEB-V2
- ScienceCast
- Xiaohongshu
- YouTube
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →