Researchers have developed Omni-Embed-Mini, a compact omni-modal embedding model designed to integrate multiple data types without compromising text retrieval quality. The model, available in 0.9B and 2.3B parameter versions, maps text, speech, audio, images, and video into a shared space by training media encoders to match the embeddings of dense captions generated from the media itself. This approach preserves the original text embedding capabilities while adding new modalities, making the smaller version suitable for on-device applications and competitive with larger, closed-source models. AI
IMPACT Enables more efficient multi-modal AI systems by reducing model size without sacrificing text retrieval performance.
RANK_REASON The cluster describes a new research paper detailing a novel model architecture and training methodology for multi-modal embeddings.
- arXiv
- BEIR-8
- Gemini
- Gemini Embedding 2
- Hugging Face
- LoRA+
- Mohamed bin Zayed University of Artificial Intelligence
- Mohammed Irfan Kurpath
- MTEB-v2
- Omni-Embed-Mini
- Omni-Embed-Mini-0.9B
- SigLIP
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →