Researchers have developed a novel cross-modal feature embedding method to improve the association between facial images and voice clips. This technique addresses the heterogeneity in audio-visual features, which often leads to inaccuracies in previous methods. By embedding voice and face features within a convex hull and utilizing cross-modal attention, the proposed approach significantly reduces false positives and negatives. Experiments on the VoxCeleb dataset show notable improvements in cross-modal verification, matching, and retrieval tasks compared to existing state-of-the-art methods. AI
IMPACT This research could lead to more accurate and robust systems for tasks requiring the matching of faces and voices, such as speaker verification and identification.
RANK_REASON The cluster contains two identical arXiv papers detailing a new research method.
- arXiv
- CORE Recommender
- Face and Voice Cross-modal Association with Learning Convex Feature Embedding
- Hugging Face
- Voxceleb: Large-scale speaker verification in the wild
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →