A new research paper published on arXiv explores how the structure of a contrastive corpus, rather than its sheer volume, dictates what attributes an audio embedding model learns. The study found that adding a lexical-speech component to a frozen multimodal embedding model significantly improved zero-shot keyword spotting but degraded speech-emotion recognition. Further experiments indicated that fine-tuning on a corpus with controlled prosody could recover emotion recognition, suggesting that the way attributes are encoded is dependent on the in-batch negatives and the separating signals available within the corpus structure. AI
RANK_REASON The cluster contains a research paper published on arXiv detailing findings about audio embedding models. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX Code Finder for Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →