Researchers have introduced the Open Screen Soundtrack Library version 2 (OSSL-v2), a reproducible dataset of 34,343 video clips totaling 246.4 hours, derived from public-domain films. This new corpus aims to address the reproducibility gap in video-to-music generation models, which are often trained on ephemeral YouTube links. The team also developed a method to use dialogue as a conditioning signal for video-to-music generation, enhancing existing models by incorporating dialogue tracks to better align music with on-screen speech. Their approach demonstrated improvements over state-of-the-art baselines when evaluated on both public-domain and commercial films. AI
IMPACT This research introduces a more robust dataset for training AI models in video-to-music generation, potentially improving the quality and emotional resonance of AI-generated soundtracks.
RANK_REASON The item describes a new dataset and a novel method for video-to-music generation presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Open Screen Soundtrack Library version 2
- OSSL-v2
- ScienceCast
- YouTube
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →