PulseAugur
EN
LIVE 09:33:07

New dataset OSSL-v2 enhances reproducible video-to-music generation

Researchers have introduced the Open Screen Soundtrack Library version 2 (OSSL-v2), a reproducible dataset of 34,343 video clips totaling 246.4 hours, derived from public-domain films. This new corpus aims to address the reproducibility gap in video-to-music generation models, which are often trained on ephemeral YouTube links. The team also developed a method to use dialogue as a conditioning signal for video-to-music generation, enhancing existing models by incorporating dialogue tracks to better align music with on-screen speech. Their approach demonstrated improvements over state-of-the-art baselines when evaluated on both public-domain and commercial films. AI

IMPACT This research introduces a more robust dataset for training AI models in video-to-music generation, potentially improving the quality and emotional resonance of AI-generated soundtracks.

RANK_REASON The item describes a new dataset and a novel method for video-to-music generation presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New dataset OSSL-v2 enhances reproducible video-to-music generation

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Haven Kim, Zachary Novack, Julian McAuley, Hao-Wen Dong ·

    Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

    arXiv:2608.11576v1 Announce Type: cross Abstract: Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reproducibility gap: models are often trained on crawled …