PulseAugur
EN
LIVE 18:33:32

WanSong model generates 5-minute songs with dual stems using diffusion

Researchers have introduced WanSong, a novel diffusion-based model designed for generating long-form, commercial-grade songs. This model directly produces high-fidelity, multilingual songs up to five minutes in length, outputting both vocals and background music in a single pass. WanSong's diffusion framework also facilitates faster inference through step-distillation and offers a streamlined process for fine-tuning and customization for editing tasks. AI

IMPACT This model could accelerate the development of AI-powered music creation tools and enhance the quality and efficiency of audio synthesis.

RANK_REASON The cluster contains a technical report detailing a new model for audio generation, which falls under research.

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

WanSong model generates 5-minute songs with dual stems using diffusion

COVERAGE [3]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    WanSong v1.0 Technical Report

    Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powe…

  2. arXiv cs.CV TIER_1 English(EN) · Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou ·

    WanSong v1.0 Technical Report

    arXiv:2607.14749v1 Announce Type: cross Abstract: Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address …

  3. arXiv cs.CV TIER_1 English(EN) · Jingren Zhou ·

    WanSong v1.0 Technical Report

    Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple…