The OpenMOSS-Team has released MOSS-SoundEffect-v2.0, a text-to-audio model capable of generating high-fidelity sound effects. This new version utilizes a Diffusion Transformer backbone trained with Flow Matching, a DAC VAE, and a Qwen3 text encoder. MOSS-SoundEffect-v2.0 can produce up to 30 seconds of audio at 48 kHz from natural language prompts in both English and Chinese, covering a range of categories including environments, creatures, and human actions. AI
IMPACT Enables generation of diverse sound effects from text, potentially impacting content creation and audio design tools.
RANK_REASON Release of a new model with technical details on architecture and training. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- DAC VAE
- diffusers
- Diffusion Transformer
- Flow Matching
- MOSS-SoundEffect-v2.0
- MOSS-TTS
- OpenMOSS-Team/MOSS-SoundEffect-v2.0
- PyTorch
- Qwen3
- transformers
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →